applsci-logo

Journal Browser

Journal Browser

Research on Deep Learning for Advanced Image Processing and Computer Vision

A special issue of Applied Sciences (ISSN 2076-3417). This special issue belongs to the section "Computing and Artificial Intelligence".

Deadline for manuscript submissions: 20 September 2026 | Viewed by 1333

Editors


E-Mail Website
Guest Editor
1. Department of Engineering, University of Trás-os-Montes and Alto Douro, Vila Real, Portugal
2. INESC TEC, 5000-801 Vila Real, Portugal
Interests: computer vision; image and video processing; machine learning; artificial intelligence
Special Issues, Collections and Topics in MDPI journals

E-Mail Website
Guest Editor
1. School of Technology and Management, Polytechnic University of Leiria, Leiria, Portugal
2. Institute for Systems Engineering and Computers at Coimbra (INESC Coimbra), Coimbra, Portugal
Interests: computer vision and image processing; artificial intelligence and deep learning in health systems; medical image analysis; biosensors; sensor-based systems; industrial automation systems; Industry 4.0
Special Issues, Collections and Topics in MDPI journals

Special Issue Information

Dear Colleagues,

Computer vision and image processing have witnessed a paradigm shift with the advent of deep learning. While object identification remains a fundamental problem, modern algorithms' capabilities extend far beyond simple detection. Today, deep learning models are essential for a vast array of applications, ranging from autonomous driving, robotics, and augmented reality to medical diagnostics, remote sensing, and industrial inspection.

Deep learning-based methods, initially popularised by Convolutional Neural Networks (CNNs) and more recently by Vision Transformers (ViTs) and generative models (e.g., GANs, Diffusion Models), have revolutionised how we extract rich representations from visual data. These models successfully address challenges not only in accurately localising items but also in image classification, semantic and instance segmentation, image restoration, registration, and synthesis across complex and varied settings.

The Special Issue covers a broad spectrum of study areas, emphasising the creation of innovative architectures, feature extraction strategies, and training approaches, as well as the application of deep learning models to solve complex image processing challenges. We invite researchers to explore a range of designs and practical implementations, moving beyond traditional boundaries to encompass the whole pipeline of visual understanding. Furthermore, integrating attention mechanisms, transfer learning, and multimodal analysis is of significant interest for enhancing performance across diverse domains.

This Special Issue aims to present a thorough summary of current developments and new directions in deep learning for image processing, compiling original research and review articles on recent advances, technologies, solutions, practical applications, and novel challenges in this field.

Potential topics include, but are not limited to, the following:

  • Novel deep learning architectures for image analysis (CNNs, Vision Transformers, Graph Neural Networks);
  • Innovative applications of deep learning in medical imaging (CT, MRI, X-ray analysis, tumour detection);
  • Remote sensing and aerial imagery applications (satellite data analysis, drone/UAV surveillance, precision agriculture);
  • Image segmentation and classification (semantic, instance, and panoptic segmentation);
  • Image restoration and enhancement (super-resolution, denoising, deblurring, and colourisation);
  • Generative AI for Image Processing (GANs and diffusion models for synthesis and data augmentation);
  • Real-time image processing for autonomous vehicles and robotics;
  • Defect detection and quality control in industrial settings;
  • Challenges in dataset annotation, bias mitigation, and domain adaptation;
  • Edge computing and mobile deployment of vision models.

Dr. Sandra Pereira
Dr. António Manuel Trigueiros Da Silva Cunha
Dr. Paulo Jorge Coelho
Guest Editors

Manuscript Submission Information

Manuscripts should be submitted online at www.mdpi.com by registering and logging in to this website. Once you are registered, click here to go to the submission form. Manuscripts can be submitted until the deadline. All submissions that pass pre-check are peer-reviewed. Accepted papers will be published continuously in the journal (as soon as accepted) and will be listed together on the special issue website. Research articles, review articles as well as short communications are invited. For planned papers, a title and short abstract (about 250 words) can be sent to the Editorial Office for assessment.

Submitted manuscripts should not have been published previously, nor be under consideration for publication elsewhere (except conference proceedings papers). All manuscripts are thoroughly refereed through a single-anonymized peer-review process. A guide for authors and other relevant information for submission of manuscripts is available on the Instructions for Authors page. Applied Sciences is an international peer-reviewed open access semimonthly journal published by MDPI.

Please visit the Instructions for Authors page before submitting a manuscript. The Article Processing Charge (APC) for publication in this open access journal is 2400 CHF (Swiss Francs). Submitted papers should be well formatted and use good English. Authors may use MDPI's English editing service prior to publication or during author revisions.

Keywords

  • object detection
  • deep learning
  • computer vision
  • image processing
  • feature extraction

Benefits of Publishing in a Special Issue

  • Ease of navigation: Grouping papers by topic helps scholars navigate broad scope journals more efficiently.
  • Greater discoverability: Special Issues support the reach and impact of scientific research. Articles in Special Issues are more discoverable and cited more frequently.
  • Expansion of research network: Special Issues facilitate connections among authors, fostering scientific collaborations.
  • External promotion: Articles in Special Issues are often promoted through the journal's social media, increasing their visibility.
  • Reprint: MDPI Books provides the opportunity to republish successful Special Issues in book format, both online and in print.

Further information on MDPI's Special Issue policies can be found here.

Published Papers (2 papers)

Order results
Result details
Select all
Export citation of selected articles as:

Research

21 pages, 32395 KB  
Article
OSM-CLIP: Enhancing Remote Sensing Image–Text Representation Learning with OpenStreetMap Data
by Alessio Pierdominici, Riccardo Ricci, Mohammed Alruqimi and Farid Melgani
Appl. Sci. 2026, 16(14), 7002; https://doi.org/10.3390/app16147002 - 13 Jul 2026
Viewed by 366
Abstract
Remote sensing vision–language models, such as RemoteCLIP and GeoRSCLIP, have advanced image–text representation learning. However, they rely on manually curated caption datasets that are expensive to scale and provide only global image-level supervision. In this paper, we introduce OSM-CLIP, a framework that exploits [...] Read more.
Remote sensing vision–language models, such as RemoteCLIP and GeoRSCLIP, have advanced image–text representation learning. However, they rely on manually curated caption datasets that are expensive to scale and provide only global image-level supervision. In this paper, we introduce OSM-CLIP, a framework that exploits the freely available, continuously growing annotations of OpenStreetMap (OSM) to provide regionally scalable, patch-level supervision for remote sensing image-text learning. We construct a large-scale dataset of over 265,000 satellite images covering the contiguous United States, each automatically paired with fine-grained geographic annotations scraped from OSM and mapped to individual image patches. A contrastive loss operating at the patch level associates each image region with its corresponding OSM textual description, enabling the model to learn spatially grounded representations without any manual labeling effort. After fine-tuning on standard remote sensing captioning datasets, OSM-CLIP achieves an average improvement of 10.81% in zero-shot classification, 5.06% in text-to-image retrieval (R@1), and 3.87% in image-to-text retrieval (R@1) over existing methods across 13 classification and 4 retrieval benchmarks. Our results demonstrate that freely available geographic annotations can serve as a powerful source of supervision for remote sensing vision–language models in regions with high-quality OSM coverage. Full article
Show Figures

Figure 1

31 pages, 9766 KB  
Article
Benchmarking Conditional GANs in Industrial Marble Texture Synthesis via a Dual-Evaluation Framework
by António Alves de Campos, Margarida Figueiredo, Carlos M. A. Diogo, Gustavo Paneiro and Pedro Amaral
Appl. Sci. 2026, 16(8), 4028; https://doi.org/10.3390/app16084028 - 21 Apr 2026
Viewed by 453
Abstract
Deploying conditional Generative Adversarial Networks (cGANs) for industrial texture synthesis faces two barriers: the prohibitive cost of manual data annotation and the uncertain alignment between automated evaluation metrics and human perception. This study addresses both challenges for marble texture synthesis using 289 high-resolution [...] Read more.
Deploying conditional Generative Adversarial Networks (cGANs) for industrial texture synthesis faces two barriers: the prohibitive cost of manual data annotation and the uncertain alignment between automated evaluation metrics and human perception. This study addresses both challenges for marble texture synthesis using 289 high-resolution industrial scans. We adapt an unsupervised segmentation pipeline combining Simple Linear Iterative Clustering (SLIC) superpixels, Gaussian Mixture Models (GMMs), and graph cut optimization to extract vein structures without manual annotation. Four cGAN architectures—baseline cGAN, Pix2Pix, BicycleGAN, and GauGAN—are benchmarked using a dual-evaluation protocol contrasting ten automated metrics with structured human-centered assessment. The results reveal a significant metric–perception discrepancy. Pix2Pix achieved the best Fréchet Inception Distance (FID = 85.3) yet received the lowest human ratings due to periodic texture artifacts. GauGAN produced textures statistically indistinguishable from real marble, achieving a Visual Turing Pass Rate (VTPR) of 0.533 and a Mean Opinion Score on Marble Authenticity (MOS-MA) of 2.89, despite an inferior FID (87.3). These findings make three contributions: an annotation-free segmentation pipeline, empirical evidence that automated metrics alone are insufficient for architecture selection, and a dual-evaluation framework that establishes human-in-the-loop assessment as essential for quality-critical industrial deployment. Full article
Show Figures

Figure 1

Back to TopTop