-
A Method for Paired Comparisons of Glo Germ Quantity in Images of Hands Before and After Washing -
Artificial Intelligence in Pulmonary Endoscopy: Current Evidence, Limitations, and Future Directions -
AI-Based Osteoporosis Detection on Dental Radiographs -
Implementation of Image-Based AI Is Associated with Increased Case Volume in a High-Acuity, 15-Room Cardiothoracic Suite at a Tertiary Academic Hospital
Journal Description
Journal of Imaging
Journal of Imaging
is an international, multi/interdisciplinary, peer-reviewed, open access journal of imaging techniques, published online monthly by MDPI.
- Open Accessfree for readers, with article processing charges (APC) paid by authors or their institutions.
- High Visibility: indexed within Scopus, ESCI (Web of Science), PubMed, PMC, dblp, Inspec, Ei Compendex, and other databases.
- Journal Rank: JCR - Q2 (Imaging Science and Photographic Technology) / CiteScore - Q1 (Radiology, Nuclear Medicine and Imaging)
- Rapid Publication: manuscripts are peer-reviewed and a first decision is provided to authors approximately 21.3 days after submission; acceptance to publication is undertaken in 3.6 days (median values for papers published in this journal in the first half of 2026).
- Recognition of Reviewers: reviewers who provide timely, thorough peer-review reports receive vouchers entitling them to a discount on the APC of their next publication in any MDPI journal, in appreciation of the work done.
Impact Factor:
3.8 (2025);
5-Year Impact Factor:
3.6 (2025)
Latest Articles
From Pixel Modification to Generative Synthesis: A Survey of Deep Learning for Image Data Hiding
J. Imaging 2026, 12(9), 417; https://doi.org/10.3390/jimaging12090417 - 4 Sep 2026
Abstract
This survey presents a structured review of deep learning-based techniques for image data hiding, proposing a three-paradigm taxonomy organized by the method’s operational relationship to the carrier image. We classify existing methods into modification-based, synthesis-based, and logic-based approaches. In the modification-based tier, we
[...] Read more.
This survey presents a structured review of deep learning-based techniques for image data hiding, proposing a three-paradigm taxonomy organized by the method’s operational relationship to the carrier image. We classify existing methods into modification-based, synthesis-based, and logic-based approaches. In the modification-based tier, we trace the architectural progression from foundational Convolutional Neural Networks and Generative Adversarial Networks to high-capacity Invertible Neural Networks and Transformers, analyzing their distinct trade-offs between embedding capacity, imperceptibility, and robustness. In the synthesis-based tier, we examine how Diffusion Probabilistic Models and generative adversarial frameworks reframe data hiding as a carrier generation problem rather than a pixel editing task. This paradigm encompasses both generative steganography (where carriers are synthesized from scratch) and proactive watermarking (where provenance is embedded during AI content generation). In the logic-based tier, we review zero-watermarking and coverless steganography, where ownership is established through feature extraction and semantic mapping without modifying any image, a critical property for sensitive domains such as medical imaging. Finally, we identify four persistent infrastructure gaps: benchmarking fragmentation, narrow robustness evaluation, domain generalization failures, and computational infeasibility that prevent real-world deployment despite architectural progress, and we propose concrete research directions.
Full article
(This article belongs to the Section Image and Video Processing)
Open AccessArticle
VentrEX: An Anatomically Guided Deep Learning Pipeline for Ventricular Segmentation in Cine Cardiac MRI
by
Abla Bedoui, Julieta Anahí Rancati, Ignacio Lugones and Mohammed Cherkaoui
J. Imaging 2026, 12(9), 416; https://doi.org/10.3390/jimaging12090416 - 3 Sep 2026
Abstract
Automated segmentation of the left and right ventricles (LVs and RVs) in cine cardiac MRI (CMR) underpins reliable volumetry and mass estimation. However, papillary muscles and trabeculae (PM/T) introduce clinically meaningful variability and exacerbate cross-dataset domain shift. We present VentrEX, an anatomically guided
[...] Read more.
Automated segmentation of the left and right ventricles (LVs and RVs) in cine cardiac MRI (CMR) underpins reliable volumetry and mass estimation. However, papillary muscles and trabeculae (PM/T) introduce clinically meaningful variability and exacerbate cross-dataset domain shift. We present VentrEX, an anatomically guided pipeline. The core segmenter, VentrEX-Seg, is a 3D encoder-decoder with parallel channel-spatial attention and a Transformer bottleneck. Training is performed exclusively on ACDC. A lightweight PM/T module automatically extracts papillary and trabecular burdens and standardizes cavity volumes. External evaluation is zero-shot (no fine-tuning) on Sunnybrook (LV) and MM-WHS MRI (RV). We report Dice, HD95 (mm); for volumetry, we use Bland-Altman analyses (LV and RV volumes). Attention/Grad-CAM visualizations support interpretability. On ACDC, VentrEX achieved higher Dice and lower boundary error than U-Net, nnU-Net, CBAM, and VentrEX-Seg. Zero-shot performance was preserved externally (e.g., Sunnybrook LV Dice 0.9053, HD95 4.95 mm; MM-WHS RV Dice 0.9236, HD95 6.61 mm). Patient-level Bland–Altman analyses characterized LV and RV volumetric agreement. Qualitative overlays and 3D reconstructions showed fewer PM/T “leaks” and anatomically plausible borders across ED/ES. Single-source training with dual zero-shot external tests demonstrates robustness under domain shift. The combination of parallel attention and a Transformer bottleneck enables accurate, transparent cine-CMR segmentation across datasets.
Full article
(This article belongs to the Special Issue Advancing Magnetic Resonance Imaging: Emerging Technologies, Computation, and Clinical Applications)
Open AccessArticle
Hybrid Fusion of Time–Intensity Curve, Deep Learning, and Fractional Zernike–Caputo Features for Accurate Liver Lesion Classification in DCE-MRI
by
Ali M. Hasan, Noor K. N. Al-Waely, Wallaa L. Alfalluji, Rabha W. Ibrahim, Hamid A. Jalab and Farid Meziane
J. Imaging 2026, 12(9), 415; https://doi.org/10.3390/jimaging12090415 - 3 Sep 2026
Abstract
Early and accurate detection of liver masses is essential for effective clinical management and improved patient outcomes, as treatment strategies differ significantly between benign and malignant lesions. Dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) is widely used for liver lesion evaluation due to its
[...] Read more.
Early and accurate detection of liver masses is essential for effective clinical management and improved patient outcomes, as treatment strategies differ significantly between benign and malignant lesions. Dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) is widely used for liver lesion evaluation due to its ability to capture temporal enhancement behavior. However, reliable interpretation remains challenging because many lesions exhibit similar enhancement patterns, leading to diagnostic uncertainty and potential misclassification. This study proposes a hybrid diagnostic framework that integrates multiple complementary feature representations for automated liver mass classification. Specifically, the model extracts time–intensity curve (TIC) characteristics to capture contrast dynamics, deep learning features to represent complex spatial patterns, and fractional Zernike–Caputo descriptors to encode advanced mathematical shape and texture information. These heterogeneous feature sets are subsequently fused to form a unified and discriminative representation of each lesion. The proposed study aims to enhance the differentiation between benign and malignant liver masses by leveraging the strengths of kinetic, data-driven, and fractional mathematical descriptors. Experimental results demonstrate that the model achieves superior diagnostic performance, reaching an accuracy of 94.05% on a dataset comprising 645 dynamic MRI cases. Overall, the proposed approach provides a robust and efficient tool for liver lesion characterization, with potential to support clinical decision-making and reduce diagnostic ambiguity in medical imaging practice.
Full article
(This article belongs to the Section Medical Imaging)
►▼
Show Figures

Figure 1
Open AccessArticle
Physics-Preserving Attention-Guided Artifact Removal for Spaceborne Optical Images
by
Shuxiang Cai, Zuoxun Hou, Haian Zhou, Zheng Pan and Dong Wang
J. Imaging 2026, 12(9), 414; https://doi.org/10.3390/jimaging12090414 - 3 Sep 2026
Abstract
Artifacts, including halos and saturated bright spots, are common in spaceborne optical images and can degrade the reliability of star extraction, space object detection, and photometric analysis. Traditional signal-processing methods, such as morphological filtering and low-rank decomposition, rely on fixed priors that may
[...] Read more.
Artifacts, including halos and saturated bright spots, are common in spaceborne optical images and can degrade the reliability of star extraction, space object detection, and photometric analysis. Traditional signal-processing methods, such as morphological filtering and low-rank decomposition, rely on fixed priors that may fail under complex artifact morphologies. Deep restoration networks can improve visual quality but do not explicitly enforce radiometric consistency. Multimodal instruction-driven editing models provide semantic localization capability, but probabilistic diffusion resampling can introduce uncontrolled pixel changes in non-target regions, compromising pixel-level physical consistency. We refer to this problem as editing-induced radiometric drift. To address this problem, we propose PARE, a physics-preserving attention-guided artifact removal framework for spaceborne optical images. Instead of directly using the edited image as the final restoration, PARE treats it as a candidate restoration and derives artifact-region constraints from the image-to-text cross-attention sub-block of the MM-DiT joint attention matrix. These constraints are combined with multi-scale fusion to restrict generative modification to localized artifact regions, thereby enabling artifact suppression while reducing unintended changes in non-target regions. Experiments on simulated and real spaceborne optical image datasets show that PARE achieves effective artifact suppression while improving radiometric preservation. On real on-orbit images, PARE reaches 92.95% artifact mean reduction (AMR) and 98.57% artifact energy reduction (AER), reduces the outer-region mean absolute error (O-MAE) to 0.54, and improves the outer-region structural similarity index (O-SSIM) to 0.997. It also yields the smallest background shifts among all compared methods. These results indicate that PARE provides a favorable trade-off between artifact suppression and radiometric fidelity and offers a practical way to apply generative models to scientific imaging tasks that require pixel-level physical consistency.
Full article
(This article belongs to the Section AI in Imaging)
►▼
Show Figures

Figure 1
Open AccessArticle
DGF-YOLO: A Degradation-Guided Feature Enhancement Method for Small-Scale Pedestrian Detection in UAV Images
by
Boyu Wang, Jingguo Lv, Shuwei Huang and Yingqi Bai
J. Imaging 2026, 12(9), 413; https://doi.org/10.3390/jimaging12090413 - 2 Sep 2026
Abstract
Small-scale pedestrians in UAV imagery often exhibit limited pixel coverage, weak texture, and severe background interference, while progressive network downsampling can further degrade their short-side structures and increase missed detections. To address this problem, we propose DGF-YOLO, a degradation-guided feature enhancement method built
[...] Read more.
Small-scale pedestrians in UAV imagery often exhibit limited pixel coverage, weak texture, and severe background interference, while progressive network downsampling can further degrade their short-side structures and increase missed detections. To address this problem, we propose DGF-YOLO, a degradation-guided feature enhancement method built on YOLOv12n. The method introduces a degradation-level criterion to identify the feature stage at which a pedestrian first undergoes significant structural degradation and, based on the resulting statistics, incorporates a high-resolution P2 detection head. It further employs a Directional Structure-Aware module to enhance local, horizontal, and vertical structural cues through adaptive multi-branch fusion, a Degradation-Guided Attention module to learn a degradation guidance map under explicit supervision and reweight degradation-sensitive regions, and a Fine-Grained Structure Preservation module to retain local contours and contextual details using depthwise and dilated convolutions. On the single-class pedestrian detection task constructed from VisDrone2019-DET, DGF-YOLO achieves 70.4% precision, 51.8% recall, 59.7% mAP50, and 26.9% mAP50-95, improving the YOLOv12n baseline by 6.6, 7.4, 10.3, and 6.5 percentage points, respectively. The results suggest that the proposed feature enhancement strategy helps reduce missed detections associated with structural degradation in small-scale pedestrians.
Full article
(This article belongs to the Special Issue AI-Driven Image Analysis and Pattern Recognition)
►▼
Show Figures

Figure 1
Open AccessArticle
A Text-Guided Lesion Mining Vision–Language Ordinal Classification Framework for Diabetic Retinopathy Grading
by
Jiawen Huang and Jinxia Shang
J. Imaging 2026, 12(9), 412; https://doi.org/10.3390/jimaging12090412 - 1 Sep 2026
Abstract
Diabetic retinopathy (DR) is a major retinal disease that can cause visual impairment and irreversible blindness. Accurate automated DR grading is essential for large-scale screening and timely clinical intervention. However, most existing methods rely primarily on visual features for classification. Moreover, they often
[...] Read more.
Diabetic retinopathy (DR) is a major retinal disease that can cause visual impairment and irreversible blindness. Accurate automated DR grading is essential for large-scale screening and timely clinical intervention. However, most existing methods rely primarily on visual features for classification. Moreover, they often overlook the ordinal structure of DR severity and the intra-class phenotypic heterogeneity arising from diverse lesion combinations. To address these issues, based on the semantic prior information provided by RetiZero, we propose a text-guided lesion mining vision–language ordinal classification framework for DR grading. The proposed framework introduces a text-guided cross-layer lesion mining module that exploits semantic response differences between normal-tissue and lesion-related textual prompts, thereby guiding multi-level visual patch features toward lesion regions relevant to DR grading. To explicitly model the ordered progression of DR severity, we design a conditional ordinal regression branch and an ordinal distribution alignment strategy that jointly encourage the predictions to follow the inherent order of DR grades. Moreover, we introduce a multi-center feature constraint to capture diverse intra-grade phenotypic patterns and enhance feature discriminability. Experiments on APTOS 2019 show that the proposed method achieves 86.3% accuracy, 90.6% AUC, and 70.9% Macro-F1, which improved by 2.4, 0.7, and 5.5 percentage points compared to RetiZero. Furthermore, under the standardized leave-one-domain-out protocol of GDRNet, the proposed method achieves the highest reported average accuracy of 58.6% across six public DR datasets, exceeding the reported result of PAF (54.6%) by 4.0 percentage points. Nevertheless, our approach is limited in F1 and AUC metrics, for which GDRNet delivers superior performance. These results suggest that the proposed framework can improve DR grading performance and the cross-dataset generalization ability of the model to a certain extent.
Full article
(This article belongs to the Section Medical Imaging)
►▼
Show Figures

Figure 1
Open AccessArticle
Task-Specific Detector Adaptation for Edge MOT: Tracking and Deployment Trade-Offs on the NVIDIA Jetson Nano
by
Bruna de Vargas Guterres, Juan Pedro de León, Pablo D. Cuña, Víctor Castelli, Silvia Silva da Costa Botelho and Marcelo Rita Pias
J. Imaging 2026, 12(9), 411; https://doi.org/10.3390/jimaging12090411 - 1 Sep 2026
Abstract
Although several MOT solutions have been proposed, limited evidence is available regarding how task-specific detector adaptation affects tracking quality and deployment requirements on low-cost edge hardware. These effects have not been jointly evaluated under fixed tracking and deployment conditions. This work evaluates MOT
[...] Read more.
Although several MOT solutions have been proposed, limited evidence is available regarding how task-specific detector adaptation affects tracking quality and deployment requirements on low-cost edge hardware. These effects have not been jointly evaluated under fixed tracking and deployment conditions. This work evaluates MOT on a 4 GB NVIDIA Jetson Nano using the MOT17 benchmark. Experiment 1 served as a baseline characterization and detector-selection stage. YOLOv8n, SSDLite320 and Faster R-CNN were evaluated with a fixed OC-SORT configuration under the same tracking framework. Based on the observed trade-offs, YOLOv8 was selected for adaptation. A task-specific fine-tuning stage was subsequently performed for YOLOv8n and YOLOv8s using more than 30,000 annotated pedestrian image records. Experiment 2 constituted the main analysis. Baseline and adapted models were evaluated under the same tracking and deployment protocol. Tracking performance was assessed through MOTA, IDF1, and HOTA. Processing throughput was measured together with system RAM usage. Board power consumption and energy per frame were also measured. In the paired YOLOv8n comparison, MOTA increased from 15.95 to 51.09 after adaptation. Throughput changed from 3.6 to 3.5 FPS. System RAM usage changed from 3.3820 to 3.3925 GB. Average board power changed from 4235 to 4227 mW. Energy per frame increased by approximately 2.7%. The adapted YOLOv8s model achieved a MOTA of 55.88 at 2.5 FPS. None of the evaluated pipelines achieved conventional real-time video throughput. These findings indicate that task-specific adaptation improved tracking performance with limited changes in the measured deployment characteristics within the evaluated configuration.
Full article
(This article belongs to the Section Computer Vision and Pattern Recognition)
►▼
Show Figures

Figure 1
Open AccessArticle
CSP-UNet: A Lightweight Network for Hand X-Ray Image Segmentation
by
Hai Wang, Jiale Gu, Junhao Wen and Chunlai Yang
J. Imaging 2026, 12(9), 410; https://doi.org/10.3390/jimaging12090410 - 1 Sep 2026
Abstract
Hand X-ray image segmentation is an important step in automated radiographic image analysis. However, conventional U-shaped segmentation networks often have relatively high model complexity, while variations in grayscale distributions across hand X-ray images may affect segmentation performance. To address these issues, this study
[...] Read more.
Hand X-ray image segmentation is an important step in automated radiographic image analysis. However, conventional U-shaped segmentation networks often have relatively high model complexity, while variations in grayscale distributions across hand X-ray images may affect segmentation performance. To address these issues, this study proposes a lightweight hand X-ray image segmentation network, CSP-UNet (Cross-Stage Partial U-Net). The network integrates cross-stage partial feature processing into the U-Net encoder–decoder framework to reduce redundant feature computation and the number of model parameters while preserving effective feature representation. In addition, an adaptive Gaussian histogram-matching strategy is employed to reduce variations in grayscale distributions across X-ray images. CSP-UNet was evaluated on a dataset comprising 2000 hand X-ray images and compared with Classic U-Net, Res-UNet, Attention U-Net, and Swin U-Net. Experimental results show that CSP-UNet maintained comparable segmentation performance, achieving a Dice coefficient of 0.9927, PA of 0.9839, MPA of 0.9810, and mIoU of 0.9542, while requiring only 20.01 M parameters. Compared with Classic U-Net, CSP-UNet maintained a comparable Dice coefficient (0.9927 vs. 0.9910) while reducing the parameter count from 69.1 M to 20.01 M, corresponding to a reduction of approximately 71.04%. These results indicate that CSP-UNet maintains comparable segmentation performance while substantially reducing model complexity, offering a favorable trade-off between segmentation performance and model size for hand X-ray image segmentation.
Full article
(This article belongs to the Topic Applications of Image and Video Processing in Medical Imaging)
►▼
Show Figures

Figure 1
Open AccessArticle
Reproducible Semi-Automated Quantification of Vascularization in Bone Sections Using CD31 Immunohistochemistry and Trainable Weka Segmentation
by
Nick Mattern, Holger Freischmidt, Matthias Schulte, Alma Aubert, Sanja Kalmus, Jan Makogon, Paul Alfred Grützner, Jonas Armbruster and Felix Lamadé-Dootz
J. Imaging 2026, 12(9), 409; https://doi.org/10.3390/jimaging12090409 - 1 Sep 2026
Abstract
Quantitative assessment of vascularization is important in bone regeneration research, but CD31-immunohistochemically stained sections are often evaluated manually or semi-quantitatively, limiting reproducibility and comparability. The aim of this study was to establish and validate a reproducible, open-source workflow for semi-automated quantification of CD31-positive
[...] Read more.
Quantitative assessment of vascularization is important in bone regeneration research, but CD31-immunohistochemically stained sections are often evaluated manually or semi-quantitatively, limiting reproducibility and comparability. The aim of this study was to establish and validate a reproducible, open-source workflow for semi-automated quantification of CD31-positive area fraction in bone sections using Fiji/ImageJ and Trainable Weka Segmentation (TWS). CD31-immunohistochemically stained rat bone sections from defect/regenerating tissue, femur, tibia, and spine were analyzed. The workflow combined standardized image acquisition, predefined regions of interest, pixel-based TWS classification, extraction of the CD31-positive class, and CD31-positive area normalized to tissue area (CD31.Ar/T.Ar). Manual reference measurements were performed by two independent observers in repeated runs. Manual CD31.Ar/T.Ar measurements showed good retest reliability, with mean coefficients of variation (CV) of 7.17% and 6.84% for observer 1 and observer 2, respectively, and good interobserver agreement. Independently trained TWS classifiers produced highly stable CD31.Ar/T.Ar values, with an overall mean CV of 1.98%. Manual assessment required 3:17 ± 1:35 min per section, whereas the TWS-based workflow separated an initial classifier training step from rapid repeated analysis of larger image sets. This study provides a transparent, reproducible, and time-efficient open-source workflow for semi-automated quantification of CD31-positive vascular area fraction in bone sections and supports its use as a scalable method for vascular histomorphometry in preclinical bone regeneration research.
Full article
(This article belongs to the Section Medical Imaging)
►▼
Show Figures

Figure 1
Open AccessArticle
Fusion of Radiomics and Gated Graph Attention Network for Pulmonary Nodule Malignancy Classification
by
Xinying Guo, Zirong Yu and Jibin Yin
J. Imaging 2026, 12(9), 408; https://doi.org/10.3390/jimaging12090408 - 30 Aug 2026
Abstract
Pulmonary nodule malignancy classification requires effective integration of heterogeneous imaging features and contextual information among nodules. Existing models mainly analyze nodules independently and may overlook inter-nodule relationships. We propose RGGA-Net, a radiomics-guided graph attention network that integrates deep imaging features and radiomics representations
[...] Read more.
Pulmonary nodule malignancy classification requires effective integration of heterogeneous imaging features and contextual information among nodules. Existing models mainly analyze nodules independently and may overlook inter-nodule relationships. We propose RGGA-Net, a radiomics-guided graph attention network that integrates deep imaging features and radiomics representations through cross-modal interaction and graph-based reasoning. A learned edge gate is introduced to adaptively modulate graph message passing, and an anchor regularization strategy is used to improve representation stability. RGGA-Net was evaluated on the LUNA25 dataset using patient-level splitting with an internal held-out test set and further assessed on the LIDC-IDRI cohort under cross-dataset evaluation. On the internal test set, RGGA-Net achieved an AUC of 0.8910 and a PR-AUC of 0.5173. External evaluation on LIDC-IDRI demonstrated moderate discrimination (AUC = 0.7023). Gate–Anchor ablation analysis showed a favorable interaction pattern between the learned gate and anchor regularization, although statistical superiority was not established. These findings suggest that radiomics-guided graph attention provides a feasible framework for incorporating inter-nodule information into malignancy classification, while further multi-center validation remains necessary.
Full article
(This article belongs to the Topic Applications of Image and Video Processing in Medical Imaging)
►▼
Show Figures

Figure 1
Open AccessArticle
Semantic Topological Multi-Scale Part Network for Fine-Grained Visual Classification
by
Xuerong Liu, Min Zhi, Yanjun Yin and Rula Sa
J. Imaging 2026, 12(9), 407; https://doi.org/10.3390/jimaging12090407 - 28 Aug 2026
Abstract
Fine-grained visual classification (FGVC) aims to distinguish highly similar subcategories, and its performance relies heavily on the accurate modeling of discriminative local parts and their structural relationships. However, existing Vision Transformer-based methods are susceptible to background noise interference, and the relationship modeling approach
[...] Read more.
Fine-grained visual classification (FGVC) aims to distinguish highly similar subcategories, and its performance relies heavily on the accurate modeling of discriminative local parts and their structural relationships. However, existing Vision Transformer-based methods are susceptible to background noise interference, and the relationship modeling approach relying on explicit spatial coordinates struggles to maintain stable structural representations when targets undergo pose variations and non-rigid deformations. To address these issues, this paper proposes a Semantic Topology Part Network (STP-Net). First, a Prior-Guided Part Aggregator (PGA) is designed, which leverages the foreground prior provided by foundation models to guide discriminative part discovery, enhancing target region responses while suppressing background interference. Second, a Topology-Informed Semantic Graph Convolutional Network (TIS-GCN) is designed to dynamically construct topological relationships among parts in an implicit semantic space, achieving robust modeling against complex structural variations. Furthermore, a Semantic–Spatial Cross-Attention (SSCA) mechanism is introduced to establish bidirectional interaction between semantic relationships and spatial features, and combined with a Global-Context Adaptive Gating mechanism to accomplish multi-scale feature fusion. On four mainstream fine-grained visual classification benchmarks, namely CUB-200-2011, Stanford Cars, Stanford Dogs, and NABirds, the proposed model achieves Top-1 accuracies of 92.7%, 94.9%, 95.2%, and 92.3%, respectively. Comprehensive ablation studies and visualization analyses further validate the effectiveness of the proposed method in background suppression, structural relationship modeling, and discriminative feature learning.
Full article
(This article belongs to the Section Computer Vision and Pattern Recognition)
►▼
Show Figures

Figure 1
Open AccessArticle
tgLang: A Domain-Specific Language for Geometry Processing and Computational Imaging Workflows
by
Vijai Kumar Suriyababu, Cornelis Vuik and Matthias Möller
J. Imaging 2026, 12(9), 406; https://doi.org/10.3390/jimaging12090406 - 27 Aug 2026
Abstract
Geometry-processing and computational-imaging workflows combine heterogeneous data structures, topology-changing edits, dense numerical fields, visualization, and repeated experimental variation. These workflows are often clear as algorithms but obscured in software by traversal boilerplate, representation conversions, build-system boundaries, and ad hoc scripting conventions. This paper
[...] Read more.
Geometry-processing and computational-imaging workflows combine heterogeneous data structures, topology-changing edits, dense numerical fields, visualization, and repeated experimental variation. These workflows are often clear as algorithms but obscured in software by traversal boilerplate, representation conversions, build-system boundaries, and ad hoc scripting conventions. This paper presents tgLang, a domain-specific language with explicit, runtime-enforced representation types that makes meshes, point clouds, curve networks, grids, two-dimensional images, and image stacks first-class executable values. The language combines manifest types, typed arrays, modules, deterministic parallel constructs, flow-oriented queries, and runtime-provided domain operations. Its current implementation uses a stack-based bytecode virtual machine for reference semantics and dispatches representation-heavy operations to optimized C++ kernels. The evaluation is organized around complete workflows: topological hole detection, distance-field-based mean camber line extraction, voxel downsampling of point clouds, curve-network generation, surface-mesh smoothing and remeshing, image-stack edge detection, morphological image processing, and image-stack surface extraction. These examples show that a domain-aware source language can express multi-representation geometry and imaging algorithms as compact, reproducible programs while preserving explicit representation choices and a path toward deployable implementations.
Full article
(This article belongs to the Section Computational Imaging and Computational Photography)
►▼
Show Figures

Figure 1
Open AccessArticle
Erasing and Refining Discriminative Features for CNV Subtype Classification in OCT Images
by
Jiayi Zhang, Qingbo Wang, Jiqiang Liu and Aixi Qu
J. Imaging 2026, 12(9), 405; https://doi.org/10.3390/jimaging12090405 - 27 Aug 2026
Abstract
Choroidal neovascularization (CNV) subtype classification from optical coherence tomography (OCT) images is clinically important because treatment response and disease prognosis vary across subtypes. However, automated classification remains challenging owing to subtle inter-class differences and considerable imaging noise. We propose an erasing–refining discriminative feature
[...] Read more.
Choroidal neovascularization (CNV) subtype classification from optical coherence tomography (OCT) images is clinically important because treatment response and disease prognosis vary across subtypes. However, automated classification remains challenging owing to subtle inter-class differences and considerable imaging noise. We propose an erasing–refining discriminative feature network (ERDF-Net) that mitigates noisy dominant activations and reveals fine-grained structural cues. The model perturbs salient regions to facilitate subtle feature learning, restores clean salient information, and fuses both representations via channel–spatial attention to form a coherent and discriminative embedding of CNV morphology. Experiments on a clinical OCT dataset show that ERDF-Net consistently surpasses state-of-the-art fine-grained and erasing-based methods across multiple metrics. Ablation and visualization analyses further confirm the benefit and interpretability of controlled salient suppression and refined feature fusion. ERDF-Net provides an effective and reliable solution for fine-grained medical image classification.
Full article
(This article belongs to the Special Issue Deep Learning for Image Analysis in Scientific and Engineering Applications)
►▼
Show Figures

Figure 1
Open AccessArticle
Hypergraph-Driven Heterogeneous Spatial Relationship Learning for Remote Sensing Segmentation
by
Qihao Zhang, Lankun Peng, Feiyang Hu and Xiaoming Xi
J. Imaging 2026, 12(9), 404; https://doi.org/10.3390/jimaging12090404 - 27 Aug 2026
Abstract
Remote sensing semantic segmentation is critical for extracting fine-grained spatial information in applications such as urban planning and environmental monitoring. However, existing methods face significant challenges in modeling heterogeneous spatial relationships within complex urban scenes, where semantically related regions are spatially dispersed yet
[...] Read more.
Remote sensing semantic segmentation is critical for extracting fine-grained spatial information in applications such as urban planning and environmental monitoring. However, existing methods face significant challenges in modeling heterogeneous spatial relationships within complex urban scenes, where semantically related regions are spatially dispersed yet functionally interdependent. Conventional convolutional neural networks exhibit limited receptive fields that fail to capture long-range dependencies, while Transformer-based approaches capture global dependencies but do not explicitly model regional heterogeneity, leading to blurred boundaries and category confusion. To address these limitations, this paper proposes a novel multi-relational-aware segmentation framework that leverages hypergraph theory to dynamically model higher-order semantic groupings across non-adjacent regions. The core innovation lies in a hypergraph structure learning unit that employs fuzzy clustering to partition multi-scale features into adaptive hyperedge sets, enabling joint representation of topological associations and functional dependencies among spatially distributed entities. Additionally, a multi-scale co-modeling strategy integrates stochastic feature masking with weighted fusion to bridge semantic abstraction and spatial localization. Experiments demonstrate that the proposed method achieves state-of-the-art mIoU performance on the LoveDA, Vaihingen, and Potsdam datasets, obtaining mIoU scores of 54.70%, 85.01%, and 87.64%, with improvements of 0.30%, 0.91%, and 0.08%, respectively, over the best existing methods.
Full article
(This article belongs to the Special Issue Deep Learning for Image Analysis in Scientific and Engineering Applications)
►▼
Show Figures

Figure 1
Open AccessArticle
A Deep Reconstruction Framework with Ringing Artifact Suppression for Overexposed Remote Sensing Image Restoration
by
Dinghao Yang, Yujie Xing, Hongmei Li, Xuquan Wang and Xiong Dun
J. Imaging 2026, 12(9), 403; https://doi.org/10.3390/jimaging12090403 - 26 Aug 2026
Abstract
Computational imaging shifts part of the aberration correction from optical hardware to algorithms, offering a viable path toward compact, simplified systems. However, overexposed regions—often caused by phenomena such as water-body reflections—can readily induce severe ringing artifacts in reconstructed images. To address this problem,
[...] Read more.
Computational imaging shifts part of the aberration correction from optical hardware to algorithms, offering a viable path toward compact, simplified systems. However, overexposed regions—often caused by phenomena such as water-body reflections—can readily induce severe ringing artifacts in reconstructed images. To address this problem, we propose a Ringing-perceptive Cooperative Reconstruction Network (RPCR-Net). This network integrates a learned Wiener filter and a field-of-view shared kernel prediction network (FOV-KPN) for feature extraction and innovatively incorporates a combined regularization mechanism that leverages a Local Maximum Gradient Prior and a multi-scale ringing measurement model within its loss function to suppress artifacts while preserving details. Validated on a constructed overexposed image dataset, RPCR-Net improves the Peak Signal-to-Noise Ratio (PSNR) from 29.08 dB to 37.06 dB and the Structural Similarity Index Measure (SSIM) from 0.8795 to 0.9549. Experiments on real-world scenes further confirm its capability to suppress ringing artifacts while maintaining visual quality. The proposed method can generate high-quality images such as image reconstruction and robustness improvement in optical systems.
Full article
(This article belongs to the Topic Image Processing, Signal Processing and Their Applications)
►▼
Show Figures

Figure 1
Open AccessArticle
Area-Driven Adaptive Sampling of Closed Droplet Contours for Vision-Based Droplet Observation
by
Xuefeng Wang, Yangting Zheng, Chenyao Bai, Yinqi Chen, Xiang Gao, Yiyue Li and Yunlong Zhu
J. Imaging 2026, 12(9), 402; https://doi.org/10.3390/jimaging12090402 - 26 Aug 2026
Abstract
Closed droplet contours provide the geometric basis for area estimation in vision-based droplet observation. In OLED inkjet printing, droplets are deposited into pixel wells with predefined geometry; projected area is therefore a primary geometric quantity for assessing whether the deposited liquid sufficiently fills
[...] Read more.
Closed droplet contours provide the geometric basis for area estimation in vision-based droplet observation. In OLED inkjet printing, droplets are deposited into pixel wells with predefined geometry; projected area is therefore a primary geometric quantity for assessing whether the deposited liquid sufficiently fills the well or risks overflow. This work formulates closed-contour sampling under a fixed sampling budget as an area-driven sampling problem. A leading-order analysis of the local arc–chord area error shows that the dominant cubic term depends jointly on curvature and segment length. Minimization of the resulting leading-order area-error functional yields an asymptotically optimal area-driven sampling density proportional to the cube root of curvature, together with a sampling-budget estimate under a target area-error tolerance. The derived sampling density is implemented on the fitted closed contour through cumulative-weight inversion. Experiments on random closed curves and the droplet dataset provide a systematic quantitative comparison with representative methods under identical fixed-budget settings, complemented by statistical analysis and evaluations of geometric fidelity, sensitivity, and computational efficiency. The proposed method achieves lower area estimation error under the tested sampling budgets, with the improvement being most pronounced at lower sampling budgets, while the reported geometric-fidelity metrics show no disproportionate degradation of contour fidelity. These results demonstrate the effectiveness of area-driven sampling for closed-contour area estimation under limited sampling budgets.
Full article
(This article belongs to the Section Image and Video Processing)
►▼
Show Figures

Figure 1
Open AccessArticle
Spectral-DETR: Learnable Frequency Decomposition with Adaptive Contrastive Regularization for Robust Underground Mine Detection
by
Yuexin Song, Lukang Dai, Xinqi Xu and Jun Yang
J. Imaging 2026, 12(9), 401; https://doi.org/10.3390/jimaging12090401 - 26 Aug 2026
Abstract
Underground mine object detection is challenged by low illumination, blur, dust scattering, and repetitive tunnel clutter, which jointly corrupt backbone features, entangle DETR queries, and weaken localization for small objects. Existing enhancement-based and detector-internal methods do not explicitly propagate degradation reliability across features,
[...] Read more.
Underground mine object detection is challenged by low illumination, blur, dust scattering, and repetitive tunnel clutter, which jointly corrupt backbone features, entangle DETR queries, and weaken localization for small objects. Existing enhancement-based and detector-internal methods do not explicitly propagate degradation reliability across features, decoder queries, and box refinement. We propose Spectral-DETR, a detector-internal reliability framework built on RF-DETR. Its central design is a cross-stage reliability pathway that connects Degradation-Aware Frequency Decomposition (DAFD), Degradation-Adaptive Query Contrastive Denoising (DQCD), and Salience-Calibrated Uncertainty with Learned Uncertainty Estimation (SCU+LUE). On Mine-Objects (14 classes, 3081 images), Spectral-DETR achieves an average precision of 0.917 at an intersection-over-union threshold of 0.5 and 0.493 when averaged over thresholds from 0.5 to 0.95, exceeding YOLOv9m by 1.6 and 0.8 percentage points, respectively, under the dataset-specific evaluation protocol. In controlled RF-DETR validation, the three reliability stages improve these two measures from 0.883 to 0.913 and from 0.472 to 0.486, respectively. Spectral-DETR obtains corresponding values of 0.848 and 0.571 on ExDark and 0.973 and 0.495 on ScienceDB. DQCD and SCU remain training-only losses with no inference cost.
Full article
(This article belongs to the Section Computer Vision and Pattern Recognition)
►▼
Show Figures

Figure 1
Open AccessArticle
PaIR: Partition-Based Information Rebalancing for Robust Text-Based Person Search
by
Luda Wang, Jiabao Li, Xinpan Yuan and Ningdan Zhang
J. Imaging 2026, 12(9), 400; https://doi.org/10.3390/jimaging12090400 - 25 Aug 2026
Abstract
Text-based person search (TPS) suffers from cross-modal informational skewness: pedestrian images are high-dimensional and redundancy-prone, while textual descriptions are sparse, incomplete, and sometimes inaccurate. To address the low alignment accuracy and poor robustness caused by the inherent uneven information distribution of visual and
[...] Read more.
Text-based person search (TPS) suffers from cross-modal informational skewness: pedestrian images are high-dimensional and redundancy-prone, while textual descriptions are sparse, incomplete, and sometimes inaccurate. To address the low alignment accuracy and poor robustness caused by the inherent uneven information distribution of visual and textual modalities in TPS, this paper proposes a unified Partition-based Information Rebalancing (PaIR) framework to realize balanced optimization and precise alignment of cross-modal information from both global content and local part dimensions. The framework adopts the CLIP dual-modal encoder for basic feature extraction and constructs a parallel global–local dual representation system to compensate for the lack of fine-grained spatial information in single global features. To eliminate modal redundancy and noise interference, a dual-modal noise suppression module is designed to filter invalid redundant information through visual foreground–background separation and textual token weight screening, while introducing adversarial constraints and orthogonal constraints to purify effective features. On this basis, a part balance alignment module is built to complete human semantic part decomposition and soft matching alignment for dual-modal features. Aiming at the common part semantic missing problem in textual descriptions, a visual part correlation affinity matrix is utilized for semantic associative completion to balance the information density of dual modalities. Finally, a global–local joint alignment strategy integrates hierarchical features and bidirectional cross-modal attention interaction to eliminate global–local semantic discontinuity and enhance fine-grained cross-modal matching capability. Extensive experiments on three public benchmarks demonstrate that PaIR consistently improves multiple baselines.
Full article
(This article belongs to the Topic Intelligent Image Processing Technology)
►▼
Show Figures

Figure 1
Open AccessArticle
Event-Guided Image Reconstruction for Nighttime Dynamic Scenes
by
Qingjiao Meng, Ji Li and Yan Jin
J. Imaging 2026, 12(9), 399; https://doi.org/10.3390/jimaging12090399 - 23 Aug 2026
Abstract
Image reconstruction in nighttime dynamic scenes is challenged by low illumination, long exposure, rapid camera or object motion, and sensor noise. Conventional RGB cameras, therefore, struggle to recover both sufficient brightness and clear structural details in nighttime dynamic scenes. To address this problem,
[...] Read more.
Image reconstruction in nighttime dynamic scenes is challenged by low illumination, long exposure, rapid camera or object motion, and sensor noise. Conventional RGB cameras, therefore, struggle to recover both sufficient brightness and clear structural details in nighttime dynamic scenes. To address this problem, we propose an event-guided image reconstruction method for nighttime dynamic visual perception. The method constructs a multi-channel event voxel representation by jointly encoding event count, event intensity, timestamp distribution, and blurred-frame intensity priors. A parameter-efficient local–global reconstruction network is then designed to restore fine-grained textures and model holistic structures. In addition, edge-alignment and blur-alignment constraints are introduced to improve geometric consistency and imaging plausibility. Experiments on the HQF and REDS datasets show that the proposed method outperforms existing methods in the MSE, PSNR, and SSIM. Compared with DeblurSR, it reduces the MSE by 14.81% on HQF and 10.00% on REDS, while improving the PSNR by 1.603 dB and 1.053 dB, respectively. Qualitative results further show sharper edges, lower structural errors, and better edge consistency. Low illumination, dynamic blur, rapid brightness variation, and event noise are also common degradation factors in nighttime UAV imaging, making the investigated problem technically relevant to that setting. However, because neither REDS nor HQF was acquired during an actual UAV flight, the reported results establish benchmark-level reconstruction performance rather than UAV-specific operational effectiveness.
Full article
(This article belongs to the Topic New Challenges in Image Processing and Pattern Recognition)
►▼
Show Figures

Figure 1
Open AccessArticle
Attention-Enhanced Multi-Scale Feature-Wise Linear Modulation for Fine-Grained Poisonous Mushroom Image Recognition
by
Yuan He, Haikun Lv, Chenyang Lu, Dengqi Yang, Xiaowei Li and Lina Zhang
J. Imaging 2026, 12(8), 398; https://doi.org/10.3390/jimaging12080398 - 21 Aug 2026
Abstract
Fine-grained poisonous mushroom recognition in natural scenes is challenging because of complex backgrounds, subtle morphological differences, and the limited interpretability of model decisions. To address these challenges, this paper proposes Att-FiLM, an attention-enhanced multi-scale Feature-Wise Linear Modulation network for poisonous mushroom image recognition.
[...] Read more.
Fine-grained poisonous mushroom recognition in natural scenes is challenging because of complex backgrounds, subtle morphological differences, and the limited interpretability of model decisions. To address these challenges, this paper proposes Att-FiLM, an attention-enhanced multi-scale Feature-Wise Linear Modulation network for poisonous mushroom image recognition. The model adopts an asymmetric dual-backbone architecture in which a frozen ConvNeXt-Base branch provides global semantic priors, while a trainable EfficientNet-B0 branch learns local discriminative features. Rather than directly concatenating heterogeneous features, Att-FiLM generates scale and shift parameters from semantic features and performs channel-wise modulation on multi-scale EfficientNet features at Stage 2 and Stage 4. This mechanism enables global semantic information to guide local feature learning while reducing feature redundancy and semantic inconsistency. Experimental results show that Att-FiLM achieves an Accuracy of 95.58% and an F1-score of 0.9455 on the poisonous/edible binary classification task. On the 190-class species-level classification task, it achieves a Top-1 Accuracy of 93.63% and a Macro-F1 of 0.9347. Interpretability analysis further shows that decision-relevant responses are frequently associated with morphologically relevant regions, including gills, annuli, volvae, and cap textures. These results indicate that Att-FiLM provides effective recognition performance together with interpretable decision evidence for mushroom recognition in complex natural scenes.
Full article
(This article belongs to the Section Image and Video Processing)
►▼
Show Figures

Figure 1
Highly Accessed Articles
Latest Books
E-Mail Alert
News
27 July 2026
Meet Us at the 29th International Conference on Medical Image Computing and Computer-Assisted Intervention, 27 September–1 October 2026, Strasbourg, France
Meet Us at the 29th International Conference on Medical Image Computing and Computer-Assisted Intervention, 27 September–1 October 2026, Strasbourg, France
1 September 2026
MDPI INSIGHTS: The CEO’s Letter #38 – 2 Million Published Articles, Outstanding Reviewers, Michele Parrinello Award, AIS 2026 & WSF-12
MDPI INSIGHTS: The CEO’s Letter #38 – 2 Million Published Articles, Outstanding Reviewers, Michele Parrinello Award, AIS 2026 & WSF-12
Topics
Topic in
BioMed, Cancers, Diagnostics, JCM, J. Imaging
Machine Learning and Deep Learning in Medical Imaging
Topic Editors: Rafał Obuchowicz, Michał Strzelecki, Adam Piórkowski, Karolina NurzynskaDeadline: 31 October 2026
Topic in
Applied Sciences, Electronics, J. Imaging, JMSE, Machines, Robotics, Sensors, Drones
Applications and Development of Underwater Robotics and Underwater Vision Technology, 2nd Edition
Topic Editors: Jingchun Zhou, Wenqi Ren, Qiuping Jiang, Yan-Tsung PengDeadline: 30 November 2026
Topic in
AI, Applied Sciences, Sensors, J. Imaging
Applied Computing and Machine Intelligence (ACMI): 2nd Edition
Topic Editors: Chuan-Ming Liu, Wei-Shinn KuDeadline: 31 December 2026
Topic in
Electronics, MAKE, Sensors, Applied Sciences, J. Imaging
Applied Computer Vision and Pattern Recognition: 3rd Edition
Topic Editors: Antonio Fernández-Caballero, Byung-Gyu KimDeadline: 28 February 2027
Conferences
Special Issues
Special Issue in
J. Imaging
Progress, Challenges, and Future Trends in Computer Vision and Pattern Recognition
Guest Editors: Qing Cai, Jinxing LiDeadline: 30 September 2026
Special Issue in
J. Imaging
Techniques in Multi-View Image Analysis
Guest Editor: Jian WeiDeadline: 30 September 2026
Special Issue in
J. Imaging
AI-Driven Medical Image Processing and Analysis
Guest Editor: Hong WangDeadline: 30 September 2026
Special Issue in
J. Imaging
Learning and Optimization for Medical Imaging–2nd Edition
Guest Editors: Simona Moldovanu, Elena Morotti, Adina CocuDeadline: 30 September 2026



