Next Issue
Volume 12, August
Previous Issue
Volume 12, June
 
 

J. Imaging, Volume 12, Issue 7 (July 2026) – 60 articles

Cover Story (view full-size image): Accurate correction of lens distortion is essential for applications such as 3D reconstruction, image registration, and camera calibration, particularly for specialised or disposable cameras, where only limited calibration data are available. We present a Deep Gaussian Process framework that models complex, spatially varying lens distortions, while also estimating the uncertainty of every corrected pixel. Unlike conventional methods that assume a single global distortion model, our approach adapts to local image behaviour and highlights regions where corrections are less reliable. Experiments on three challenging real-world camera systems demonstrate improved performance for highly irregular distortions, making the method well suited for next-generation scientific, medical, and industrial imaging devices. View this paper
  • Issues are regarded as officially published after their release is announced to the table of contents alert mailing list.
  • You may sign up for e-mail alerts to receive table of contents of newly released issues.
  • PDF is the official format for papers published in both, html and pdf forms. To view the papers in pdf format, click on the "PDF Full-text" link, and use the free Adobe Reader to open them.
Order results
Result details
Section
Select all
Export citation of selected articles as:
45 pages, 1738 KB  
Systematic Review
Structuring Variability in Human Gait Datasets: A Covariate-Centered Taxonomy and Systematic Review of Image- and Depth-Based Collections
by João Ferreira Nunes, Pedro Miguel Moreira and João Manuel R. S. Tavares
J. Imaging 2026, 12(7), 334; https://doi.org/10.3390/jimaging12070334 - 22 Jul 2026
Viewed by 188
Abstract
Human gait datasets play a central role in the development and evaluation of computer vision models. However, the current dataset landscape remains highly heterogeneous, with inconsistent reporting of acquisition conditions, user variability, and sensing configurations, which limits reproducibility and hinders principled cross-dataset comparability. [...] Read more.
Human gait datasets play a central role in the development and evaluation of computer vision models. However, the current dataset landscape remains highly heterogeneous, with inconsistent reporting of acquisition conditions, user variability, and sensing configurations, which limits reproducibility and hinders principled cross-dataset comparability. In this work, we propose a covariate-centered, modality-agnostic taxonomy for gait datasets, explicitly structuring variability across scene-level, user-level, and sensor-level factors. The proposed framework enables consistent characterization of datasets through a standardized set of covariates (A–R), bridging differences across application domains and sensing modalities. Following a systematic review protocol aligned with PRISMA 2020, we analyze 47 publicly available image- and depth-based human gait datasets spanning healthcare, biometric, and attribute-recognition application domains. Using the proposed taxonomy, we derive a quantitative analysis of covariate coverage, revealing systematic biases in current dataset design. Full article
(This article belongs to the Section Computer Vision and Pattern Recognition)
Show Figures

Figure 1

14 pages, 2132 KB  
Article
Decadal Changes in Institutional Diagnostic Reference Levels for X-Ray Angiography: A Retrospective Comparative Study
by Ioannis Antonakos, Emmanouil Anousis, Tatiana Roko, Antonia Alexiadou, Maria Dimitropoulou, Dimitris Filippiadis, Stavros Spiliopoulos, Konstantinos Palialexis, Athanasios Giannakis, Niki Parmenidou and Efstathios Efstathopoulos
J. Imaging 2026, 12(7), 333; https://doi.org/10.3390/jimaging12070333 - 22 Jul 2026
Viewed by 258
Abstract
Angiography is a key imaging modality for the diagnosis and treatment of vascular diseases, and the growing sophistication of interventional procedures has heightened the need for radiation dose optimization. Diagnostic Reference Levels (DRLs) are widely used to monitor patient exposure and to support [...] Read more.
Angiography is a key imaging modality for the diagnosis and treatment of vascular diseases, and the growing sophistication of interventional procedures has heightened the need for radiation dose optimization. Diagnostic Reference Levels (DRLs) are widely used to monitor patient exposure and to support optimization in accordance with the ALARA principle. This study compared radiation dose metrics from a newly installed angiographic system at Attikon University Hospital with those obtained from the institution’s previous system and with values reported in the published literature. Radiation dose and procedural parameters were retrospectively collected for digital cerebral subtraction angiography (DSA), embolization, nephrostomy, vertebroplasty, transjugular intrahepatic portosystemic shunt (TIPS), chemoembolization, and injection procedures. Dose area product (DAP), fluoroscopy-related DAP, patient entrance dose indicators, and fluoroscopy time were analyzed. Median DAP values ranged from 5.72 to 349.60 Gy·cm2 depending on the procedure. Compared with data acquired approximately a decade earlier, DAP values increased by an average of 96.4%, whereas fluoroscopy times remained largely unchanged. Despite this increase, dose levels were generally lower than those reported in the international literature. Median DAP values ranged from 5.72 to 349.60 Gy·cm2 depending on the procedure. Compared with data acquired approximately a decade earlier, DAP values increased by an average of 96.4%, whereas fluoroscopy times remained largely unchanged. Although these differences may reflect the combined influence of technological developments, evolving procedural complexity, operator-related factors, and changes in clinical practice over time, these variables were not directly assessed in the present retrospective study. Nevertheless, the updated institutional Diagnostic Reference Levels provide a valuable benchmark for radiation dose optimization, quality assurance, and future multicenter studies aimed at supporting national DRL establishment. Full article
(This article belongs to the Special Issue Diagnostic Imaging: From Basic Knowledge to Latest Advancements)
Show Figures

Figure 1

15 pages, 16367 KB  
Article
LABFNet: A Restoration Network Guided by the LAB Colour Space and Frequency-Domain Constraints
by Yaqian Zhang, Guanjun Wang, Quan Zhang and Bochao Zhou
J. Imaging 2026, 12(7), 332; https://doi.org/10.3390/jimaging12070332 - 22 Jul 2026
Viewed by 259
Abstract
In the restoration of mural images with rich colour information and complex texture structures, existing techniques typically extract the spatial-domain features in the Red–Green–Blue (RGB) colour space. However, the three RGB channels are physically decoupled without unified perceptual colour correlation constraints, which often [...] Read more.
In the restoration of mural images with rich colour information and complex texture structures, existing techniques typically extract the spatial-domain features in the Red–Green–Blue (RGB) colour space. However, the three RGB channels are physically decoupled without unified perceptual colour correlation constraints, which often leads to noticeable colour deviation in damaged regions with large colour variations. In addition, restoring both high-frequency texture details and low-frequency global structures in a mixed-frequency spatial domain can create conflicts between frequencies, making it difficult to generate realistic high-frequency details. To address these issues, we propose the laboratory frequency network (LABFNet), a restoration network guided by the laboratory (LAB) colour space and frequency-domain constraints. Our model has two key improvements: (1) it incorporates colour parameters from the LAB space to model colour loss in murals, and (2) it decomposes the image into low- and high-frequency components and enforces frequency consistency during restoration. In the Dunhuang 20–40% mask-ratio setting, the Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index (SSIM) improved by 1.58% and 0.27%, respectively, while the Mean Absolute Error (MAE), Learned Perceptual Image Patch Similarity (LPIPS) and CIEDE2000 decreased by 5.87%, 6.5%, and 21.65%, respectively. Experimental results on benchmark datasets show that LABFNet reduces colour deviation and structural defects. Full article
(This article belongs to the Topic Computer Vision and Image Processing, 3rd Edition)
Show Figures

Figure 1

19 pages, 13479 KB  
Article
Longitudinal CT Scanning for Explainable Early Detection of Postharvest Disorders: The ‘Braeburn’ Browning Case
by Dirk Elias Schut, Rachael Maree Wood, Rob Schouten, Robert van Liere, Tristan van Leeuwen and Kees Joost Batenburg
J. Imaging 2026, 12(7), 331; https://doi.org/10.3390/jimaging12070331 - 21 Jul 2026
Viewed by 324
Abstract
This study presents two workflows for leveraging longitudinal computed tomography (CT) datasets when developing deep learning-based detection systems for gradually developing postharvest disorders. Workflow 1 (Longitudinal Benchmarking) benchmarks neural networks by training and testing them on images from different stages of disorder progression. [...] Read more.
This study presents two workflows for leveraging longitudinal computed tomography (CT) datasets when developing deep learning-based detection systems for gradually developing postharvest disorders. Workflow 1 (Longitudinal Benchmarking) benchmarks neural networks by training and testing them on images from different stages of disorder progression. It examines the trade-off between detecting a disorder early or accurately and evaluates whether neural networks can generalize across time points. Workflow 2 (Longitudinal eXplainable Artificial Intelligence (XAI) Heatmaps) provides heatmaps that indicate how changes over time affect the outcomes of neural networks. It uses image registration to align an earlier-acquired image and then uses it as a baseline when calculating the heatmap. The workflows are demonstrated on a dataset of ‘Braeburn’ apples that were CT-scanned multiple times while developing internal browning during controlled-atmosphere (CA) storage and shelf life. The Longitudinal Benchmarking workflow was used to investigate whether images acquired immediately after CA storage can be used to predict the eventual browning after a shelf-life period, which is highly relevant in industrial practice. Moreover, the longitudinal XAI heatmaps avoided artifacts caused by out-of-distribution baselines or identical baseline regions, which occurred with conventional black or zero baselines. Full article
(This article belongs to the Section AI in Imaging)
Show Figures

Figure 1

16 pages, 5055 KB  
Article
Application of Machine Learning for Mean Glandular Dose Prediction Utilizing DICOM Mammography Images
by Ali A. A. Alghamdi
J. Imaging 2026, 12(7), 330; https://doi.org/10.3390/jimaging12070330 - 21 Jul 2026
Viewed by 319
Abstract
The growing demand for raw and processed scientific data has encouraged many researchers and research institutions to adopt an open-source data policy. At present, data accessibility is of paramount importance due to the growing demand for artificial intelligence (AI) and machine learning (ML) [...] Read more.
The growing demand for raw and processed scientific data has encouraged many researchers and research institutions to adopt an open-source data policy. At present, data accessibility is of paramount importance due to the growing demand for artificial intelligence (AI) and machine learning (ML) applications in various scientific fields, particularly medicine. Medium- to large-scale mammography datasets are widely used in breast cancer research to develop and evaluate computer-aided detection methods. However, there are only a few studies on using mammogram datasets for the prediction of the breast mean glandular dose (MGD) with AI or ML models. The aim of this study was to investigate the feasibility of using ML and deep ML for MGD prediction based on DICOM images and retrieved dosimetric data from DICOM mammogram images. A total of 26,988 mammography images in DICOM format were obtained from the Federated Research Data Repository (FRDR). Eleven regression algorithms and three neural network-based models were evaluated using five-fold cross-validation. In addition, a deep ML fusion model based on Vision Transformer (ViT) and tabular data was developed for the prediction of the MGD normalized conversion factor CF(DgN). A mean breast thickness of 61.37 mm and a mean MGD of 1.53 mGy (0.55–6.33 mGy) were calculated using this dataset. Regarding tabular data, the artificial neural network (ANN) sequential models outperformed other linear and tree-based models. The ViT deep ML fusion model was tested with three configuration versions differing on the number of features included. A comparison of the three versions revealed that the version with six features achieved the best overall predictor performance. This study demonstrates that ML and deep ML can effectively predict the MGD using dosimetric tabular data and mammography DICOM images. The use of ML with tabular data extracted from DICOM images can be further strengthened by incorporating larger and more diverse datasets. Full article
(This article belongs to the Section Medical Imaging)
Show Figures

Figure 1

20 pages, 2542 KB  
Article
Dynamic Convolution Enhanced Attention Network for Pulmonary Nodule Detection
by Shengqun Zhang, Annie Anak Joseph and Kho Lee Chin
J. Imaging 2026, 12(7), 329; https://doi.org/10.3390/jimaging12070329 - 21 Jul 2026
Viewed by 313
Abstract
Pulmonary nodules are circular or irregular lesions visible on chest computed tomography (CT), and their early detection is critical for lung cancer screening. Deep learning detection algorithms have been widely adopted for pulmonary nodule diagnosis; existing lightweight models suffer from redundant network parameters [...] Read more.
Pulmonary nodules are circular or irregular lesions visible on chest computed tomography (CT), and their early detection is critical for lung cancer screening. Deep learning detection algorithms have been widely adopted for pulmonary nodule diagnosis; existing lightweight models suffer from redundant network parameters and low detection accuracy for tiny lesions. To address these limitations, this study proposes an improved detection model based on YOLOv8n. First, Omni-Dimensional Dynamic Convolution (ODConv) replaces static convolution in the backbone to enhance multi-morphology nodule feature extraction. Second, the Convolutional Block Attention Module (CBAM) is embedded at multiple positions of the neck network to suppress background interference from blood vessels and normal lung parenchyma. Third, Complete Intersection over Union (CIoU) loss is substituted by Wise Intersection over Union (W-IoU) to optimize bounding box regression for hard samples with blurred boundaries. Experiments on the LUNA16 dataset show that compared with the original YOLOv8n, the proposed model improves Precision by 6.3%, Recall by 8.6%, mAP50 by 3.4%, and mAP50-95% by 2.7% while maintaining high inference speed. Additional generalization verification on the LIDC-IDRI multi-center dataset further proves the robustness of the proposed lightweight architecture, which achieves balanced accuracy and real-time performance compared with mainstream detection models. Full article
(This article belongs to the Section Medical Imaging)
Show Figures

Figure 1

20 pages, 4332 KB  
Article
Contrastive and Transfer Learning for Aligned Multimodal Neuroimaging Classification of Autism Spectrum Disorder
by Raja Vavekanand, Ganesh Kumar, Muhammad Moazzam Jawaid, Shafiya Qadeer Memon and Teerath Kumar
J. Imaging 2026, 12(7), 328; https://doi.org/10.3390/jimaging12070328 - 20 Jul 2026
Viewed by 347
Abstract
Autism Spectrum Disorder (ASD) assessment remains challenging because behavioural instruments are partly observer-dependent and neuroimaging data are heterogeneous. This paper presents FAA (Fuse After Aligned), which is a multimodal classification framework that combines transfer learning for structural MRI (sMRI) representation learning with a [...] Read more.
Autism Spectrum Disorder (ASD) assessment remains challenging because behavioural instruments are partly observer-dependent and neuroimaging data are heterogeneous. This paper presents FAA (Fuse After Aligned), which is a multimodal classification framework that combines transfer learning for structural MRI (sMRI) representation learning with a contrastive objective for the pre-fusion alignment of sMRI and resting-state functional MRI-derived functional connectivity (FC) features. Evaluation was restricted to the single-site ABIDE-I New York University subset comprising 75 participants with ASD and 98 typically developing controls. Under the reported five-fold internal cross-validation protocol, FAA achieved a mean accuracy of 92.6% compared with 87.4% for naive fusion and 90.9% for the sMRI-only baseline. Ablation analyses indicate that adding the contrastive objective is associated with improved classification performance and that ResNet-18 outperforms the evaluated ViT-16 configurations in this small-sample setting. These findings support the methodological value of pre-fusion feature alignment within the evaluated cohort. The framework offers a robust, computationally efficient, and clinically viable approach for objective ASD diagnosis with strong potential for generalisation to multi-site neuroimaging applications. Full article
(This article belongs to the Section Medical Imaging)
Show Figures

Figure 1

13 pages, 1136 KB  
Article
A Simplified CT Score for Thrombus Burden in Acute Pulmonary Embolism: Clinical Correlation and Reproducibility
by Ignacio Díaz-Lorenzo, Rio Jorge Aguilar Torres, Paloma Caballero Sanchez-Robles, Raquel Caminero Garcia, Alfonso Canabal Berlanga, Alfonsa Friera Reyes and Alberto Alonso-Burgos
J. Imaging 2026, 12(7), 327; https://doi.org/10.3390/jimaging12070327 - 19 Jul 2026
Viewed by 421
Abstract
(1) Objectives: In acute pulmonary embolism (PE), detailed thrombus burden scores are often complex and time-consuming, limiting their integration into urgent radiology reports. We evaluated a simplified modified Ghanima score (GmScore and GmS) designed to provide a structured estimate of thrombus burden and [...] Read more.
(1) Objectives: In acute pulmonary embolism (PE), detailed thrombus burden scores are often complex and time-consuming, limiting their integration into urgent radiology reports. We evaluated a simplified modified Ghanima score (GmScore and GmS) designed to provide a structured estimate of thrombus burden and assessed its clinical correlation and reproducibility. (2) Methods: In this retrospective single-center study, 132 consecutive patients with confirmed acute PE were classified according to the modified GmScore: GmS1 (segmental), GmS2 (lobar), and GmS3 (main pulmonary arteries), considering luminal obstruction ≥ 50%. European Society of Cardiology (ESC) risk category, simplified Pulmonary Embolism Severity Index (sPESI), CT right-to-left ventricular (RV/LV) ratio, echocardiographic right ventricular dysfunction, and 30-day mortality were recorded. Inter- and intraobserver agreement were assessed using weighted kappa. (3) Results: In 132 patients (mean age 64.8 ± 16.5 years; 77 men), a significant clinical gradient was observed across GmScore categories. ESC intermediate–high/high risk occurred in 0% of GmS1 and 95.6% of GmS2–3 patients (p < 0.001). The median RV/LV ratio increased progressively (0.76, 1.58, and 1.79 for GmS1–3; p < 0.001), with a strong correlation between the GmScore and RV/LV (Spearman ρ = 0.75). GmS2 and GmS3 showed no significant difference in ventricular repercussion (p = 0.938), whereas GmS1 differed markedly. Using GmS ≥ 2 to identify ESC intermediate–high/high risk yielded 100% sensitivity and negative predictive value. Interobserver agreement was excellent (κ = 0.92). Thirty-day mortality was 0% in GmS1, 2.0% in GmS2, and 14.6% in GmS3 (p = 0.005). (4) Conclusions: The modified GmScore is a simple, reproducible CT-based descriptor that aligns closely with right ventricular repercussion and ESC risk stratification. Full article
(This article belongs to the Section Medical Imaging)
Show Figures

Figure 1

37 pages, 2713 KB  
Article
Microscopic Pollen Image Classification via Contour-Signal Representation, Wavelet Analysis, and CNN
by Abror Shavkatovich Buriboev, Akhram Nishanov, Shuxrat Isroilov, Inomjon Narzullaev, Umidjon Djumayozov, Shavkat Buriboyev, Temur Azamov, Parda Yuldashov, Davron Shodmonov, Djamshid Sultanov and Abbos Abduvaytov
J. Imaging 2026, 12(7), 326; https://doi.org/10.3390/jimaging12070326 - 18 Jul 2026
Viewed by 261
Abstract
Accurate classification of pollen grains in microscopic images remains challenging because of noise, structural variability, background complexity, weak texture, and intra-class similarity. To address these issues, this study proposes a hybrid framework that integrates contour-signal modeling, spectral–wavelet analysis, and deep learning for robust [...] Read more.
Accurate classification of pollen grains in microscopic images remains challenging because of noise, structural variability, background complexity, weak texture, and intra-class similarity. To address these issues, this study proposes a hybrid framework that integrates contour-signal modeling, spectral–wavelet analysis, and deep learning for robust microscopic pollen image recognition. In the proposed approach, microscopic pollen images are first converted into contour-based point-signal representations, allowing object boundaries to be analyzed as structured one-dimensional signals. To improve signal quality under real imaging conditions, the framework incorporates Gaussian, median, and contour-aware filtering together with defect-point detection and correction. The processed contour signals are then analyzed using Fourier transform, continuous wavelet transform, and discrete wavelet transform to extract complementary global and local descriptors. These enriched representations are provided to a convolutional neural network for final classification. Experiments conducted on a seven-class microscopic pollen-image dataset demonstrate that the proposed method outperforms conventional computer-vision and baseline deep-learning approaches. The best-performing hybrid configuration achieved an error rate of 6.4%, while the overall classification accuracy reached 0.977 with an F1-score of 0.966, compared with 0.837 for a traditional computer-vision pipeline. These results confirm that combining contour-based signal processing with hierarchical deep feature learning provides an effective and noise-robust strategy for microscopic pollen image recognition. However, the present validation is limited to pollen images, and further experiments on broader microscopic object datasets are required to assess generalization to other micro-object categories such as nanoparticles, fibers, rods, and synthetic microstructures. Full article
(This article belongs to the Section Computer Vision and Pattern Recognition)
Show Figures

Figure 1

29 pages, 3314 KB  
Article
Efficient Object Detection in Compressed Domain by Exploiting Knowledge Distillation from Pixel Domain
by Serhat Dikyar and Behcet Ugur Toreyin
J. Imaging 2026, 12(7), 325; https://doi.org/10.3390/jimaging12070325 - 18 Jul 2026
Viewed by 375
Abstract
The proliferation of high-definition video data necessitates highly efficient processing pipelines for real-time edge analytics. However, traditional object detection architectures rely exclusively on pixel-domain inputs, which renders the computationally prohibitive decoding phase a latency bottleneck. In this paper, we propose a novel dual-phase [...] Read more.
The proliferation of high-definition video data necessitates highly efficient processing pipelines for real-time edge analytics. However, traditional object detection architectures rely exclusively on pixel-domain inputs, which renders the computationally prohibitive decoding phase a latency bottleneck. In this paper, we propose a novel dual-phase framework designed to achieve fast and efficient object detection directly within the partially decoded compressed-domain data. First, we introduce a partial decoding paradigm featuring the Low-Frequency Spectral Prioritization method on the encoder side. By systematically discarding high-frequency residual coefficients and retaining only a sparse subset of fundamental spatial frequencies, this method dramatically reduces transmission payloads and accelerates the standard decoding process. Second, to recover the structural fidelity lost due to the intentional omission of residual data, we employ a multi-granularity cross-domain knowledge distillation architecture. This strategy aligns global contextual features, foreground boundary attention maps, and final response logits, transferring rich representational capacities from a high-performing pixel-domain teacher network to a lightweight compressed-domain student network. Comprehensive experiments utilizing RetinaNet, FCOS, and GFL object detection networks on the COCO-mini dataset demonstrate the superiority of the proposed framework. By retaining fundamental residual coefficients within the HEVC pipeline, the proposed method reduces average decoding latency while improving the mAP score by +0.96% over the conventional fully decoded pixel-domain baseline on the COCO-mini dataset. Full article
(This article belongs to the Section Computer Vision and Pattern Recognition)
Show Figures

Figure 1

38 pages, 3059 KB  
Review
Review: Techniques in Egocentric Multi-View Image Analysis: Advances, Challenges, and Future Directions
by Duc Tri Phan and Hong Duc Nguyen
J. Imaging 2026, 12(7), 324; https://doi.org/10.3390/jimaging12070324 - 17 Jul 2026
Viewed by 351
Abstract
Egocentric multi-view image analysis refers to the processing of utilizing synchronized video streams captured from multiple wearable cameras worn on the head or body, providing complementary first-person perspectives of dynamic, real-world interactions. Unlike single-view egocentric vision, which may suffer from severe occlusions, motion [...] Read more.
Egocentric multi-view image analysis refers to the processing of utilizing synchronized video streams captured from multiple wearable cameras worn on the head or body, providing complementary first-person perspectives of dynamic, real-world interactions. Unlike single-view egocentric vision, which may suffer from severe occlusions, motion blur, and limited field-of-view or traditional fixed-camera multi-view setups (assuming static geometry and controlled environments), egocentric multi-view systems leverage body-worn rigs to enable a more robust and flexible 3D understanding in open-world, mobile scenarios. In this work, we present a systematic survey of advancements in cross-view feature fusion, geometric consistency enforcement, open-world detection, human–object interaction (HOI) modeling, action segmentation, 3D reconstruction, and novel-view synthesis specifically tailored to wearable multi-camera platforms. Key datasets released between 2024 and 2026—including HOT3D (833 min of synchronized multi-view hand/object interactions from Project Aria and Quest 3), MultiEgo (first multi-egocentric dataset for 4D social scene reconstruction), and Ego-1K (large-scale 12-camera rig for dynamic 3D video synthesis) are thoroughly examined alongside an analysis of integrations with large language models (LLMs) and vision–language models that drive performance gains, typically in the 15–30% range over single-view baselines in hand tracking, HOI recognition, and reconstruction fidelity, although we show through a consolidated meta-analysis that this gain is task-dependent: larger for geometry-bottlenecked tasks such as in-hand object lifting, and smaller, method-dependent, or occasionally negative for semantic-recognition tasks such as keystep recognition under naive view fusion. These methods cover work in multi-view stereo, cross-view learning, and novel-view synthesis while addressing several real-time wearable constraints. Practical applications such as immersive Augmented Reality/Virtual Reality (AR/VR), assistive robotics, and healthcare monitoring are also discussed together with the challenges in motion calibration, benchmark diversity, and edge deployment ability. Thus, in this review, we attempt to fill a critical gap by focusing exclusively on wearable multi-view systems in an open-world setting, synthesizing the latest literature to chart future directions toward more embodied and continual learning agents. Full article
(This article belongs to the Special Issue Techniques in Multi-View Image Analysis)
Show Figures

Figure 1

26 pages, 14712 KB  
Article
Magnetic Resonance Imaging Preprocessing for Robust Spinal Cord Segmentation in Cervical Myelopathy
by Hediyeh Toufani, Richard M. Dansereau, Philippe Phan, Jefferson R. Wilson and Eve C. Tsai
J. Imaging 2026, 12(7), 323; https://doi.org/10.3390/jimaging12070323 - 17 Jul 2026
Viewed by 312
Abstract
Accurate spinal cord segmentation is important for quantitative analysis of spinal cord magnetic resonance imaging, including measurement of cross-sectional area and diffusion-based microstructural characterization. In pathological conditions like cervical myelopathy, the shape deformation induced by cord compression is extreme, rendering automated segmentation particularly [...] Read more.
Accurate spinal cord segmentation is important for quantitative analysis of spinal cord magnetic resonance imaging, including measurement of cross-sectional area and diffusion-based microstructural characterization. In pathological conditions like cervical myelopathy, the shape deformation induced by cord compression is extreme, rendering automated segmentation particularly challenging. While deep learning-based methods yield good results in healthy or mildly pathological cases, their reliability suffers when anatomical assumptions fail under compression. In this work, we introduce a pathology-aware, boundary-focused preprocessing framework that directly aims to mitigate failure modes imposed by cord compression. Instead of generic preprocessing, each component aims to enhance intensity homogeneity, suppress noise and improve boundary visibility. At the core of this approach is a multi-representation input derived from a single T2*-weighted scan, whereby complementary intensity-, contrast- and edge-enhanced representations are fed to the U-Net model. The proposed framework is evaluated on spinal cord MRI data from three clinical centers (194 cervical myelopathy cases). The results demonstrate that the proposed preprocessing framework improves segmentation accuracy, robustness, and stability, particularly in anatomically challenging regions affected by compression. These findings highlight the importance of pathology-aware preprocessing for reliable spinal cord segmentation in cervical myelopathy. Full article
(This article belongs to the Section Image and Video Processing)
Show Figures

Figure 1

19 pages, 2930 KB  
Article
Sex Estimation Based on the Cranial Base of Three-Dimensional Skull Models from the Bosnia and Herzegovina Population Using Geometric Morphometrics
by Zurifa Ajanović, Saleha Redžepi, Uzeir Ajanović, Naida Spahović, Amina Zorlak-Čavčić, Emina Dervišević, Admir Terzić and Mirza Pojskić
J. Imaging 2026, 12(7), 322; https://doi.org/10.3390/jimaging12070322 - 16 Jul 2026
Viewed by 423
Abstract
Sex estimation is a fundamental component of biological profiling in forensic anthropology, particularly when skeletal remains are incomplete or fragmented. This study aimed to evaluate sex estimation of the cranial base using geometric morphometrics and to assess the predictive value of cranial base [...] Read more.
Sex estimation is a fundamental component of biological profiling in forensic anthropology, particularly when skeletal remains are incomplete or fragmented. This study aimed to evaluate sex estimation of the cranial base using geometric morphometrics and to assess the predictive value of cranial base morphology for sex estimation. The study included 211 adult skulls (139 male, 72 female) from the Bosnian population. Each skull was digitized to generate 3D models, and 27 anatomical landmarks were recorded. Landmark coordinates were standardized using Generalized Procrustes Analysis, Principal Component Analysis, Discriminant Function Analysis with permutation testing, and regression of shape on centroid size. Statistically significant sex estimation was observed at both the form (shape and size) and shape levels. Classification accuracy based on cranial base form reached 92.81% for males and 86.11% for females. Shape-based classification, after removal of size effects, also showed high accuracy (90.65% for males and 81.94% for females). Regression analysis indicated that size contributed significantly but modestly to shape variation. The cranial base exhibits stable sexually dimorphic patterns and may represent a reliable anatomical region for sex estimation. These findings contribute to population-specific standards for the Bosnia and Herzegovina population and support the forensic applicability of 3D geometric morphometric approaches. Full article
(This article belongs to the Section Biometrics, Forensics, and Security)
Show Figures

Figure 1

17 pages, 11624 KB  
Article
An Adaptive Attention-Driven Quadruplet Deep Hashing Method for Retrieving Histopathological Images
by Seyed Mohammad Alizadeh, Henning Müller and Mohammad Sadegh Helfroush
J. Imaging 2026, 12(7), 321; https://doi.org/10.3390/jimaging12070321 - 15 Jul 2026
Viewed by 286
Abstract
Retrieving histopathological images can assist in the recognition and treatment planning of several diseases. Nevertheless, high-dimensional features can make this process complex and inefficient. These challenges can be addressed by encoding the feature domain into binary codes of different lengths utilizing deep hashing [...] Read more.
Retrieving histopathological images can assist in the recognition and treatment planning of several diseases. Nevertheless, high-dimensional features can make this process complex and inefficient. These challenges can be addressed by encoding the feature domain into binary codes of different lengths utilizing deep hashing approaches. Still, the vanishing gradient challenge remains a concern in these approaches. According to several studies, quadruplet deep hashing models have exhibited promising performance in retrieving images from multi-category datasets. Furthermore, adding an attention module to a convolutional neural network architecture can increase the efficiency of feature extraction. Thus, we introduce an adaptive quadruplet deep hashing model to retrieve histopathological images. Four designed deep hashing models with matching structures and parameters are utilized to produce hash codes. The resulting codes are trained according to a novel adaptive quadruplet loss function. The adaptive structure is capable of improving retrieval performance. The presented approach also suggests a novel hash layer for the vanishing gradient issue. In addition, a simple yet effective attention module is implemented to enhance feature extraction performance. Our model is evaluated on three publicly available histopathology datasets: Kather, Kimia Path960, and Kimia Path24C. The results indicate that the suggested approach achieves the highest mean average precision (MAP) of approximately 0.9940, 0.9983, and 0.9968 for the respective datasets. Based on experiments performed on the datasets, our model surpasses current hashing techniques. Full article
(This article belongs to the Section AI in Imaging)
Show Figures

Figure 1

19 pages, 7862 KB  
Article
Fast-CenLaneNet: A Lightweight Instance Segmentation-Based Network for Real-Time Lane Detection
by Qidong Han, Shuo Feng, Yang Gao, Mengyao Li, Teng Meng, Ke Li and Yuhao Yang
J. Imaging 2026, 12(7), 320; https://doi.org/10.3390/jimaging12070320 - 13 Jul 2026
Viewed by 410
Abstract
Lane detection is a critical component of autonomous driving systems, requiring both high accuracy and real-time performance under complex driving scenarios. Unlike current methods that rely on predefined lane counts, instance segmentation methods can handle an arbitrary number of lanes, making them more [...] Read more.
Lane detection is a critical component of autonomous driving systems, requiring both high accuracy and real-time performance under complex driving scenarios. Unlike current methods that rely on predefined lane counts, instance segmentation methods can handle an arbitrary number of lanes, making them more adaptable in real-world applications. However, this flexibility typically relies on dense pixel-level predictions, which necessitate large-scale networks and result in prohibitively high computational costs, hindering deployment on embedded platforms. To address these challenges, we present Fast-CenLaneNet, a lightweight architecture that improves inference efficiency while maintaining detection accuracy. Specifically, we design a lightweight backbone to reduce model parameters and computational cost, propose a learnable spatial similarity attention module to capture spatial dependencies within lane regions and enhance feature discriminability, and construct multi-branch output heads with Ghost convolutions to refine lane-related features with low computational overhead. Experiments on the TuSimple and CULane benchmarks demonstrate that Fast-CenLaneNet achieves a favorable accuracy–efficiency trade-off. On TuSimple, Fast-CenLaneNet obtains 96.40 ± 0.06% accuracy and 162.7 ± 6.8 FPS with 4.7 M parameters and 9.9 GFLOPs. Compared with CenLaneNet, it reduces the number of parameters by 89.1% and improves forward inference speed by 107.5%, with an accuracy decrease of only 0.08 percentage points. Full article
(This article belongs to the Special Issue Computer Vision and Image Processing: Advances and Challenges)
Show Figures

Figure 1

13 pages, 2988 KB  
Article
Exploring Adipose Tissue Behavior in CT: Impact of Age, Sex, and Contrast Media on Body Composition, Liver and Skeletal Muscle
by Emil Matthisson, Hanns-Christian Breit, Markus Obmann, Jakob Wasserthal, Martin Segeroth and Daniel Boll
J. Imaging 2026, 12(7), 319; https://doi.org/10.3390/jimaging12070319 - 13 Jul 2026
Viewed by 294
Abstract
Objectives: To evaluate the impact of contrast phase, age, and sex on CT-derived body composition metrics—specifically attenuation and volume of subcutaneous adipose tissue (SAT), visceral adipose tissue (VAT), liver, and skeletal muscle. The potential of the proportion of muscle voxels below 0 Hounsfield [...] Read more.
Objectives: To evaluate the impact of contrast phase, age, and sex on CT-derived body composition metrics—specifically attenuation and volume of subcutaneous adipose tissue (SAT), visceral adipose tissue (VAT), liver, and skeletal muscle. The potential of the proportion of muscle voxels below 0 Hounsfield units (HU) as a surrogate for fatty infiltration was also explored. Materials and Methods: A retrospective analysis of 866 multiphasic abdominal CT scans (non-enhanced [NE], arterial [ART], portal venous [PV]) from 2012 to 2022 was performed. Segmentation of SAT, VAT, liver, and skeletal muscle was conducted using the AI-based TotalSegmentator. Wilcoxon signed-rank tests and Bland–Altman analysis (mean bias and 95% limits of agreement) were applied to assess contrast-related effects; Spearman’s correlation coefficient was used to assess demographic associations. Results: Significant variation in attenuation and volume of SAT, VAT, and muscle was observed across contrast phases (p < 0.001). SAT attenuation was higher in NE and PV than in ART, while VAT attenuation was highest in PV. SAT volume increased and VAT volume decreased in contrast-enhanced phases. Attenuation and volume showed strong inter-phase correlation (ρ > 0.9). VAT attenuation was significantly higher in females, whereas VAT volume was significantly greater in males. VAT volume negatively correlated with liver attenuation (ρ = −0.33). Muscle voxels <0 HU were significantly reduced in contrast-enhanced scans. Conclusions: Contrast phase, age, and sex significantly influence CT-based body composition parameters. These confounding factors should be considered when using quantitative imaging biomarkers in clinical and research settings. Full article
(This article belongs to the Section Medical Imaging)
Show Figures

Figure 1

11 pages, 1107 KB  
Article
The Use of High-Frequency Skin Ultrasound in the Evaluation of Psoriatic Plaques—A Pilot Comparative Study Between Conventional and Biological Therapy
by Adelina Filofteia Ghilencea, Daniel Octavian Costache, Constantin Căruntu, Maria Moga and Raluca Simona Costache
J. Imaging 2026, 12(7), 318; https://doi.org/10.3390/jimaging12070318 - 13 Jul 2026
Viewed by 335
Abstract
Introduction. Psoriasis is a chronic inflammatory disease histologically characterized by epidermal hyperproliferation, altered keratinocyte differentiation and dermal vascular remodeling. Although the diagnosis is mainly clinical, non-invasive imaging methods, such as high-frequency skin ultrasound, allow an objective assessment of skin changes and disease [...] Read more.
Introduction. Psoriasis is a chronic inflammatory disease histologically characterized by epidermal hyperproliferation, altered keratinocyte differentiation and dermal vascular remodeling. Although the diagnosis is mainly clinical, non-invasive imaging methods, such as high-frequency skin ultrasound, allow an objective assessment of skin changes and disease activity. Material and Methods. We conducted a pilot, observational, cross-sectional and comparative study, conducted within the Dermatovenerology Department of the Central Military Emergency Hospital “Dr. Carol Davila”, Bucharest, which included 40 patients diagnosed with psoriasis vulgaris, of whom 22 received conventional systemic treatment (methotrexate 15 mg/week), and 18 received biological therapy. For each patient, a representative, clinically active and recently appeared psoriatic plaque was evaluated with ultrasound, and the thickness of the epidermis, the thickness of the dermis, the thickness of the hypoechoic subepidermal band (SLEB) and the Doppler signal were analyzed. Statistical analysis was performed using SPSS v26. Results. Patients under biologic therapy had significantly lower ultrasound parameters compared to those under conventional systemic therapy, especially regarding epidermis thickness, hypoechoic subepidermal band thickness and Doppler signal. The PASI score was significantly higher in the conventionally treated group. Also, significant positive correlations were found between the PASI score and the hypoechoic subepidermal band thickness and the Doppler signal, Conclusions. Ultrasound parameters represent useful objective markers in the evaluation of psoriasis, reflecting disease activity. Patients under biologic therapy presented, at the time of evaluation, imaging parameters suggestive of reduced skin inflammation compared to those treated conventionally. Full article
(This article belongs to the Section Medical Imaging)
Show Figures

Figure 1

33 pages, 6785 KB  
Review
Pedestrian Detection Techniques for Advanced Driver Assistance Systems: A Comprehensive Review
by Dănuţ-Ovidiu Pop and Adrian-Silviu Roman
J. Imaging 2026, 12(7), 317; https://doi.org/10.3390/jimaging12070317 - 10 Jul 2026
Viewed by 489
Abstract
Pedestrian detection is a fundamental component of Advanced Driver Assistance Systems (ADAS) and plays a key role in collision avoidance and the safety of vulnerable road users. This paper presents a structured review of pedestrian detection methodologies developed between 2000 and 2025, spanning [...] Read more.
Pedestrian detection is a fundamental component of Advanced Driver Assistance Systems (ADAS) and plays a key role in collision avoidance and the safety of vulnerable road users. This paper presents a structured review of pedestrian detection methodologies developed between 2000 and 2025, spanning classical vision techniques and modern deep learning architectures. We organize the review into two phases. First, we examine classical methods, including Histogram of Oriented Gradients (HOG)+Support Vector Machine (SVM), Viola–Jones, Deformable Part Models, and Integral Channel Features, which established the conceptual foundations of the field. Then, we analyze state-of-the-art deep learning architectures, categorized by detector stage (one-stage vs. two-stage), localization strategy (anchor-based vs. anchor-free), feature extraction paradigm (Convolutional Neural Network (CNN)-based vs. transformer-based), output representation (bounding box vs. instance segmentation), and computational profile (lightweight vs. heavyweight). Several design principles introduced by classical methods remain visible in modern architectures, indicating that they were not fully superseded. The review also examines publicly available benchmark datasets and compares the strengths and limitations of camera-, Light Detection And Ranging (LiDAR)-, radar-, and multi-sensor-fusion-based systems for ADAS deployment. We close by identifying six open problems for the field: adversarial robustness, real-time inference under embedded constraints, detection under adverse weather, dataset bias and demographic fairness, the deployment of Bird’s-Eye View (BEV) and unified perception on automotive hardware, and explainability for safety-critical use. Full article
Show Figures

Figure 1

22 pages, 3083 KB  
Article
NS-GUSL: Green U-Shaped Learning for Nuclei Segmentation from Histopathology Images
by Catherine Aurelia Christie Alexander, Vasileios Magoulianitis, Jiaxin Yang and C.-C. Jay Kuo
J. Imaging 2026, 12(7), 316; https://doi.org/10.3390/jimaging12070316 - 10 Jul 2026
Viewed by 334
Abstract
Nuclei segmentation is a key task in digital histopathology, highlighting important aspects of nuclear morphology and topology in many cancer-related evaluations and studies. Variability in nuclear appearance both within and across different organs, stain heterogeneity, and inconsistencies in acquisition procedures contribute to the [...] Read more.
Nuclei segmentation is a key task in digital histopathology, highlighting important aspects of nuclear morphology and topology in many cancer-related evaluations and studies. Variability in nuclear appearance both within and across different organs, stain heterogeneity, and inconsistencies in acquisition procedures contribute to the complexity of the task. The existing nuclei segmentation methods apply deep learning to address these challenges, using models with millions of parameters, thereby significantly increasing computational complexity. They also face limitations in generalizing to unseen organs and slide preparations. In this paper, we propose a transparent and lightweight Green U-Shaped Learning model for nuclei segmentation (NS-GUSL). NS-GUSL features a multi-scale architecture for coarse-to-fine refinement of probability maps, which are subsequently binarized using a novel low-confidence sample binarization (LCSB) technique. The model features a modular, feed-forward feature learning scheme with unsupervised representation learning and supervised feature selection and generation. A final morphological post-processing step refines the segmentation maps to improve instance separation while preserving nuclei convexity. The model was trained and tested on the MoNuSeg dataset and compared against other deep learning baselines for segmentation performance. In addition, external validation experiments were conducted to evaluate the proposed model’s generalizability to unseen organs and staining procedures. NS-GUSL exhibits the best panoptic segmentation performance and competitive detection quality across all datasets. Moreover, our model is shown to be compact, low in computational complexity, and to have a minimal carbon footprint, compared to other deep learning models, making it a suitable choice for deployment on edge devices. Full article
(This article belongs to the Special Issue AI-Driven Medical Image Processing and Analysis)
Show Figures

Figure 1

19 pages, 4550 KB  
Article
B-Mode Ultrasound Radiomics for Differentiating Benign and Malignant Small Hyperechoic Renal Masses: An Exploratory Single-Center Experience
by Fabrizio Urraro, Nicoletta Giordano, Vittorio Patanè, Marco Piscopo, Giovanni Ciani, Giovanni Balestrucci, Maria Chiara Brunese, Anna Russo, Mario Sansone and Alfonso Reginelli
J. Imaging 2026, 12(7), 315; https://doi.org/10.3390/jimaging12070315 - 10 Jul 2026
Viewed by 527
Abstract
Introduction: Small hyperechoic renal masses are frequently detected incidentally on conventional ultrasound and are often presumed to represent benign lesions, particularly angiomyolipomas. However, malignant renal tumors, including renal cell carcinoma, may also appear hyperechoic when small, creating a diagnostic challenge at first-line [...] Read more.
Introduction: Small hyperechoic renal masses are frequently detected incidentally on conventional ultrasound and are often presumed to represent benign lesions, particularly angiomyolipomas. However, malignant renal tumors, including renal cell carcinoma, may also appear hyperechoic when small, creating a diagnostic challenge at first-line imaging. This study aimed to evaluate the feasibility and exploratory diagnostic performance of B-mode ultrasound radiomics for differentiating benign and malignant small hyperechoic renal masses. Methods: This retrospective single-center study included adult patients with incidentally detected small hyperechoic renal masses measuring ≤3 cm and examined between July 2022 and April 2025. All lesions underwent standardized B-mode ultrasound assessment and multidisciplinary review. Final diagnosis was established by histopathology when available or by longitudinal ultrasound follow-up stability for lesions considered benign. Lesions were manually segmented on representative B-mode DICOM images, and original radiomic features were extracted using PyRadiomics version 3.0 according to standardized definitions compatible with the Image Biomarker Standardisation Initiative framework. A total of 114 original radiomic features were extracted from each lesion. The primary comparison was benign versus malignant lesions. Diagnostic performance was assessed using feature-level receiver operating characteristic analysis. Results: Forty-two lesions were included in the final radiomic cohort, including 26 malignant renal cell carcinomas and 16 benign angiomyolipomas. Malignant lesions included papillary renal cell carcinoma, chromophobe renal cell carcinoma, and clear-cell renal cell carcinoma. All malignant lesions were histologically confirmed. Among benign lesions, 14 angiomyolipomas were classified based on longitudinal ultrasound stability, whereas 2 were confirmed by ultrasound-guided percutaneous biopsy after mild dimensional increase during imaging surveillance. Among the extracted radiomic features, firstorder_Variance and firstorder_MeanAbsoluteDeviation showed the highest exploratory discriminatory performance, each achieving an area under the receiver operating characteristic curve of 0.837. Both features are first-order measures of gray-level dispersion within the segmented lesion. Higher values were observed in malignant lesions, suggesting greater intralesional grayscale heterogeneity compared with benign angiomyolipomas. Conclusions: B-mode ultrasound radiomics is feasible for the quantitative assessment of small hyperechoic renal masses and may provide complementary information for differentiating benign angiomyolipomas from malignant renal cell carcinomas. firstorder_Variance emerged as a representative candidate imaging biomarker of grayscale dispersion, with firstorder_MeanAbsoluteDeviation showing concordant performance as a related dispersion measure. These findings should be considered preliminary and hypothesis-generating and require validation in larger multicenter cohorts before clinical implementation. Full article
(This article belongs to the Section Medical Imaging)
Show Figures

Figure 1

11 pages, 6737 KB  
Article
Non-Destructive Neutron Tomography Analysis of Ceramic Vessels from the Shubarat-1 Archeological Site
by Kuanysh Nazarov, Veronica Smirnova, Murat Kenessarin, Yekaterina Dubyagina, Yeldos Kariyev, Sergey Kichanov, Ayazhan Zhomartova, Bagdaulet Mukhametuly and Elmira Myrzabekova
J. Imaging 2026, 12(7), 314; https://doi.org/10.3390/jimaging12070314 - 10 Jul 2026
Viewed by 275
Abstract
We studied the structural features in internal pores and mineral inclusions of several vessels from the Shubarat-1 archeological site in the Republic of Kazakhstan, dating to the late first millennium BC, using neutron tomography. Differences in the neutron attenuation coefficients of the constituent [...] Read more.
We studied the structural features in internal pores and mineral inclusions of several vessels from the Shubarat-1 archeological site in the Republic of Kazakhstan, dating to the late first millennium BC, using neutron tomography. Differences in the neutron attenuation coefficients of the constituent elements of pottery objects, as well as the high penetration capability of neutron tomography, make it possible to conduct non-destructive studies of rare ceramic vessels. By analyzing the three-dimensional tomography data, we can reconstruct the size and morphological parameters of internal pores and minerals. Based on these structural findings, we can clarify past pottery production processes. Full article
(This article belongs to the Section Image and Video Processing)
Show Figures

Figure 1

20 pages, 3571 KB  
Article
Concentration- and Sequence-Dependent MRI Signal Intensity Behavior of Ilex paraguariensis Aqueous Extract in MRCP-like Sequences: A Preclinical Phantom Study
by Mario J. Noh-Burgos, Juan B. Chalé-Dzul, Leticia Olivera-Castillo, César Puerto-Castillo, Nina Méndez-Domínguez and Rosa E. Moo-Puc
J. Imaging 2026, 12(7), 313; https://doi.org/10.3390/jimaging12070313 - 10 Jul 2026
Viewed by 687
Abstract
Magnetic resonance cholangiopancreatography (MRCP) is widely used for biliopancreatic imaging; however, hyperintense gastrointestinal fluids in heavily T2-weighted sequences may interfere with visualization of the biliary and pancreatic ducts. Natural manganese-containing beverages have been investigated in MRCP-related imaging contexts, and yerba mate (Ilex [...] Read more.
Magnetic resonance cholangiopancreatography (MRCP) is widely used for biliopancreatic imaging; however, hyperintense gastrointestinal fluids in heavily T2-weighted sequences may interfere with visualization of the biliary and pancreatic ducts. Natural manganese-containing beverages have been investigated in MRCP-related imaging contexts, and yerba mate (Ilex paraguariensis A. St.-Hil.) has been studied to this end. However, its concentration- and sequence-dependent signal behavior under MRCP-like phantom conditions remains insufficiently characterized. This preclinical phantom study evaluated the concentration- and sequence-dependent MRI signal intensity behavior of an aqueous extract of Ilex paraguariensis. The extract was characterized by means of elemental analysis, total manganese and iron quantification, total phenolic content, antioxidant capacity, and LC-ESI-MS analysis. MRI phantom experiments were run at different extract concentrations using T1-weighted, T2-weighted, and single-shot turbo spin echo (SSHTSE) sequences. The dried extract contained 1.22 ± 0.04 mg/g total manganese and 0.40 ± 0.01 mg/g total iron. Calculated total Mn concentrations in phantom dilutions ranged from 0.06 to 0.97 mg/dL. The extract showed concentration- and sequence-dependent signal behavior, with T1-weighted signal enhancement and progressive signal suppression in T2-weighted and SSHTSE sequences. No T1/T2 mapping or r1/r2 relaxivity measurements were performed. LC-ESI-MS identified MS1-based putatively assigned phenolic features without MS/MS confirmation of extract peaks. Ilex paraguariensis aqueous extract showed preliminary concentration- and sequence-dependent MRI signal intensity changes under phantom conditions, including signal suppression in MRCP-like heavily T2-weighted sequences. These findings do not establish clinical applicability, safety, tolerability, comparative efficacy, or improved duct visualization. Further studies are needed, incorporating relaxometric measurements, comparator agents, formulation assessment, in vivo evaluation, and clinical validation. Full article
(This article belongs to the Section Medical Imaging)
Show Figures

Graphical abstract

25 pages, 4900 KB  
Article
A2S2C-Det: Dual-Path Adaptive Aggregation with Spatial-Semantic Compensation for Strip Steel Surface Defect Detection
by Yange Sun, Mengdi Wang, Chenglong Xu, Huaping Guo, Li Zhang, Hongzhou Yue and Yan Feng
J. Imaging 2026, 12(7), 312; https://doi.org/10.3390/jimaging12070312 - 9 Jul 2026
Viewed by 301
Abstract
Accurate identification of surface defects on steel strips is critical for manufacturing quality assurance and operational reliability. Although deep learning has greatly advanced defect detection, precise recognition remains challenging due to significant background texture interference, loss of spatial details, and semantic imbalance across [...] Read more.
Accurate identification of surface defects on steel strips is critical for manufacturing quality assurance and operational reliability. Although deep learning has greatly advanced defect detection, precise recognition remains challenging due to significant background texture interference, loss of spatial details, and semantic imbalance across multiscale features. To address these challenges, we propose A2S2C-Det, a novel detector that integrates dual-path adaptive aggregation with spatial–semantic compensation to enhance feature representation for defect detection. First, we design a plug-and-play semantic refinement bottleneck (SRB) that augments backbone features through multiscale perception and a feature-screening bottleneck, enabling the model to suppress background interference while capturing subtle defect shapes. We further introduce a dual-path adaptive aggregation (DPAA) module that fuses complementary information from cross-level semantic consistency and fine-grained structural cues via two coordinated pathways, alleviating semantic imbalance across scales. Finally, we develop a spatial-semantic gated compensation (SSGC) module that adaptively supplies semantic information to low-level features while delivering spatial details to high-level features, recovering lost spatial details in high-level features. Extensive experiments on three benchmark datasets demonstrate that our A2S2C-Det achieves mAP50 of 82.0%, 73.0%, and 91.2%, and mAP of 47.1%, 36.5%, and 60.6%, respectively, comparing favorably against current state-of-the-art methods. Full article
Show Figures

Figure 1

17 pages, 7625 KB  
Article
Hitting the Gym with Fit3D: Benchmarking and Improving Monocular 3D Human Reconstruction on Extreme Fitness Motions
by Mihai Fieraru
J. Imaging 2026, 12(7), 311; https://doi.org/10.3390/jimaging12070311 - 9 Jul 2026
Viewed by 457
Abstract
Fitness motions present some of the most challenging cases for monocular 3D human reconstruction: extreme articulations, heavy self-occlusion, and frequent self-contact. The Fit3D dataset captures these motions at large scale via a 12-camera VICON motion capture system synchronized with 4 RGB cameras, but [...] Read more.
Fitness motions present some of the most challenging cases for monocular 3D human reconstruction: extreme articulations, heavy self-occlusion, and frequent self-contact. The Fit3D dataset captures these motions at large scale via a 12-camera VICON motion capture system synchronized with 4 RGB cameras, but in its original release provides only 3D skeletons. This paper contributes three studies built on top of Fit3D. First, we design and validate a methodology for constructing dense GHUM and SMPL-X pseudo-ground-truth shape and pose annotations on top of the raw MoCap: an optimization-based fitting pipeline that combines markers, multi-view 2D keypoints, separate body and hand normalizing-flow priors, and a self-collision loss, which we show improves on the marker-only MoSh++ baseline on the hands and extremities. Second, we define a standardized evaluation protocol—metric set, frame sampling, and coordinate conventions—for monocular 3D human reconstruction on Fit3D, served through the IMAR-hosted Fit3D evaluation resource, and use it to conduct a comparative benchmark study of 19 representative methods spanning optimization-based, single-frame, and video-based families; the two trained on Fit3D (NLF and SMPLest-X) lead the position and orientation metrics, respectively. Third, a controlled fine-tuning experiment shows that adding Fit3D to the training mixture of a strong baseline (HMR2.0) sharply lowers error on the hardest fitness poses without degrading out-of-domain generalization. The Fit3D dataset and the GHUM/SMPL-X annotations are available, under a non-commercial research license, through the IMAR Fit3D resource; as of June 2026, 1088 academics have registered for access. Full article
(This article belongs to the Section Computer Vision and Pattern Recognition)
Show Figures

Figure 1

22 pages, 5248 KB  
Article
Echo Model Analysis and Frequency-Domain Imaging Algorithm for Geosynchronous Spaceborne–Airborne FMCW Bistatic SAR with High-Maneuvering Receiver
by Xinyu Liu, Li Ding, Chenlei Lu, Wenlong Yang and Ping Li
J. Imaging 2026, 12(7), 310; https://doi.org/10.3390/jimaging12070310 - 8 Jul 2026
Viewed by 197
Abstract
Geosynchronous spaceborne–airborne frequency-modulated continuous-wave bistatic synthetic aperture radar (GEO SA FMCW BiSAR) offers cost-effective and persistent target monitoring. However, both the maneuvers of the receiver during the signal propagation delay and the continuous movements of the radar platforms within the sweep complicate the [...] Read more.
Geosynchronous spaceborne–airborne frequency-modulated continuous-wave bistatic synthetic aperture radar (GEO SA FMCW BiSAR) offers cost-effective and persistent target monitoring. However, both the maneuvers of the receiver during the signal propagation delay and the continuous movements of the radar platforms within the sweep complicate the received echo signal. These factors invalidate the “stop-and-go” assumption, which presumes constant-velocity motion. This paper proposes an echo model that simultaneously considers intra-pulse motion and accelerated motion of the high-maneuvering receiver. The introduction of receiver acceleration leads to nonlinear range terms in the bistatic range history, which will degrade the focusing performance if not properly compensated. Since the acceleration term is a small second-order quantity relative to the time delay, it is approximated by segmenting the aperture and applying the “stop-and-go” assumption within each sub-aperture. After dechirp, the two-dimensional (2-D) spectrum for imaging is derived by applying the principle of stationary phase and determining the azimuth stationary phase point via series reversion. Finally, imaging is achieved by azimuth compression, range cell migration correction, and secondary range compression. Simulation results demonstrate that the proposed algorithm achieves well-focused images while maintaining computational efficiency. Full article
(This article belongs to the Section Image and Video Processing)
Show Figures

Figure 1

27 pages, 3323 KB  
Review
Hybrid Imaging in Industrial Applications: A Review of Principles and Deployment
by Andrzej Burghardt, Piotr Garbacz and Magdalena Muszyńska
J. Imaging 2026, 12(7), 309; https://doi.org/10.3390/jimaging12070309 - 8 Jul 2026
Viewed by 317
Abstract
Hybrid imaging methods are emerging as one of the most dynamically evolving research areas in industrial inspection systems. This paper presents a literature review covering relevant scientific publications and official reports on the use of multimodal approaches in quality inspection and NDT systems. [...] Read more.
Hybrid imaging methods are emerging as one of the most dynamically evolving research areas in industrial inspection systems. This paper presents a literature review covering relevant scientific publications and official reports on the use of multimodal approaches in quality inspection and NDT systems. Hybrid imaging involves combining two or more imaging techniques to enhance the detection, characterization, and interpretation of features in inspected objects. The paper describes the physical foundations of vision-based inspection systems, including the interaction of optical radiation with matter. It also introduces a classification of optical methods and discusses the role of image fusion in multimodal data processing, with particular emphasis on high-speed quality control systems. The review outlines the current capabilities, limitations, and industrial applications of hybrid imaging, as well as future research directions, including integration with real-time systems and the use of artificial intelligence for automated defect interpretation. Full article
(This article belongs to the Section Computer Vision and Pattern Recognition)
Show Figures

Figure 1

22 pages, 6335 KB  
Article
PIP-PACA: An Interpretable Image Classification Framework via Prototype-Aware Clustering Attention
by Xinyuan Jia, Yanling Li and Yihui Wang
J. Imaging 2026, 12(7), 308; https://doi.org/10.3390/jimaging12070308 - 8 Jul 2026
Viewed by 317
Abstract
Image classification interpretability remains a fundamental challenge in the field of computer vision. Despite the remarkable improvements achieved by deep neural networks in classification accuracy, their decision-making processes are often opaque, which limits their applicability in high-stakes scenarios requiring reliability and transparency. Prototype-based [...] Read more.
Image classification interpretability remains a fundamental challenge in the field of computer vision. Despite the remarkable improvements achieved by deep neural networks in classification accuracy, their decision-making processes are often opaque, which limits their applicability in high-stakes scenarios requiring reliability and transparency. Prototype-based methods, such as PIP-Net, address this issue by establishing explicit correspondences between input images and semantic prototypes, thereby enabling an intuitive, evidence-based reasoning paradigm. However, these approaches still suffer from insufficient global context modeling and underutilization of structural relationships among prototypes. To address these limitations, this paper proposes an interpretable image classification model termed PIP-PACA, which is built upon a prototype-aware clustering attention mechanism. In contrast to conventional Transformer architectures based on self-attention, the proposed PACA module introduces a set of learnable cluster centers to project feature representations into a prototype space. Global information is then captured via a bidirectional attention mechanism between features and prototypes. This design is inherently aligned with the principles of prototype learning while reducing the computational complexity from quadratic to linear. Furthermore, a normalization operation is incorporated during the feature extraction stage to enhance the stability of feature distributions and improve the reliability of prototype matching. Extensive experimental results demonstrate that the proposed method not only preserves the interpretability of the original framework but also achieves notable improvements in classification accuracy, sparsity, and prototype purity. These findings validate the effectiveness and superiority of the clustering-based attention mechanism within the prototype learning paradigm. Full article
(This article belongs to the Topic Computer Vision and Image Processing, 3rd Edition)
Show Figures

Figure 1

16 pages, 5773 KB  
Article
Deep Learning-Based Multi-Class Pediatric Wrist Fracture Subtype Classification: A Pilot Study Comparing Convolutional Neural Network Architectures
by Rohan A. Phadke, Samer G. Salman, Zane G. Salman, Sai M. Yedupati, Joshua Ong, Alireza Tavakkoli, Sainyam Galhotra, Ajay Tripuraneni and James Rizkalla
J. Imaging 2026, 12(7), 307; https://doi.org/10.3390/jimaging12070307 - 8 Jul 2026
Viewed by 465
Abstract
Pediatric wrist fractures are among the most prevalent musculoskeletal injuries in children. Fracture subtype, including buckle/torus, greenstick, and Salter–Harris physeal injuries, directly influences management and prognosis. Subspecialty radiographic expertise required for subtype classification is not universally available in emergency or resource-limited settings. Deep [...] Read more.
Pediatric wrist fractures are among the most prevalent musculoskeletal injuries in children. Fracture subtype, including buckle/torus, greenstick, and Salter–Harris physeal injuries, directly influences management and prognosis. Subspecialty radiographic expertise required for subtype classification is not universally available in emergency or resource-limited settings. Deep learning (DL) offers an automated approach to fracture subtype recognition from plain radiographs. This pilot study evaluated convolutional neural network (CNN)-based five-class pediatric wrist fracture classification using the GRAZPEDWRI-DX dataset.A total of 940 pediatric wrist radiographs from GRAZPEDWRI-DX (figshare ID 14825193) were labeled using Arbeitsgemeinschaft fur Osteosynthesefragen (AO) pediatric codes into five classes: no fracture, buckle/torus, greenstick, Salter–Harris physeal fracture, and other fracture. Contrast-limited adaptive histogram equalization (CLAHE) and letterbox resizing to 224 × 224 pixels were applied. Patient-level stratified splits (70/15/15%) prevented data leakage. Three ImageNet-pretrained architectures (DenseNet-169, ResNet-50, and EfficientNet-B4) underwent two-phase transfer learning. Performance was assessed by balanced accuracy, macro F1, macro area under the receiver operating characteristic curve (AUROC), and Cohen’s kappa.DenseNet-169 achieved the highest balanced accuracy (0.371; 95% confidence interval [CI]: 0.289–0.448), macro F1 (0.334; 95% CI: 0.251–0.416), and macro AUROC (0.669), with Cohen’s kappa of 0.269 on the held-out test set (n = 139) under initial five-epoch pilot training conditions. All three networks exceeded a majority-class (no-information) baseline (balanced accuracy 0.20). Extending training to 50 epochs (approximately 2100 mini-batch iterations) with GPU acceleration substantially improved DenseNet-169 to a balanced accuracy of 0.532 (95% CI: 0.451–0.614), macro F1 of 0.516, and macro AUROC of 0.815, with statistically significant pairwise architecture differences (McNemar p < 0.01); per-class sensitivity was highest for no-fracture detection (0.969) and lowest for buckle/torus fractures (0.393). Gradient-weighted class activation mapping (Grad-CAM) confirmed anatomically coherent model saliency at the distal radial metaphysis and physeal plate.DenseNet-169 achieved the best five-class classification performance among evaluated architectures under pilot training conditions, and extended training substantially improved accuracy, although classification accuracy remained below clinically usable thresholds. These results establish a reproducible, patient-stratified DL pipeline and a benchmark for full-dataset training and future methodological development, rather than a clinically deployable tool. Full article
(This article belongs to the Special Issue Medical Computer Vision: Innovations and Clinical Impact)
Show Figures

Graphical abstract

18 pages, 1618 KB  
Article
BDKD-Net: Boundary-Probability Knowledge Distillation for Compact Polyp Segmentation
by Tian Xia, Jianhua Li and Liping Sun
J. Imaging 2026, 12(7), 306; https://doi.org/10.3390/jimaging12070306 - 8 Jul 2026
Viewed by 326
Abstract
Accurate polyp segmentation in colonoscopy supports early detection of colorectal cancer, but compact models under a student-only inference budget tend to lose boundary fidelity. BDKD-Net is a 3.72M-parameter compact student trained under a composite knowledge-distillation loss that combines a response signal, a boundary-probability [...] Read more.
Accurate polyp segmentation in colonoscopy supports early detection of colorectal cancer, but compact models under a student-only inference budget tend to lose boundary fidelity. BDKD-Net is a 3.72M-parameter compact student trained under a composite knowledge-distillation loss that combines a response signal, a boundary-probability signal restricted to a static teacher-derived boundary band, and an auxiliary detail-feature alignment; the teacher is used only during training and discarded at inference. In a uniform five-seed re-run on a locked Kvasir-SEG and CVC-ClinicDB development split, the same-architecture scratch student reaches Dev Dice 0.9210 ± 0.0029 and boundary F1 at 3-pixel tolerance 0.7726 ± 0.0057, while full BDKD-Net reaches 0.9300 ± 0.0029 and 0.8008 ± 0.0094. Boundary-probability distillation is the load-bearing signal: it has the highest mean Dev BF1t3 among the single KD signals and carries the boundary gain at no inference cost, with the full model reaching four-external Dice 0.8225 ± 0.0071 at 1.76 GFLOPs and 140.9 FPS. Among the directly reproduced baselines, the higher-Dice Polyp-PVT (0.8417 ± 0.0091) needs 6.8× the parameters and 5.7× the compute at roughly half the frame rate. BDKD-Net thus delivers a compact, boundary-faithful student that keeps most of the accuracy of much larger models at a fraction of their inference cost. Full article
(This article belongs to the Section Medical Imaging)
Show Figures

Graphical abstract

22 pages, 2316 KB  
Article
Attention-Enhanced Pedestrian Trajectory Prediction via Compressed Point Cloud Representation
by Yuting Han, Shuyu Li and Yunfei Tan
J. Imaging 2026, 12(7), 305; https://doi.org/10.3390/jimaging12070305 - 7 Jul 2026
Viewed by 292
Abstract
To address the high storage overhead and inadequate spatial geometric representation associated with raw point cloud data in multi-pedestrian trajectory prediction, a compressed point cloud-based and attention-enhanced trajectory prediction method (CPCAE) is proposed in the paper. First, for input raw point cloud, a [...] Read more.
To address the high storage overhead and inadequate spatial geometric representation associated with raw point cloud data in multi-pedestrian trajectory prediction, a compressed point cloud-based and attention-enhanced trajectory prediction method (CPCAE) is proposed in the paper. First, for input raw point cloud, a lossy compression module is designed, which improves the Depoco framework by introducing a multi-feature extraction component and employing a coordinate decomposition strategy to optimize compression quality and spatial representation. For input video frames of pedestrians, spatial features are extracted using a 2D convolutional network, and dynamic interactions among pedestrians are captured by a Transformer-based encoder. Then, both spatial attention and modal attention mechanisms are incorporated to dynamically balance the contributions of two modal features and precisely identify key regions and positions. Experimental results evaluate the proposed framework from the perspectives of point cloud compression and downstream trajectory prediction. The results demonstrate that compressed point cloud representations can support competitive trajectory prediction performance in CPCAE. Full article
(This article belongs to the Section Computer Vision and Pattern Recognition)
Show Figures

Figure 1

Previous Issue
Next Issue
Back to TopTop