-
A Method for Paired Comparisons of Glo Germ Quantity in Images of Hands Before and After Washing -
Artificial Intelligence in Pulmonary Endoscopy: Current Evidence, Limitations, and Future Directions -
AI-Based Osteoporosis Detection on Dental Radiographs -
Implementation of Image-Based AI Is Associated with Increased Case Volume in a High-Acuity, 15-Room Cardiothoracic Suite at a Tertiary Academic Hospital
Journal Description
Journal of Imaging
Journal of Imaging
is an international, multi/interdisciplinary, peer-reviewed, open access journal of imaging techniques, published online monthly by MDPI.
- Open Accessfree for readers, with article processing charges (APC) paid by authors or their institutions.
- High Visibility: indexed within Scopus, ESCI (Web of Science), PubMed, PMC, dblp, Inspec, Ei Compendex, and other databases.
- Journal Rank: JCR - Q2 (Imaging Science and Photographic Technology) / CiteScore - Q1 (Radiology, Nuclear Medicine and Imaging)
- Rapid Publication: manuscripts are peer-reviewed and a first decision is provided to authors approximately 21.3 days after submission; acceptance to publication is undertaken in 3.6 days (median values for papers published in this journal in the first half of 2026).
- Recognition of Reviewers: reviewers who provide timely, thorough peer-review reports receive vouchers entitling them to a discount on the APC of their next publication in any MDPI journal, in appreciation of the work done.
Impact Factor:
3.8 (2025);
5-Year Impact Factor:
3.6 (2025)
Latest Articles
Attention-Enhanced Multi-Scale Feature-Wise Linear Modulation for Fine-Grained Poisonous Mushroom Image Recognition
J. Imaging 2026, 12(8), 398; https://doi.org/10.3390/jimaging12080398 - 21 Aug 2026
Abstract
Fine-grained poisonous mushroom recognition in natural scenes is challenging because of complex backgrounds, subtle morphological differences, and the limited interpretability of model decisions. To address these challenges, this paper proposes Att-FiLM, an attention-enhanced multi-scale Feature-Wise Linear Modulation network for poisonous mushroom image recognition.
[...] Read more.
Fine-grained poisonous mushroom recognition in natural scenes is challenging because of complex backgrounds, subtle morphological differences, and the limited interpretability of model decisions. To address these challenges, this paper proposes Att-FiLM, an attention-enhanced multi-scale Feature-Wise Linear Modulation network for poisonous mushroom image recognition. The model adopts an asymmetric dual-backbone architecture in which a frozen ConvNeXt-Base branch provides global semantic priors, while a trainable EfficientNet-B0 branch learns local discriminative features. Rather than directly concatenating heterogeneous features, Att-FiLM generates scale and shift parameters from semantic features and performs channel-wise modulation on multi-scale EfficientNet features at Stage 2 and Stage 4. This mechanism enables global semantic information to guide local feature learning while reducing feature redundancy and semantic inconsistency. Experimental results show that Att-FiLM achieves an Accuracy of 95.58% and an F1-score of 0.9455 on the poisonous/edible binary classification task. On the 190-class species-level classification task, it achieves a Top-1 Accuracy of 93.63% and a Macro-F1 of 0.9347. Interpretability analysis further shows that decision-relevant responses are frequently associated with morphologically relevant regions, including gills, annuli, volvae, and cap textures. These results indicate that Att-FiLM provides effective recognition performance together with interpretable decision evidence for mushroom recognition in complex natural scenes.
Full article
(This article belongs to the Section Image and Video Processing)
Open AccessArticle
Landmark Recognition Beyond Curated Benchmarks: Cross-Domain Evaluation of a Multi-Threshold Selective YOLO11 Ensemble on User-Generated Imagery, with a Zero-Shot Multimodal LLM Baseline
by
Ulugbek Hudayberdiev, Abdimumin Alikulov, Adkham Israilov, Muhiddin Xidirov and Javokhir Musaev
J. Imaging 2026, 12(8), 397; https://doi.org/10.3390/jimaging12080397 - 21 Aug 2026
Abstract
Landmark recognition for smart tourism is usually validated on curated benchmark images. In deployment, however, the classifier must handle user-generated photographs whose viewpoint, lighting, resolution, occlusion, and compression differ sharply from curated data. This paper evaluates a previously published multi-threshold enhancement and selective
[...] Read more.
Landmark recognition for smart tourism is usually validated on curated benchmark images. In deployment, however, the classifier must handle user-generated photographs whose viewpoint, lighting, resolution, occlusion, and compression differ sharply from curated data. This paper evaluates a previously published multi-threshold enhancement and selective YOLO11n-cls ensemble under this shift, and provides a preliminary zero-shot comparison of three general-purpose multimodal large language models (MLLMs) on the same task. To measure the shift, we build Samarkand v2-SNS, a 300-image out-of-distribution test set of social-media photographs of 12 Samarkand landmarks, disjoint from the training and validation data. Under the shift, four supervised baselines fall by 12.73–22.08 percentage points to 73–80% accuracy, and their in-distribution ranking does not hold. The selective ensemble degrades least (99.24% to 93.00%, −6.24 points) and outperforms the strongest baseline by 13 points. A capacity-matched ablation shows that most of this robustness comes from enhancement diversity, not from generic ensembling. In a preliminary comparison, zero-shot MLLMs (GPT-5, Claude Sonnet 4.5, Gemini 2.5) reach only 24.81–54.26%, far below deployment needs. The results argue for reporting out-of-distribution accuracy alongside curated benchmarks, and for hybrid systems that pair compact specialised recognisers with MLLM-based interpretation.
Full article
(This article belongs to the Section Computer Vision and Pattern Recognition)
►▼
Show Figures

Figure 1
Open AccessArticle
Identity Document Presentation Attack Detection in Visible Light with Illumination-Controlled Scanner
by
Lada Tolstenko, Alexey Popkov, Irina Kunina, Dmitry Polevoy and Sergey Usilin
J. Imaging 2026, 12(8), 396; https://doi.org/10.3390/jimaging12080396 - 20 Aug 2026
Abstract
A reliable sign of the absence of a document presentation attack, when a print copy is presented instead of the original document, is the presence of such security features as OVDs (Optical Variable Devices), for example, holograms. To check the presence of holograms,
[...] Read more.
A reliable sign of the absence of a document presentation attack, when a print copy is presented instead of the original document, is the presence of such security features as OVDs (Optical Variable Devices), for example, holograms. To check the presence of holograms, it is sufficient to use the visible light and a series of document images captured with a varying angle of incidence and reflection of light. It can be achieved either by changing the position of the document or by changing the position of the illumination source. This work proposes a method for detecting holograms on identity documents using a scanner with controlled illumination. The method is based on obtaining a series of document images in various illumination modes and identifying features characteristic of holograms. To test the method, a dataset MIDV-Holo-Scan was collected by scanning physical documents used in the creation of the open dataset MIDV-Holo. It includes both documents with holograms, accepted in this work as originals, and documents without holograms, simulating an attack on document presentation. The proposed method for detecting attacks on document presentation achieves a quality of Accuracy = 100%, which surpasses the quality of the baseline method published with the MIDV-Holo dataset.
Full article
(This article belongs to the Topic Image Processing, Signal Processing and Their Applications)
►▼
Show Figures

Figure 1
Open AccessArticle
Imaging-Based Parallel Seismic Test for In-Service Bridge Pile Foundations
by
Zhichao Luo, Weibin Luo, Peng Wang, Yuan Gu and Peimin Zhu
J. Imaging 2026, 12(8), 395; https://doi.org/10.3390/jimaging12080395 - 20 Aug 2026
Abstract
The conventional parallel seismic test (PST) is a widely used and highly reliable method for determining lengths of in-service bridge pile foundations. However, it relies solely on the manual interpretation of first-arrival in seismic records and often fails to detect, characterize, or geometrically
[...] Read more.
The conventional parallel seismic test (PST) is a widely used and highly reliable method for determining lengths of in-service bridge pile foundations. However, it relies solely on the manual interpretation of first-arrival in seismic records and often fails to detect, characterize, or geometrically define internal defects within the piles. To overcome these limitations and enable the intuitive identification of internal defects and damage within piles, this study introduces an elastic reverse time migration (ERTM) imaging algorithm based on the spectral element method, achieving high-resolution imaging of existing bridge pile foundations and their defects and damage. To suppress crosstalk between P- and S-waves during the ERTM process, a wavefield decoupling method is employed to separate the elastic wavefield into P- and S-wave components for independent imaging. Two-dimensional numerical testing on a bridge pile model with a necking defect demonstrates that ERTM can effectively image both the pile geometry and its defects, significantly improving the capability of defect detection and characterization. This approach provides more intuitive visualization for assessing the structural integrity of in-service bridge pile foundations.
Full article
(This article belongs to the Section Image and Video Processing)
►▼
Show Figures

Figure 1
Open AccessArticle
Structure-Prior-Guided Multi-Stage Cross-Modal Collaborative Network for RGB-D Semantic Segmentation
by
Yifan Yu, Zhiwei Zhong, Fan Min and Song Deng
J. Imaging 2026, 12(8), 394; https://doi.org/10.3390/jimaging12080394 - 20 Aug 2026
Abstract
Red–green–blue and depth (RGB-D) semantic segmentation combines appearance cues from RGB images with geometric information from depth maps, but sensor noise, missing measurements, and boundary-inconsistent depth responses can introduce conflicting evidence during cross-modal fusion. We propose the Structure-Prior-Guided Network (SPGNet), a dual-branch, multi-stage
[...] Read more.
Red–green–blue and depth (RGB-D) semantic segmentation combines appearance cues from RGB images with geometric information from depth maps, but sensor noise, missing measurements, and boundary-inconsistent depth responses can introduce conflicting evidence during cross-modal fusion. We propose the Structure-Prior-Guided Network (SPGNet), a dual-branch, multi-stage framework that follows a correction-before-fusion strategy. At each feature scale, SPGNet estimates a learned structure prior from cross-modal agreement and discrepancy. The Cross-Modal Correction Module (CCM) uses this prior to regulate bidirectional information transfer, suppressing unreliable responses while retaining complementary cues. The Dual-branch Enhancement Fusion Module (DEF) then enhances the corrected RGB and depth features and integrates them through shared-representation-guided interaction, after which a lightweight multi-scale decoder produces the segmentation output. Under a unified training and evaluation protocol, SPGNet achieved three-run mean Intersection over Union (mIoU) scores of 50.845% on NYU Depth V2 and 48.457% on SUN RGB-D. Compared with the best reproduced baseline on each dataset, SPGNet improved mean mIoU by 2.111 and 0.899 percentage points, respectively. These results suggest that separating reliability-oriented correction from multimodal fusion can limit the propagation of unreliable cross-modal responses and improve indoor RGB-D semantic segmentation performance.
Full article
(This article belongs to the Section AI in Imaging)
►▼
Show Figures

Figure 1
Open AccessArticle
Turning Immersive Viewers into Analytical Workspaces: ASCRIBE-XR and Agent-Driven Scientific Visualization
by
Ronald Pandolfi, Luke Weidner, James Sethian, Jeffrey Donatelli and Daniela Ushizima
J. Imaging 2026, 12(8), 393; https://doi.org/10.3390/jimaging12080393 - 20 Aug 2026
Abstract
Scientific visualization is changing from passive observation to active, AI-assisted collaboration. While Extended Reality (XR) has proven valuable for comprehending dense 3D arrays, traditional VR applications are typically deployed in rigid, single-purpose, and monolithic architectures. In this paper, we present the evolution of
[...] Read more.
Scientific visualization is changing from passive observation to active, AI-assisted collaboration. While Extended Reality (XR) has proven valuable for comprehending dense 3D arrays, traditional VR applications are typically deployed in rigid, single-purpose, and monolithic architectures. In this paper, we present the evolution of ASCRIBE-XR: a virtual reality platform backed by remote computation that has been re-engineered into a dynamic, service-oriented ecosystem. We introduce three core innovations that make immersive data analysis easier, faster, and more flexible when using multimodal scientific imaging. First, a lightweight Python REST interface decouples XR logic from the rendering engine, enabling real-time, programmable scene customization and on-demand data generation. Second, we present a Specimen Catalog architecture that lets the platform pivot between radically different disciplines, ranging from archaeological heterogeneous concrete and fuel-cell membranes to the root system of a bioenergy grass, by describing each dataset through portable metadata rather than hard-coded application logic. Finally, we introduce a prompt-driven layer powered by the Claude Agent SDK, allowing researchers to generate, segment, and manipulate volumetric and mesh data through natural language dialogue within the virtual space. For example, applying foundation models such as the Segment Anything Model (SAM) to perform zero-shot segmentation on demand. By bridging human intent with remote computation, ASCRIBE-XR relaxes the constraints of conventional visualization tools, offering a highly adaptable, conversational platform for scientific discovery with human auditing.
Full article
(This article belongs to the Section AI in Imaging)
►▼
Show Figures

Figure 1
Open AccessArticle
A Stability Atlas for IBSI Radiomics Features Using Synthetic Digital Phantoms, with Proof-of-Concept Physics-Based Normalisation
by
Shuji Yamamoto
J. Imaging 2026, 12(8), 392; https://doi.org/10.3390/jimaging12080392 - 20 Aug 2026
Abstract
Radiomics features are strongly sensitive to image acquisition, and separating that sensitivity from biological signal usually requires repeated patient scans that cannot be shared. We present an open, fully synthetic framework (radiomics-phantom) that maps and, as a proof of concept, corrects radiomics feature
[...] Read more.
Radiomics features are strongly sensitive to image acquisition, and separating that sensitivity from biological signal usually requires repeated patient scans that cannot be shared. We present an open, fully synthetic framework (radiomics-phantom) that maps and, as a proof of concept, corrects radiomics feature instability without any patient data. Deterministic three-dimensional texture phantoms are generated as anisotropic Gaussian random fields with known ground truth and an optional embedded lesion. An independently implemented feature core aligned with the Image Biomarker Standardization Initiative (IBSI) covers all eleven IBSI-1 feature families and matched all 482 published digital-phantom benchmark values within the applicable tolerances. An image-domain acquisition simulator applies point-spread blur, slice-profile averaging, dose-scaled correlated noise, resampling, and quantisation. Per-feature reproducibility across a sweep of fifteen textures (varying correlation length, anisotropy, and intensity scale) by nine acquisition conditions, with five independent noise realisations per stochastic setting, is summarised by the absolute-agreement intraclass correlation ICC(2,1), with a realisation-aware percentile-bootstrap 95% confidence interval for every estimate; constant features are excluded from estimation. Values span nearly the full range (median 0.13, 95% CI 0.03–0.19), and a hierarchical variance decomposition attributes a median 77% of per-feature variance to the acquisition condition and under 1% to stochastic realisation; the values are interpreted as exploratory rankings within this acquisition envelope. As a proof of concept, intensity variance and grey-level co-occurrence contrast under additive Gaussian noise were normalised using calibrated, invertible response models, returning them to their noiseless values on held-out data (median error below 4% across five textures and repeated noise realisations, and about 11% when the noise level is estimated from the degraded image itself), while features the models cannot describe are refused rather than corrected. All code and a 716-test suite are released openly and archived on Zenodo. The result is a reproducible, patient-data-free testbed for radiomics feature stability.
Full article
(This article belongs to the Section Medical Imaging)
►▼
Show Figures

Figure 1
Open AccessArticle
Research on Medical Image Super-Resolution Reconstruction Algorithm Based on Dilated Convolution and Multi-Module Fusion
by
Zhuye Xu and Yucong Guo
J. Imaging 2026, 12(8), 391; https://doi.org/10.3390/jimaging12080391 - 19 Aug 2026
Abstract
Medical image resolution plays a crucial role in early disease detection and fine-structure observation. Super-resolution reconstruction technology can restore low-resolution images to high-resolution versions, thereby assisting physicians in making accurate diagnoses. To address challenges in medical image super-resolution reconstruction, including insufficient global information
[...] Read more.
Medical image resolution plays a crucial role in early disease detection and fine-structure observation. Super-resolution reconstruction technology can restore low-resolution images to high-resolution versions, thereby assisting physicians in making accurate diagnoses. To address challenges in medical image super-resolution reconstruction, including insufficient global information acquisition, excessive network complexity, and suboptimal loss function adaptation for medical imaging data, this paper proposes an image super-resolution reconstruction algorithm named IDCASR-MMF based on improved dilated convolution and multi-module fusion. First, multi-dilation-rate dilated convolution is introduced to expand the receptive field and integrated with a spatial attention mechanism to dynamically calibrate high-frequency features after feature extraction. Subsequently, the Squeeze-and-Excitation module is fused with dilated convolution as a channel attention mechanism to streamline the network architecture. Finally, a weighted fusion strategy combining adversarial loss and MSE loss is adopted, where the dynamic adjustment of weighting coefficients balances pixel-level structural accuracy and high-frequency detail authenticity, achieving synergistic optimization of objective precision and subjective quality for medical images. To validate the effectiveness of the proposed algorithm, IDCASR-MMF is compared with 11 state-of-the-art methods across five datasets (Set5, Set14, BSD100, Urban100, and Bone FD). Experimental results demonstrate that the proposed algorithm achieves superior PSNR and SSIM values on multiple datasets, confirming that IDCASR-MMF can effectively reconstruct high-resolution medical images from low-resolution inputs.
Full article
(This article belongs to the Section Medical Imaging)
►▼
Show Figures

Figure 1
Open AccessCommunication
Four-Dimensional Cine Cinematic Rendering of Structural Heart and Mechanical Circulatory Support Devices: An Illustrative Technical Experience
by
Amy Avakian and Muhammad Umair
J. Imaging 2026, 12(8), 390; https://doi.org/10.3390/jimaging12080390 - 19 Aug 2026
Abstract
Patients with implanted cardiac devices are a rapidly growing imaging population, and electrocardiogram-gated cardiac computed tomography (CT) is increasingly used to characterize device geometry, multi-device relationships, and dynamic behavior across the cardiac cycle. Cinematic rendering (CR) is a photorealistic three-dimensional (3D) visualization technique
[...] Read more.
Patients with implanted cardiac devices are a rapidly growing imaging population, and electrocardiogram-gated cardiac computed tomography (CT) is increasingly used to characterize device geometry, multi-device relationships, and dynamic behavior across the cardiac cycle. Cinematic rendering (CR) is a photorealistic three-dimensional (3D) visualization technique for cardiac CT whose established contribution in this population is communicative: it conveys 3D device geometry and material distinctions within a single rendered volume. We describe a demonstrative case series extending CR across the cardiac cycle—time-resolved “4D cine” CR—to depict dynamic device behavior and time-resolved multi-device interaction in a single volume; this is an illustrative technical experience rather than a systematic evaluation of diagnostic performance. Illustrative examples include an EVOQUE transcatheter tricuspid valve rendered together with concurrent surgical mitral and transcatheter aortic valves, a left atrial appendage occlusion device, a normally positioned Impella catheter, and a HeartMate 3 left ventricular assist device (LVAD). Across cases, 4D cine CR feasibility scaled inversely with metallic burden—the aggregate volume and radiodensity of metallic device components within the scan field—with renderings informative for low-metal nitinol and catheter devices but substantially degraded by streak artifact in high-metal LVAD housings. This relationship was observed qualitatively in a small selected series and is offered as an initial observation rather than an established characteristic of the technique. We discuss current limitations and emerging directions such as photon-counting detector CT, metal artifact reduction, and artificial-intelligence-assisted post-processing that may extend 4D cine CR in this population.
Full article
(This article belongs to the Section Medical Imaging)
►▼
Show Figures

Figure 1
Open AccessArticle
PRMEFNet: A Real-Time Unsupervised Multi-Exposure Fusion Network Driven by Prior Knowledge
by
Junwei Qi, Hangdong Wang, Xu Xiao and Jingpeng Gao
J. Imaging 2026, 12(8), 389; https://doi.org/10.3390/jimaging12080389 - 19 Aug 2026
Abstract
Due to the limited dynamic range of imaging sensors, most cameras can only capture low-dynamic-range (LDR) images. Multi-exposure fusion (MEF) is an effective technique for generating high-dynamic-range (HDR) images. However, to simultaneously preserve texture details and global exposure, most existing methods primarily rely
[...] Read more.
Due to the limited dynamic range of imaging sensors, most cameras can only capture low-dynamic-range (LDR) images. Multi-exposure fusion (MEF) is an effective technique for generating high-dynamic-range (HDR) images. However, to simultaneously preserve texture details and global exposure, most existing methods primarily rely on more complex models to improve performance, resulting in higher computational costs and longer processing times. To address this issue, we propose a real-time unsupervised MEF network driven by prior knowledge. To this end, a hierarchical feature extraction module is designed that utilizes filtering operations to decompose the source images into base layers and detail layers. Features are extracted from each layer separately to reduce the difficulty of extracting effective features. Then, the receptive field of feature maps is expanded by dilated convolutions, and a window-based self-attention mechanism is applied to perform context modeling, achieving effective contextual modeling with low computational cost. Subsequently, texture features and global features are extracted separately to enable the model to maintain both local texture clarity and global smoothness. In addition, a one-dimensional lookup table is utilized to accelerate the inference process. Comprehensive experiments are conducted to verify the effectiveness of the proposed method. The subjective evaluation results demonstrate that the fused images exhibit superior visual quality, while objective experiments further quantify its superior performance, demonstrating that the proposed method effectively reduces computation time.
Full article
(This article belongs to the Topic Computational Imaging)
►▼
Show Figures

Figure 1
Open AccessArticle
Ordinal Deep Learning for Lumbar Foraminal Stenosis Grading on Sagittal MRI
by
Rohan A. Phadke, Samer G. Salman, Zane G. Salman, Akhil Marupudi, Kirtan Patel, Joshua Ong, Alireza Tavakkoli, Sainyam Galhotra, Ajay Tripuraneni, James Rizkalla and Nathan J. Lee
J. Imaging 2026, 12(8), 388; https://doi.org/10.3390/jimaging12080388 - 19 Aug 2026
Abstract
Lumbar foraminal stenosis grading contributes to surgical-level selection, but automated four-grade classification remains challenging. Published pipelines for this dataset reach approximately 65% four-class accuracy, and deep classifiers offer no anatomical rationale. We investigated whether interpretable, millimeter-scale morphometry from segmentation masks improves grading beyond
[...] Read more.
Lumbar foraminal stenosis grading contributes to surgical-level selection, but automated four-grade classification remains challenging. Published pipelines for this dataset reach approximately 65% four-class accuracy, and deep classifiers offer no anatomical rationale. We investigated whether interpretable, millimeter-scale morphometry from segmentation masks improves grading beyond a deep image model. We analyzed the LSS-MRI-AISSLab sagittal T2-weighted dataset (469 patients, 2979 expert-graded foramina spanning L1-L2 through L5-S1 bilaterally on a four-grade scale). Ten morphometric descriptors were computed from mid-sagittal polygon segmentations and scaled to millimeters using each patient’s recorded pixel spacing. A dual-branch network combined a fine-tuned ResNet-18 embedding of each foraminal region of interest with the morphometric vector through an ordinal regression head. Foraminal regions were supplied from expert bounding-box annotations; automated localization within the full sagittal examination was not evaluated. Four configurations (nominal softmax, appearance-only, anatomy-only, and fusion) were compared on a locked patient-level test set of 94 patients after five-fold cross-validation, with quadratic weighted kappa (QWK) as the primary endpoint and patient-clustered bootstrap inference. Feature-grade correlations were reported pooled and adjusted for lumbar level. Fusion achieved QWK 0.813 (95% confidence interval [CI] 0.769–0.847) and 75.3% four-class accuracy. Appearance-only was statistically indistinguishable (QWK 0.806; delta QWK +0.006, 95% CI −0.023 to 0.036, p = 0.68), whereas anatomy-only reached 0.444, and a level-and-side-only reference reached 0.314. Boundary discrimination was strong (area under the curve 0.92–0.99), 98.3% of predictions fell within one grade, and performance was consistent across scanner vendors. Morphometric associations were confounded by level: the apparent spondylolisthesis effect (rho −0.349) disappeared after adjustment (rho −0.000), while disc height, null when pooled (rho +0.021), emerged as a genuine within-level effect (rho −0.108). A fine-tuned ordinal image classifier achieved strong agreement for four-grade lumbar foraminal stenosis classification. The evaluated segmentation-derived morphometric features did not improve performance beyond imaging alone, and several apparent anatomic associations reflected confounding by lumbar level. External and prospective validation in complete clinical MRI workflows are needed before implementation.
Full article
(This article belongs to the Special Issue Medical Computer Vision: Innovations and Clinical Impact)
►▼
Show Figures

Figure 1
Open AccessCommunication
Quantitative MRI Susceptibility: Mapping of Transient Ischemic Attack—A Preliminary Study
by
Philipp Gruber, Michael Diepers, Markus Gschwind, Paul G. Unschuld, Luca Remonda, Franca Wagner, Pasquale Mordasini and Jatta Berberat
J. Imaging 2026, 12(8), 387; https://doi.org/10.3390/jimaging12080387 - 17 Aug 2026
Abstract
A transient ischemic attack (TIA) is a transient episode of neurological dysfunction without evidence of acute infarct demarcation on neuroimaging. The radiographic features of TIAs are uncertain, as since there is currently no perfect imaging method with which to differentiate a stroke from
[...] Read more.
A transient ischemic attack (TIA) is a transient episode of neurological dysfunction without evidence of acute infarct demarcation on neuroimaging. The radiographic features of TIAs are uncertain, as since there is currently no perfect imaging method with which to differentiate a stroke from a TIA. We aimed to evaluate whether regional abnormalities are associated with TIA by identifying tissue with increased susceptibility (χ) values using quantitative susceptibility mapping (QSM). A total of 39 adults with clinical TIA symptoms (66 ± 12 years) and 39 age- and sex matched controls (65 ± 12 years) underwent MRI. Based on the QSM-maps, local susceptibility changes were tested for statistically significant differences between patients with TIA and healthy controls. The susceptibility maps of TIA patients showed that magnetic susceptibility differed from that of healthy volunteers, with the strongest effects observable in the right lingual gyrus (p < 0.001) and bilaterally in the caudal anterior cingulate (cACC, p < 0.01). TIA is often difficult to diagnose at the time of presentation in the emergency department, and QSM could show an association between regional QSM abnormalities and TIA.
Full article
(This article belongs to the Special Issue Quantitative Neuroimaging: Advancing Diagnosis, Biomarkers, and Multimodal Fusion)
►▼
Show Figures

Figure 1
Open AccessArticle
Information Retention and Feature Screening Synergistic Network for Aviation Ground Safety and Protective Devices
by
Enming Wu, Mingxuan Wang, Runxia Guo, Jiusheng Chen, Jiaren Li, Fuyu Sun and Liyuan Ye
J. Imaging 2026, 12(8), 386; https://doi.org/10.3390/jimaging12080386 - 17 Aug 2026
Abstract
Aviation ground safety and protective devices are critical for flight safety; however, their unintentional retention on aircraft after maintenance remains a persistent risk. Existing deep learning-based approaches for aviation safety have predominantly followed a reactive paradigm, detecting FOD on runways or inspecting the
[...] Read more.
Aviation ground safety and protective devices are critical for flight safety; however, their unintentional retention on aircraft after maintenance remains a persistent risk. Existing deep learning-based approaches for aviation safety have predominantly followed a reactive paradigm, detecting FOD on runways or inspecting the aircraft for inadvertently retained tools post-maintenance. In contrast, this paper advocates a proactive philosophy: using a neural network to recognize and inventory all ground safety and protective devices immediately after maintenance closure, thereby preventing retention incidents at their source. However, realizing this proactive verification is technically challenging—object detection for these devices often suffers from loss of fine-grained detail due to downsampling and inherently sparse semantic information of the targets. To this end, we propose an Information Retention and Feature Screening Synergistic Network (RS-Net) grounded in information bottleneck theory. The network comprises a main branch that enhances discriminative features through attention-guided screening, and an auxiliary branch, used only during training, that preserves fine-grained spatial details via information-retentive convolutions. A Dual-State Region Refinement Module (DRM) provides configurable support for both branches, decoupling the conflicting objectives of background compression and detail preservation. Experiments on a self-constructed dataset collected from real airline maintenance operations demonstrate that RS-Net substantially outperforms the strong YOLOv9 baseline, achieving gains of 4.531% in F1-score, 2.533% in mAP0.5, and 1.429% in mAP0.5:0.95. Cross-dataset experiments further validate its strong generalization capability.
Full article
(This article belongs to the Topic Image Processing, Signal Processing and Their Applications)
►▼
Show Figures

Figure 1
Open AccessArticle
VLM-Assisted Routing and HU-Traceable DICOM Adaptation for Lumbar CT HU Measurement: A Deployment-Oriented Pilot Technical Evaluation
by
Zhe-Yu Ye, Jun-Mu Peng and Tamotsu Kamishima
J. Imaging 2026, 12(8), 385; https://doi.org/10.3390/jimaging12080385 - 14 Aug 2026
Abstract
Automated lumbar computed tomography (CT) Hounsfield unit (HU) measurement can support opportunistic osteoporosis screening, but heterogeneous Digital Imaging and Communications in Medicine (DICOM) inputs often disrupt automated workflows before measurement. We evaluated a local Qwen2.5-VL-7B vision-language model (VLM)-assisted front end for suitability routing,
[...] Read more.
Automated lumbar computed tomography (CT) Hounsfield unit (HU) measurement can support opportunistic osteoporosis screening, but heterogeneous Digital Imaging and Communications in Medicine (DICOM) inputs often disrupt automated workflows before measurement. We evaluated a local Qwen2.5-VL-7B vision-language model (VLM)-assisted front end for suitability routing, input planning, and HU-traceable DICOM-derived field-of-view/orientation adaptation upstream of an unchanged single-slice lumbar CT HU workflow. The VLM was restricted to routing and planning. In a 20-case FUJIFILM pilot cohort, observed HU output availability was higher with the front-end-assisted route than with direct processing (12/20 versus 5/20), while paired automatic-success cases showed excellent agreement with post hoc manual ImageJ (version 1.54g) measurements (ICC(A,1) = 0.997; MAE = 4.64 HU). Public DICOM stress testing further demonstrated improved workflow robustness across heterogeneous datasets. These findings support the use of a deployment-oriented front-end strategy to improve the auditability and robustness of automated lumbar CT HU measurement while preserving an unchanged downstream measurement workflow.
Full article
(This article belongs to the Section Medical Imaging)
►▼
Show Figures

Figure 1
Open AccessArticle
Understanding Trade-Offs in Continuous Neural Representations for Diffeomorphic Image Registration: A Comparative Study of Implicit Neural Representations and Neural Ordinary Differential Equations
by
Salvador Rodriguez-Sanz, Carlos Paesa-Lia and Monica Hernandez
J. Imaging 2026, 12(8), 384; https://doi.org/10.3390/jimaging12080384 - 14 Aug 2026
Abstract
Non-rigid image registration is a fundamental problem in medical imaging and a representative example of continuous transformation modeling in image processing. Diffeomorphic registration methods, such as Large Deformation Diffeomorphic Metric Mapping (LDDMM) and its PDE-constrained variants (PDE-LDDMM), provide mathematically grounded formulations with strong
[...] Read more.
Non-rigid image registration is a fundamental problem in medical imaging and a representative example of continuous transformation modeling in image processing. Diffeomorphic registration methods, such as Large Deformation Diffeomorphic Metric Mapping (LDDMM) and its PDE-constrained variants (PDE-LDDMM), provide mathematically grounded formulations with strong geometric guarantees for transformation quality. However, existing approaches face persistent trade-offs between numerical stability, accuracy, and computational efficiency. Recent work has explored implicit neural representations (INRs) and neural ordinary differential equations (NODEs) as flexible neural representations for modeling continuous transformations. Despite their increasing adoption, their practical behavior and limitations in diffeomorphic registration remain insufficiently understood. In this paper, we present a unified formulation of INR- and NODE-based registration methods within LDDMM and PDE-LDDMM, enabling a systematic and controlled comparison across architectures, sampling strategies, and numerical solvers. Our analysis reveals fundamental trade-offs between these approaches. In particular, we show that MLP-based INR formulations introduce significant computational overhead and rely on sampling strategies that can degrade smoothness and lead to the increased occurrence of non-diffeomorphic transformations at higher resolutions. Moreover, these approximations do not fully alleviate the computational cost, with some variants exceeding the costs of expensive classical optimization-based methods. In contrast, NODE-based formulations and downsampling strategies consistently provide transformations with more controlled Jacobian extrema while maintaining competitive computational performance. Among the evaluated methods, the original NODE-LDDMM and NODE-PDE-LDDMM formulations achieve the most favorable trade-offs between registration accuracy, geometric consistency, and computational efficiency. These findings provide clear insights into the design of neural representations for continuous transformation modeling, with practical implications for diffeomorphic registration and computational anatomy applications.
Full article
(This article belongs to the Section Medical Imaging)
►▼
Show Figures

Figure 1
Open AccessArticle
Clinical Determinants of Dose–Length Product in Pediatric Brain CT: Implications for Age-Specific Imaging Protocols
by
Kangmin Lee, Jina Shim and Youngjin Lee
J. Imaging 2026, 12(8), 383; https://doi.org/10.3390/jimaging12080383 - 14 Aug 2026
Abstract
Systematic optimization of pediatric brain computed tomography (CT) protocols remains challenging because clinical data are limited for children younger than 5 years. Sixty-nine non-contrast brain CT examinations in children aged 5 years or younger were retrospectively analyzed. Multiple linear regression was used to
[...] Read more.
Systematic optimization of pediatric brain computed tomography (CT) protocols remains challenging because clinical data are limited for children younger than 5 years. Sixty-nine non-contrast brain CT examinations in children aged 5 years or younger were retrospectively analyzed. Multiple linear regression was used to assess associations between dose–length product (DLP) and age group, sex, body mass index (BMI), and scanner group. In univariable analyses, DLP differed significantly by age group (p < 0.001) and scanner group (p < 0.001). In the multivariable model, only age group remained significantly associated with DLP; examinations in children aged 1–5 years showed an adjusted DLP increase of 158.71 mGy·cm compared with examinations in children younger than 1 year (p < 0.001). BMI and scanner group were not independently associated with DLP. Model stability was supported by residual normality and absence of multicollinearity. This study found that age group was the only measured variable significantly associated with DLP after adjustment, whereas BMI, sex, and scanner group were not significant. These findings highlight the importance of considering age when interpreting dose variation in pediatric brain CT.
Full article
(This article belongs to the Section Medical Imaging)
►▼
Show Figures

Figure 1
Open AccessArticle
GANCIU—Geospatial Analysis with Neural Classification and Image Understanding
by
Amedeo Ganciu, Giovannangela Ricci and Margherita Solci
J. Imaging 2026, 12(8), 382; https://doi.org/10.3390/jimaging12080382 - 14 Aug 2026
Abstract
Accurate and up-to-date knowledge of land use and land cover represents one of the central challenges in spatial planning and landscape sciences. In this context, the present work introduces GANCIU (Geospatial Analysis with Neural Classification and Image Understanding), an original hybrid pipeline for
[...] Read more.
Accurate and up-to-date knowledge of land use and land cover represents one of the central challenges in spatial planning and landscape sciences. In this context, the present work introduces GANCIU (Geospatial Analysis with Neural Classification and Image Understanding), an original hybrid pipeline for the automatic extraction of man-made infrastructure from high-resolution satellite imagery. The primary methodological contribution lies in the sequential integration of four technologically heterogeneous components: a per-pixel Random Forest classifier, a guided image modulation step, edge detection via the Mumford–Shah variational functional solved through the Ambrosio–Tortorelli approximation, and final object delineation via the Segment Anything Model (SAM). Each component does not operate independently but conditions and informs the next: The RF probability map guides the modulation, which in turn directs the sensitivity of the variational step exclusively towards regions of interest; the AT edges provide spatial prompts to SAM, for which its masks are finally filtered by the RF probability in an adaptive manner through a Gaussian Mixture Model. This progressive conditioning scheme constitutes the architectural core of GANCIU and distinguishes it from approaches that combine classification and segmentation in parallel or in purely sequential fashion with each stage conditioning the next but without any reverse correction between them. The Random Forest classifier was trained on 44 manually annotated scenes, geographically disjoint from the twelve independent scenes used for quantitative validation. This validation, based on an instance matching protocol (precision, recall, F1 score, and IoU), confirms the contribution of the full pipeline over a Random-Forest-only baseline: Pooled false positives fall by close to two orders of magnitude (from 8320 to 209), while true positives rise nearly twentyfold (from 5 to 95), with a mean IoU of 0.742 ± 0.060 on correctly matched objects. Notably, the entire pipeline—including SAM-based segmentation—runs end-to-end on a modest, GPU-free consumer laptop (four logical CPU cores, under 16 GB RAM), demonstrating that competitive infrastructure-extraction performance does not require specialised computing hardware.
Full article
(This article belongs to the Section Image and Video Processing)
►▼
Show Figures

Graphical abstract
Open AccessArticle
Domain Generalization of Histopathology Foundation Models in Multicenter, Multi-Scanner Cohorts: A Comparative Benchmark
by
Hafsa Akebli and Vincenzo Della Mea
J. Imaging 2026, 12(8), 381; https://doi.org/10.3390/jimaging12080381 - 13 Aug 2026
Abstract
Histopathology foundation models (FMs) have become widely used as patch-level feature extractors in computational pathology (CPath), where domain shift is a central challenge, yet their generalization ability across acquisition centers and scanning platforms remains insufficiently studied. In this work, we evaluate the domain
[...] Read more.
Histopathology foundation models (FMs) have become widely used as patch-level feature extractors in computational pathology (CPath), where domain shift is a central challenge, yet their generalization ability across acquisition centers and scanning platforms remains insufficiently studied. In this work, we evaluate the domain generalization of ten state-of-the-art FMs on two multi-source datasets with different supervision settings: SemiCOL, a colorectal cancer cohort of 499 whole-slide images (WSIs) for weakly labeled slide-level binary tumor classification, and BEETLE, a breast cancer cohort of 583 WSIs for patch-level four-class tissue classification. In an ablation-style setting, FMs are used as patch-level feature extractors, with patch embeddings mean-pooled into slide-level representations for SemiCOL, and a lightweight multi-layer perceptron trained for slide-level and patch-level classification on SemiCOL and BEETLE, respectively. To test FM domain generalization, we use three evaluation protocols: a Baseline source-mixed 5-fold cross-validation (CV) and two leave-source-out CV settings that assess cross-center and cross-scanner performance. On SemiCOL, all FMs achieve near-saturated performance, indicating stable performance under acquisition-source domain shift for slide-level Tumor vs. Benign classification. In contrast, BEETLE reveals clear generalization gaps, with scanner-induced domain shift more challenging than center-induced domain shift, and class-wise results showing that performance losses concentrate in epithelial discrimination. Overall, Virchow2 shows the strongest robustness across all evaluated protocols. These findings show that standard source-mixed CV can overestimate domain generalization across centers and scanners, and that FM choice matters in multi-source cohorts, especially for more challenging CPath tasks, where cross-domain failures are more visible.
Full article
(This article belongs to the Section Medical Imaging)
►▼
Show Figures

Figure 1
Open AccessArticle
Medical Textile Stain Detection Based on Chemically Enhanced Visualization and Deep Semantic Segmentation
by
Wenjie Min, Junfeng He, Zhenping Wan, Jinde Chen, Zhixiang Zou and Yuandong Mo
J. Imaging 2026, 12(8), 380; https://doi.org/10.3390/jimaging12080380 - 13 Aug 2026
Abstract
Pre-wash sorting of medical textiles is essential for hospital infection control, yet accurate stain detection remains challenging because visually apparent stains often have blurred boundaries, whereas dried urine stains lack distinguishable optical features. This study proposes a medical textile stain detection method integrating
[...] Read more.
Pre-wash sorting of medical textiles is essential for hospital infection control, yet accurate stain detection remains challenging because visually apparent stains often have blurred boundaries, whereas dried urine stains lack distinguishable optical features. This study proposes a medical textile stain detection method integrating chemically enhanced visualization with deep semantic segmentation. Dimethylaminocinnamaldehyde (DMACA) was used to convert latent urine stains into chemically developed stains with orange–red visual features. Based on the spatial color difference ΔE in the L*a*b* color space, 0.0183 mol/L was selected as the most suitable DMACA concentration among those tested. A dataset of 1974 images was constructed, including blood stains, chemically developed urine stains, medication stains, and uncontaminated textiles. A cascaded preprocessing strategy was applied to enhance stain boundaries and suppress textile texture noise, after which an Enhanced semantic segmentation model incorporating residual feature extraction, multiscale feature fusion, and transfer learning was used for pixel-level recognition. The IoU values for blood stains, chemically developed urine stains, and medication stains were 88.11%, 82.67%, and 89.62%, respectively. The average time required for image preprocessing and network inference was 15.39 ms per image. An input-level ablation comparison showed that DMACA-based color development increased the urine-stain IoU from 3.07% to 86.23%, demonstrating its substantial contribution to latent urine-stain detection. These results support the feasibility of integrating front-end chemical feature enhancement with back-end semantic segmentation for multiclass medical textile stain recognition under the current experimental conditions.
Full article
(This article belongs to the Section Image and Video Processing)
►▼
Show Figures

Figure 1
Open AccessArticle
Realistic Ultrasound Simulations of Healthy and Osteoarthritic Cartilage
by
Roby Weeteling, Yuexin Qi, Rob P. A. Janssen, Keita Ito, Corrinus C. van Donkelaar, Richard G. P. Lopata and Min Wu
J. Imaging 2026, 12(8), 379; https://doi.org/10.3390/jimaging12080379 - 12 Aug 2026
Abstract
Osteoarthritis (OA) causes irreversible cartilage damage, highlighting the need for early and sensitive assessment. Current imaging modalities are limited in detecting early-stage changes. Ultrasound (US) provides a non-invasive and accessible alternative, but its clinical adoption is limited by the lack of standardized protocols
[...] Read more.
Osteoarthritis (OA) causes irreversible cartilage damage, highlighting the need for early and sensitive assessment. Current imaging modalities are limited in detecting early-stage changes. Ultrasound (US) provides a non-invasive and accessible alternative, but its clinical adoption is limited by the lack of standardized protocols and reliable cartilage assessment. Simulations can be used to address these challenges by enabling system design, acquisition optimization and validation by providing ground truth when in vivo ground truth is unavailable. The aim of this study is to develop an in silico framework for realistic US imaging of healthy and OA cartilage by combining accurate acoustic wave modeling with a 2D microstructural cartilage phantom. The model was calibrated to healthy cartilage using first-order speckle statistics and extended to simulate degeneration through changes in structural and acoustic properties. As a proof-of-concept study, simulations were evaluated against limited ex vivo US data from healthy and OA cartilage and compared with literature data. The simulations reproduced key OA-related features and trends, including changes in reflection coefficient (R), integrated reflection coefficient (IRC), apparent integrated backscatter (AIB), and gray level distributions. These findings demonstrate the feasibility of using microstructure-based tissue phantoms to model healthy and OA cartilage. The framework provides a platform for systematic investigation of cartilage microstructure and US-derived features and may support future generation of synthetic datasets for data-driven and AI-based OA assessment.
Full article
(This article belongs to the Section Medical Imaging)
►▼
Show Figures

Figure 1
Highly Accessed Articles
Latest Books
E-Mail Alert
News
27 July 2026
Meet Us at the 29th International Conference on Medical Image Computing and Computer-Assisted Intervention, 27 September–1 October 2026, Strasbourg, France
Meet Us at the 29th International Conference on Medical Image Computing and Computer-Assisted Intervention, 27 September–1 October 2026, Strasbourg, France
5 August 2026
MDPI INSIGHTS: The CEO’s Letter #37 – Canada Summit, Sciforum Relaunch, 30 Years of Impactful Research & ISPRS 2026
MDPI INSIGHTS: The CEO’s Letter #37 – Canada Summit, Sciforum Relaunch, 30 Years of Impactful Research & ISPRS 2026
Topics
Topic in
Applied Sciences, Bioengineering, Diagnostics, J. Imaging, Signals
Signal Analysis and Biomedical Imaging for Precision Medicine
Topic Editors: Surbhi Bhatia Khan, Mo SaraeeDeadline: 31 August 2026
Topic in
BioMed, Cancers, Diagnostics, JCM, J. Imaging
Machine Learning and Deep Learning in Medical Imaging
Topic Editors: Rafał Obuchowicz, Michał Strzelecki, Adam Piórkowski, Karolina NurzynskaDeadline: 31 October 2026
Topic in
Applied Sciences, Electronics, J. Imaging, JMSE, Machines, Robotics, Sensors, Drones
Applications and Development of Underwater Robotics and Underwater Vision Technology, 2nd Edition
Topic Editors: Jingchun Zhou, Wenqi Ren, Qiuping Jiang, Yan-Tsung PengDeadline: 30 November 2026
Topic in
AI, Applied Sciences, Sensors, J. Imaging
Applied Computing and Machine Intelligence (ACMI): 2nd Edition
Topic Editors: Chuan-Ming Liu, Wei-Shinn KuDeadline: 31 December 2026
Conferences
Special Issues
Special Issue in
J. Imaging
Artificial Intelligence in Medical Imaging: Innovations and Challenges
Guest Editor: Zhao YaoDeadline: 30 August 2026
Special Issue in
J. Imaging
Diagnostic Imaging: From Basic Knowledge to Latest Advancements
Guest Editor: Paolo SpinnatoDeadline: 31 August 2026
Special Issue in
J. Imaging
Unified and Multimodal Segmentation: Foundations, Methods, and Applications
Guest Editors: Xiaoqi Zhao, Youwei Pang, Kailai ZhouDeadline: 31 August 2026
Special Issue in
J. Imaging
Trustworthy Multimodal Vision Models: Generalization, Robustness, and Explainability
Guest Editors: Xun Gong, Junzhou Chen, Chong MaDeadline: 31 August 2026



