Next Issue
Volume 12, September
Previous Issue
Volume 12, July
 
 

J. Imaging, Volume 12, Issue 8 (August 2026) – 64 articles

Cover Story (view full-size image): How do you teach a computer to reliably distinguish a roof from a field, anywhere, without a supercomputer? GANCIU answers by combining four seemingly incompatible elements into a single self-correcting pipeline: a Random Forest classifier, a variational edge detector inspired by physics, and Meta’s “Segment Anything” model, each conditioning the next rather than operating independently. On twelve independent satellite images of Sardinia, this chain eliminates nearly all false alarms, reducing them by almost a hundredfold, while detecting approximately twenty times more actual infrastructure objects than classification alone, with an average intersection-over-union accuracy above 74%. The entire pipeline, segmentation included, runs end-to-end on a modest laptop without a GPU, demonstrating that reliable, high-resolution land-cover mapping does not require specialized computing hardware. View this paper
  • Issues are regarded as officially published after their release is announced to the table of contents alert mailing list.
  • You may sign up for e-mail alerts to receive table of contents of newly released issues.
  • PDF is the official format for papers published in both, html and pdf forms. To view the papers in pdf format, click on the "PDF Full-text" link, and use the free Adobe Reader to open them.
Order results
Result details
Section
Select all
Export citation of selected articles as:
19 pages, 3894 KB  
Article
Attention-Enhanced Multi-Scale Feature-Wise Linear Modulation for Fine-Grained Poisonous Mushroom Image Recognition
by Yuan He, Haikun Lv, Chenyang Lu, Dengqi Yang, Xiaowei Li and Lina Zhang
J. Imaging 2026, 12(8), 398; https://doi.org/10.3390/jimaging12080398 - 21 Aug 2026
Viewed by 218
Abstract
Fine-grained poisonous mushroom recognition in natural scenes is challenging because of complex backgrounds, subtle morphological differences, and the limited interpretability of model decisions. To address these challenges, this paper proposes Att-FiLM, an attention-enhanced multi-scale Feature-Wise Linear Modulation network for poisonous mushroom image recognition. [...] Read more.
Fine-grained poisonous mushroom recognition in natural scenes is challenging because of complex backgrounds, subtle morphological differences, and the limited interpretability of model decisions. To address these challenges, this paper proposes Att-FiLM, an attention-enhanced multi-scale Feature-Wise Linear Modulation network for poisonous mushroom image recognition. The model adopts an asymmetric dual-backbone architecture in which a frozen ConvNeXt-Base branch provides global semantic priors, while a trainable EfficientNet-B0 branch learns local discriminative features. Rather than directly concatenating heterogeneous features, Att-FiLM generates scale and shift parameters from semantic features and performs channel-wise modulation on multi-scale EfficientNet features at Stage 2 and Stage 4. This mechanism enables global semantic information to guide local feature learning while reducing feature redundancy and semantic inconsistency. Experimental results show that Att-FiLM achieves an Accuracy of 95.58% and an F1-score of 0.9455 on the poisonous/edible binary classification task. On the 190-class species-level classification task, it achieves a Top-1 Accuracy of 93.63% and a Macro-F1 of 0.9347. Interpretability analysis further shows that decision-relevant responses are frequently associated with morphologically relevant regions, including gills, annuli, volvae, and cap textures. These results indicate that Att-FiLM provides effective recognition performance together with interpretable decision evidence for mushroom recognition in complex natural scenes. Full article
(This article belongs to the Section Image and Video Processing)
Show Figures

Figure 1

22 pages, 6472 KB  
Article
Landmark Recognition Beyond Curated Benchmarks: Cross-Domain Evaluation of a Multi-Threshold Selective YOLO11 Ensemble on User-Generated Imagery, with a Zero-Shot Multimodal LLM Baseline
by Ulugbek Hudayberdiev, Abdimumin Alikulov, Adkham Israilov, Muhiddin Xidirov and Javokhir Musaev
J. Imaging 2026, 12(8), 397; https://doi.org/10.3390/jimaging12080397 - 21 Aug 2026
Viewed by 248
Abstract
Landmark recognition for smart tourism is usually validated on curated benchmark images. In deployment, however, the classifier must handle user-generated photographs whose viewpoint, lighting, resolution, occlusion, and compression differ sharply from curated data. This paper evaluates a previously published multi-threshold enhancement and selective [...] Read more.
Landmark recognition for smart tourism is usually validated on curated benchmark images. In deployment, however, the classifier must handle user-generated photographs whose viewpoint, lighting, resolution, occlusion, and compression differ sharply from curated data. This paper evaluates a previously published multi-threshold enhancement and selective YOLO11n-cls ensemble under this shift, and provides a preliminary zero-shot comparison of three general-purpose multimodal large language models (MLLMs) on the same task. To measure the shift, we build Samarkand v2-SNS, a 300-image out-of-distribution test set of social-media photographs of 12 Samarkand landmarks, disjoint from the training and validation data. Under the shift, four supervised baselines fall by 12.73–22.08 percentage points to 73–80% accuracy, and their in-distribution ranking does not hold. The selective ensemble degrades least (99.24% to 93.00%, −6.24 points) and outperforms the strongest baseline by 13 points. A capacity-matched ablation shows that most of this robustness comes from enhancement diversity, not from generic ensembling. In a preliminary comparison, zero-shot MLLMs (GPT-5, Claude Sonnet 4.5, Gemini 2.5) reach only 24.81–54.26%, far below deployment needs. The results argue for reporting out-of-distribution accuracy alongside curated benchmarks, and for hybrid systems that pair compact specialised recognisers with MLLM-based interpretation. Full article
(This article belongs to the Section Computer Vision and Pattern Recognition)
Show Figures

Figure 1

14 pages, 65828 KB  
Article
Identity Document Presentation Attack Detection in Visible Light with Illumination-Controlled Scanner
by Lada Tolstenko, Alexey Popkov, Irina Kunina, Dmitry Polevoy and Sergey Usilin
J. Imaging 2026, 12(8), 396; https://doi.org/10.3390/jimaging12080396 - 20 Aug 2026
Viewed by 321
Abstract
A reliable sign of the absence of a document presentation attack, when a print copy is presented instead of the original document, is the presence of such security features as OVDs (Optical Variable Devices), for example, holograms. To check the presence of holograms, [...] Read more.
A reliable sign of the absence of a document presentation attack, when a print copy is presented instead of the original document, is the presence of such security features as OVDs (Optical Variable Devices), for example, holograms. To check the presence of holograms, it is sufficient to use the visible light and a series of document images captured with a varying angle of incidence and reflection of light. It can be achieved either by changing the position of the document or by changing the position of the illumination source. This work proposes a method for detecting holograms on identity documents using a scanner with controlled illumination. The method is based on obtaining a series of document images in various illumination modes and identifying features characteristic of holograms. To test the method, a dataset MIDV-Holo-Scan was collected by scanning physical documents used in the creation of the open dataset MIDV-Holo. It includes both documents with holograms, accepted in this work as originals, and documents without holograms, simulating an attack on document presentation. The proposed method for detecting attacks on document presentation achieves a quality of Accuracy = 100%, which surpasses the quality of the baseline method published with the MIDV-Holo dataset. Full article
Show Figures

Figure 1

13 pages, 4740 KB  
Article
Imaging-Based Parallel Seismic Test for In-Service Bridge Pile Foundations
by Zhichao Luo, Weibin Luo, Peng Wang, Yuan Gu and Peimin Zhu
J. Imaging 2026, 12(8), 395; https://doi.org/10.3390/jimaging12080395 - 20 Aug 2026
Viewed by 231
Abstract
The conventional parallel seismic test (PST) is a widely used and highly reliable method for determining lengths of in-service bridge pile foundations. However, it relies solely on the manual interpretation of first-arrival in seismic records and often fails to detect, characterize, or geometrically [...] Read more.
The conventional parallel seismic test (PST) is a widely used and highly reliable method for determining lengths of in-service bridge pile foundations. However, it relies solely on the manual interpretation of first-arrival in seismic records and often fails to detect, characterize, or geometrically define internal defects within the piles. To overcome these limitations and enable the intuitive identification of internal defects and damage within piles, this study introduces an elastic reverse time migration (ERTM) imaging algorithm based on the spectral element method, achieving high-resolution imaging of existing bridge pile foundations and their defects and damage. To suppress crosstalk between P- and S-waves during the ERTM process, a wavefield decoupling method is employed to separate the elastic wavefield into P- and S-wave components for independent imaging. Two-dimensional numerical testing on a bridge pile model with a necking defect demonstrates that ERTM can effectively image both the pile geometry and its defects, significantly improving the capability of defect detection and characterization. This approach provides more intuitive visualization for assessing the structural integrity of in-service bridge pile foundations. Full article
(This article belongs to the Section Image and Video Processing)
Show Figures

Figure 1

29 pages, 17754 KB  
Article
Structure-Prior-Guided Multi-Stage Cross-Modal Collaborative Network for RGB-D Semantic Segmentation
by Yifan Yu, Zhiwei Zhong, Fan Min and Song Deng
J. Imaging 2026, 12(8), 394; https://doi.org/10.3390/jimaging12080394 - 20 Aug 2026
Viewed by 257
Abstract
Red–green–blue and depth (RGB-D) semantic segmentation combines appearance cues from RGB images with geometric information from depth maps, but sensor noise, missing measurements, and boundary-inconsistent depth responses can introduce conflicting evidence during cross-modal fusion. We propose the Structure-Prior-Guided Network (SPGNet), a dual-branch, multi-stage [...] Read more.
Red–green–blue and depth (RGB-D) semantic segmentation combines appearance cues from RGB images with geometric information from depth maps, but sensor noise, missing measurements, and boundary-inconsistent depth responses can introduce conflicting evidence during cross-modal fusion. We propose the Structure-Prior-Guided Network (SPGNet), a dual-branch, multi-stage framework that follows a correction-before-fusion strategy. At each feature scale, SPGNet estimates a learned structure prior from cross-modal agreement and discrepancy. The Cross-Modal Correction Module (CCM) uses this prior to regulate bidirectional information transfer, suppressing unreliable responses while retaining complementary cues. The Dual-branch Enhancement Fusion Module (DEF) then enhances the corrected RGB and depth features and integrates them through shared-representation-guided interaction, after which a lightweight multi-scale decoder produces the segmentation output. Under a unified training and evaluation protocol, SPGNet achieved three-run mean Intersection over Union (mIoU) scores of 50.845% on NYU Depth V2 and 48.457% on SUN RGB-D. Compared with the best reproduced baseline on each dataset, SPGNet improved mean mIoU by 2.111 and 0.899 percentage points, respectively. These results suggest that separating reliability-oriented correction from multimodal fusion can limit the propagation of unreliable cross-modal responses and improve indoor RGB-D semantic segmentation performance. Full article
(This article belongs to the Section AI in Imaging)
Show Figures

Figure 1

26 pages, 6887 KB  
Article
Turning Immersive Viewers into Analytical Workspaces: ASCRIBE-XR and Agent-Driven Scientific Visualization
by Ronald Pandolfi, Luke Weidner, James Sethian, Jeffrey Donatelli and Daniela Ushizima
J. Imaging 2026, 12(8), 393; https://doi.org/10.3390/jimaging12080393 - 20 Aug 2026
Viewed by 226
Abstract
Scientific visualization is changing from passive observation to active, AI-assisted collaboration. While Extended Reality (XR) has proven valuable for comprehending dense 3D arrays, traditional VR applications are typically deployed in rigid, single-purpose, and monolithic architectures. In this paper, we present the evolution of [...] Read more.
Scientific visualization is changing from passive observation to active, AI-assisted collaboration. While Extended Reality (XR) has proven valuable for comprehending dense 3D arrays, traditional VR applications are typically deployed in rigid, single-purpose, and monolithic architectures. In this paper, we present the evolution of ASCRIBE-XR: a virtual reality platform backed by remote computation that has been re-engineered into a dynamic, service-oriented ecosystem. We introduce three core innovations that make immersive data analysis easier, faster, and more flexible when using multimodal scientific imaging. First, a lightweight Python REST interface decouples XR logic from the rendering engine, enabling real-time, programmable scene customization and on-demand data generation. Second, we present a Specimen Catalog architecture that lets the platform pivot between radically different disciplines, ranging from archaeological heterogeneous concrete and fuel-cell membranes to the root system of a bioenergy grass, by describing each dataset through portable metadata rather than hard-coded application logic. Finally, we introduce a prompt-driven layer powered by the Claude Agent SDK, allowing researchers to generate, segment, and manipulate volumetric and mesh data through natural language dialogue within the virtual space. For example, applying foundation models such as the Segment Anything Model (SAM) to perform zero-shot segmentation on demand. By bridging human intent with remote computation, ASCRIBE-XR relaxes the constraints of conventional visualization tools, offering a highly adaptable, conversational platform for scientific discovery with human auditing. Full article
(This article belongs to the Section AI in Imaging)
Show Figures

Figure 1

14 pages, 7396 KB  
Article
A Stability Atlas for IBSI Radiomics Features Using Synthetic Digital Phantoms, with Proof-of-Concept Physics-Based Normalisation
by Shuji Yamamoto
J. Imaging 2026, 12(8), 392; https://doi.org/10.3390/jimaging12080392 - 20 Aug 2026
Viewed by 263
Abstract
Radiomics features are strongly sensitive to image acquisition, and separating that sensitivity from biological signal usually requires repeated patient scans that cannot be shared. We present an open, fully synthetic framework (radiomics-phantom) that maps and, as a proof of concept, corrects radiomics feature [...] Read more.
Radiomics features are strongly sensitive to image acquisition, and separating that sensitivity from biological signal usually requires repeated patient scans that cannot be shared. We present an open, fully synthetic framework (radiomics-phantom) that maps and, as a proof of concept, corrects radiomics feature instability without any patient data. Deterministic three-dimensional texture phantoms are generated as anisotropic Gaussian random fields with known ground truth and an optional embedded lesion. An independently implemented feature core aligned with the Image Biomarker Standardization Initiative (IBSI) covers all eleven IBSI-1 feature families and matched all 482 published digital-phantom benchmark values within the applicable tolerances. An image-domain acquisition simulator applies point-spread blur, slice-profile averaging, dose-scaled correlated noise, resampling, and quantisation. Per-feature reproducibility across a sweep of fifteen textures (varying correlation length, anisotropy, and intensity scale) by nine acquisition conditions, with five independent noise realisations per stochastic setting, is summarised by the absolute-agreement intraclass correlation ICC(2,1), with a realisation-aware percentile-bootstrap 95% confidence interval for every estimate; constant features are excluded from estimation. Values span nearly the full range (median 0.13, 95% CI 0.03–0.19), and a hierarchical variance decomposition attributes a median 77% of per-feature variance to the acquisition condition and under 1% to stochastic realisation; the values are interpreted as exploratory rankings within this acquisition envelope. As a proof of concept, intensity variance and grey-level co-occurrence contrast under additive Gaussian noise were normalised using calibrated, invertible response models, returning them to their noiseless values on held-out data (median error below 4% across five textures and repeated noise realisations, and about 11% when the noise level is estimated from the degraded image itself), while features the models cannot describe are refused rather than corrected. All code and a 716-test suite are released openly and archived on Zenodo. The result is a reproducible, patient-data-free testbed for radiomics feature stability. Full article
(This article belongs to the Section Medical Imaging)
Show Figures

Figure 1

25 pages, 5700 KB  
Article
Research on Medical Image Super-Resolution Reconstruction Algorithm Based on Dilated Convolution and Multi-Module Fusion
by Zhuye Xu and Yucong Guo
J. Imaging 2026, 12(8), 391; https://doi.org/10.3390/jimaging12080391 - 19 Aug 2026
Viewed by 244
Abstract
Medical image resolution plays a crucial role in early disease detection and fine-structure observation. Super-resolution reconstruction technology can restore low-resolution images to high-resolution versions, thereby assisting physicians in making accurate diagnoses. To address challenges in medical image super-resolution reconstruction, including insufficient global information [...] Read more.
Medical image resolution plays a crucial role in early disease detection and fine-structure observation. Super-resolution reconstruction technology can restore low-resolution images to high-resolution versions, thereby assisting physicians in making accurate diagnoses. To address challenges in medical image super-resolution reconstruction, including insufficient global information acquisition, excessive network complexity, and suboptimal loss function adaptation for medical imaging data, this paper proposes an image super-resolution reconstruction algorithm named IDCASR-MMF based on improved dilated convolution and multi-module fusion. First, multi-dilation-rate dilated convolution is introduced to expand the receptive field and integrated with a spatial attention mechanism to dynamically calibrate high-frequency features after feature extraction. Subsequently, the Squeeze-and-Excitation module is fused with dilated convolution as a channel attention mechanism to streamline the network architecture. Finally, a weighted fusion strategy combining adversarial loss and MSE loss is adopted, where the dynamic adjustment of weighting coefficients balances pixel-level structural accuracy and high-frequency detail authenticity, achieving synergistic optimization of objective precision and subjective quality for medical images. To validate the effectiveness of the proposed algorithm, IDCASR-MMF is compared with 11 state-of-the-art methods across five datasets (Set5, Set14, BSD100, Urban100, and Bone FD). Experimental results demonstrate that the proposed algorithm achieves superior PSNR and SSIM values on multiple datasets, confirming that IDCASR-MMF can effectively reconstruct high-resolution medical images from low-resolution inputs. Full article
(This article belongs to the Section Medical Imaging)
Show Figures

Figure 1

9 pages, 1186 KB  
Communication
Four-Dimensional Cine Cinematic Rendering of Structural Heart and Mechanical Circulatory Support Devices: An Illustrative Technical Experience
by Amy Avakian and Muhammad Umair
J. Imaging 2026, 12(8), 390; https://doi.org/10.3390/jimaging12080390 - 19 Aug 2026
Viewed by 243
Abstract
Patients with implanted cardiac devices are a rapidly growing imaging population, and electrocardiogram-gated cardiac computed tomography (CT) is increasingly used to characterize device geometry, multi-device relationships, and dynamic behavior across the cardiac cycle. Cinematic rendering (CR) is a photorealistic three-dimensional (3D) visualization technique [...] Read more.
Patients with implanted cardiac devices are a rapidly growing imaging population, and electrocardiogram-gated cardiac computed tomography (CT) is increasingly used to characterize device geometry, multi-device relationships, and dynamic behavior across the cardiac cycle. Cinematic rendering (CR) is a photorealistic three-dimensional (3D) visualization technique for cardiac CT whose established contribution in this population is communicative: it conveys 3D device geometry and material distinctions within a single rendered volume. We describe a demonstrative case series extending CR across the cardiac cycle—time-resolved “4D cine” CR—to depict dynamic device behavior and time-resolved multi-device interaction in a single volume; this is an illustrative technical experience rather than a systematic evaluation of diagnostic performance. Illustrative examples include an EVOQUE transcatheter tricuspid valve rendered together with concurrent surgical mitral and transcatheter aortic valves, a left atrial appendage occlusion device, a normally positioned Impella catheter, and a HeartMate 3 left ventricular assist device (LVAD). Across cases, 4D cine CR feasibility scaled inversely with metallic burden—the aggregate volume and radiodensity of metallic device components within the scan field—with renderings informative for low-metal nitinol and catheter devices but substantially degraded by streak artifact in high-metal LVAD housings. This relationship was observed qualitatively in a small selected series and is offered as an initial observation rather than an established characteristic of the technique. We discuss current limitations and emerging directions such as photon-counting detector CT, metal artifact reduction, and artificial-intelligence-assisted post-processing that may extend 4D cine CR in this population. Full article
(This article belongs to the Section Medical Imaging)
Show Figures

Figure 1

28 pages, 102280 KB  
Article
PRMEFNet: A Real-Time Unsupervised Multi-Exposure Fusion Network Driven by Prior Knowledge
by Junwei Qi, Hangdong Wang, Xu Xiao and Jingpeng Gao
J. Imaging 2026, 12(8), 389; https://doi.org/10.3390/jimaging12080389 - 19 Aug 2026
Viewed by 167
Abstract
Due to the limited dynamic range of imaging sensors, most cameras can only capture low-dynamic-range (LDR) images. Multi-exposure fusion (MEF) is an effective technique for generating high-dynamic-range (HDR) images. However, to simultaneously preserve texture details and global exposure, most existing methods primarily rely [...] Read more.
Due to the limited dynamic range of imaging sensors, most cameras can only capture low-dynamic-range (LDR) images. Multi-exposure fusion (MEF) is an effective technique for generating high-dynamic-range (HDR) images. However, to simultaneously preserve texture details and global exposure, most existing methods primarily rely on more complex models to improve performance, resulting in higher computational costs and longer processing times. To address this issue, we propose a real-time unsupervised MEF network driven by prior knowledge. To this end, a hierarchical feature extraction module is designed that utilizes filtering operations to decompose the source images into base layers and detail layers. Features are extracted from each layer separately to reduce the difficulty of extracting effective features. Then, the receptive field of feature maps is expanded by dilated convolutions, and a window-based self-attention mechanism is applied to perform context modeling, achieving effective contextual modeling with low computational cost. Subsequently, texture features and global features are extracted separately to enable the model to maintain both local texture clarity and global smoothness. In addition, a one-dimensional lookup table is utilized to accelerate the inference process. Comprehensive experiments are conducted to verify the effectiveness of the proposed method. The subjective evaluation results demonstrate that the fused images exhibit superior visual quality, while objective experiments further quantify its superior performance, demonstrating that the proposed method effectively reduces computation time. Full article
(This article belongs to the Topic Computational Imaging)
Show Figures

Figure 1

11 pages, 2088 KB  
Article
Ordinal Deep Learning for Lumbar Foraminal Stenosis Grading on Sagittal MRI
by Rohan A. Phadke, Samer G. Salman, Zane G. Salman, Akhil Marupudi, Kirtan Patel, Joshua Ong, Alireza Tavakkoli, Sainyam Galhotra, Ajay Tripuraneni, James Rizkalla and Nathan J. Lee
J. Imaging 2026, 12(8), 388; https://doi.org/10.3390/jimaging12080388 - 19 Aug 2026
Cited by 1 | Viewed by 302
Abstract
Lumbar foraminal stenosis grading contributes to surgical-level selection, but automated four-grade classification remains challenging. Published pipelines for this dataset reach approximately 65% four-class accuracy, and deep classifiers offer no anatomical rationale. We investigated whether interpretable, millimeter-scale morphometry from segmentation masks improves grading beyond [...] Read more.
Lumbar foraminal stenosis grading contributes to surgical-level selection, but automated four-grade classification remains challenging. Published pipelines for this dataset reach approximately 65% four-class accuracy, and deep classifiers offer no anatomical rationale. We investigated whether interpretable, millimeter-scale morphometry from segmentation masks improves grading beyond a deep image model. We analyzed the LSS-MRI-AISSLab sagittal T2-weighted dataset (469 patients, 2979 expert-graded foramina spanning L1-L2 through L5-S1 bilaterally on a four-grade scale). Ten morphometric descriptors were computed from mid-sagittal polygon segmentations and scaled to millimeters using each patient’s recorded pixel spacing. A dual-branch network combined a fine-tuned ResNet-18 embedding of each foraminal region of interest with the morphometric vector through an ordinal regression head. Foraminal regions were supplied from expert bounding-box annotations; automated localization within the full sagittal examination was not evaluated. Four configurations (nominal softmax, appearance-only, anatomy-only, and fusion) were compared on a locked patient-level test set of 94 patients after five-fold cross-validation, with quadratic weighted kappa (QWK) as the primary endpoint and patient-clustered bootstrap inference. Feature-grade correlations were reported pooled and adjusted for lumbar level. Fusion achieved QWK 0.813 (95% confidence interval [CI] 0.769–0.847) and 75.3% four-class accuracy. Appearance-only was statistically indistinguishable (QWK 0.806; delta QWK +0.006, 95% CI −0.023 to 0.036, p = 0.68), whereas anatomy-only reached 0.444, and a level-and-side-only reference reached 0.314. Boundary discrimination was strong (area under the curve 0.92–0.99), 98.3% of predictions fell within one grade, and performance was consistent across scanner vendors. Morphometric associations were confounded by level: the apparent spondylolisthesis effect (rho −0.349) disappeared after adjustment (rho −0.000), while disc height, null when pooled (rho +0.021), emerged as a genuine within-level effect (rho −0.108). A fine-tuned ordinal image classifier achieved strong agreement for four-grade lumbar foraminal stenosis classification. The evaluated segmentation-derived morphometric features did not improve performance beyond imaging alone, and several apparent anatomic associations reflected confounding by lumbar level. External and prospective validation in complete clinical MRI workflows are needed before implementation. Full article
(This article belongs to the Special Issue Medical Computer Vision: Innovations and Clinical Impact)
Show Figures

Figure 1

9 pages, 958 KB  
Communication
Quantitative MRI Susceptibility: Mapping of Transient Ischemic Attack—A Preliminary Study
by Philipp Gruber, Michael Diepers, Markus Gschwind, Paul G. Unschuld, Luca Remonda, Franca Wagner, Pasquale Mordasini and Jatta Berberat
J. Imaging 2026, 12(8), 387; https://doi.org/10.3390/jimaging12080387 - 17 Aug 2026
Viewed by 233
Abstract
A transient ischemic attack (TIA) is a transient episode of neurological dysfunction without evidence of acute infarct demarcation on neuroimaging. The radiographic features of TIAs are uncertain, as since there is currently no perfect imaging method with which to differentiate a stroke from [...] Read more.
A transient ischemic attack (TIA) is a transient episode of neurological dysfunction without evidence of acute infarct demarcation on neuroimaging. The radiographic features of TIAs are uncertain, as since there is currently no perfect imaging method with which to differentiate a stroke from a TIA. We aimed to evaluate whether regional abnormalities are associated with TIA by identifying tissue with increased susceptibility (χ) values using quantitative susceptibility mapping (QSM). A total of 39 adults with clinical TIA symptoms (66 ± 12 years) and 39 age- and sex matched controls (65 ± 12 years) underwent MRI. Based on the QSM-maps, local susceptibility changes were tested for statistically significant differences between patients with TIA and healthy controls. The susceptibility maps of TIA patients showed that magnetic susceptibility differed from that of healthy volunteers, with the strongest effects observable in the right lingual gyrus (p < 0.001) and bilaterally in the caudal anterior cingulate (cACC, p < 0.01). TIA is often difficult to diagnose at the time of presentation in the emergency department, and QSM could show an association between regional QSM abnormalities and TIA. Full article
Show Figures

Figure 1

25 pages, 17984 KB  
Article
Information Retention and Feature Screening Synergistic Network for Aviation Ground Safety and Protective Devices
by Enming Wu, Mingxuan Wang, Runxia Guo, Jiusheng Chen, Jiaren Li, Fuyu Sun and Liyuan Ye
J. Imaging 2026, 12(8), 386; https://doi.org/10.3390/jimaging12080386 - 17 Aug 2026
Viewed by 275
Abstract
Aviation ground safety and protective devices are critical for flight safety; however, their unintentional retention on aircraft after maintenance remains a persistent risk. Existing deep learning-based approaches for aviation safety have predominantly followed a reactive paradigm, detecting FOD on runways or inspecting the [...] Read more.
Aviation ground safety and protective devices are critical for flight safety; however, their unintentional retention on aircraft after maintenance remains a persistent risk. Existing deep learning-based approaches for aviation safety have predominantly followed a reactive paradigm, detecting FOD on runways or inspecting the aircraft for inadvertently retained tools post-maintenance. In contrast, this paper advocates a proactive philosophy: using a neural network to recognize and inventory all ground safety and protective devices immediately after maintenance closure, thereby preventing retention incidents at their source. However, realizing this proactive verification is technically challenging—object detection for these devices often suffers from loss of fine-grained detail due to downsampling and inherently sparse semantic information of the targets. To this end, we propose an Information Retention and Feature Screening Synergistic Network (RS-Net) grounded in information bottleneck theory. The network comprises a main branch that enhances discriminative features through attention-guided screening, and an auxiliary branch, used only during training, that preserves fine-grained spatial details via information-retentive convolutions. A Dual-State Region Refinement Module (DRM) provides configurable support for both branches, decoupling the conflicting objectives of background compression and detail preservation. Experiments on a self-constructed dataset collected from real airline maintenance operations demonstrate that RS-Net substantially outperforms the strong YOLOv9 baseline, achieving gains of 4.531% in F1-score, 2.533% in mAP0.5, and 1.429% in mAP0.5:0.95. Cross-dataset experiments further validate its strong generalization capability. Full article
Show Figures

Figure 1

25 pages, 3965 KB  
Article
VLM-Assisted Routing and HU-Traceable DICOM Adaptation for Lumbar CT HU Measurement: A Deployment-Oriented Pilot Technical Evaluation
by Zhe-Yu Ye, Jun-Mu Peng and Tamotsu Kamishima
J. Imaging 2026, 12(8), 385; https://doi.org/10.3390/jimaging12080385 - 14 Aug 2026
Viewed by 223
Abstract
Automated lumbar computed tomography (CT) Hounsfield unit (HU) measurement can support opportunistic osteoporosis screening, but heterogeneous Digital Imaging and Communications in Medicine (DICOM) inputs often disrupt automated workflows before measurement. We evaluated a local Qwen2.5-VL-7B vision-language model (VLM)-assisted front end for suitability routing, [...] Read more.
Automated lumbar computed tomography (CT) Hounsfield unit (HU) measurement can support opportunistic osteoporosis screening, but heterogeneous Digital Imaging and Communications in Medicine (DICOM) inputs often disrupt automated workflows before measurement. We evaluated a local Qwen2.5-VL-7B vision-language model (VLM)-assisted front end for suitability routing, input planning, and HU-traceable DICOM-derived field-of-view/orientation adaptation upstream of an unchanged single-slice lumbar CT HU workflow. The VLM was restricted to routing and planning. In a 20-case FUJIFILM pilot cohort, observed HU output availability was higher with the front-end-assisted route than with direct processing (12/20 versus 5/20), while paired automatic-success cases showed excellent agreement with post hoc manual ImageJ (version 1.54g) measurements (ICC(A,1) = 0.997; MAE = 4.64 HU). Public DICOM stress testing further demonstrated improved workflow robustness across heterogeneous datasets. These findings support the use of a deployment-oriented front-end strategy to improve the auditability and robustness of automated lumbar CT HU measurement while preserving an unchanged downstream measurement workflow. Full article
(This article belongs to the Section Medical Imaging)
Show Figures

Figure 1

43 pages, 31425 KB  
Article
Understanding Trade-Offs in Continuous Neural Representations for Diffeomorphic Image Registration: A Comparative Study of Implicit Neural Representations and Neural Ordinary Differential Equations
by Salvador Rodriguez-Sanz, Carlos Paesa-Lia and Monica Hernandez
J. Imaging 2026, 12(8), 384; https://doi.org/10.3390/jimaging12080384 - 14 Aug 2026
Viewed by 259
Abstract
Non-rigid image registration is a fundamental problem in medical imaging and a representative example of continuous transformation modeling in image processing. Diffeomorphic registration methods, such as Large Deformation Diffeomorphic Metric Mapping (LDDMM) and its PDE-constrained variants (PDE-LDDMM), provide mathematically grounded formulations with strong [...] Read more.
Non-rigid image registration is a fundamental problem in medical imaging and a representative example of continuous transformation modeling in image processing. Diffeomorphic registration methods, such as Large Deformation Diffeomorphic Metric Mapping (LDDMM) and its PDE-constrained variants (PDE-LDDMM), provide mathematically grounded formulations with strong geometric guarantees for transformation quality. However, existing approaches face persistent trade-offs between numerical stability, accuracy, and computational efficiency. Recent work has explored implicit neural representations (INRs) and neural ordinary differential equations (NODEs) as flexible neural representations for modeling continuous transformations. Despite their increasing adoption, their practical behavior and limitations in diffeomorphic registration remain insufficiently understood. In this paper, we present a unified formulation of INR- and NODE-based registration methods within LDDMM and PDE-LDDMM, enabling a systematic and controlled comparison across architectures, sampling strategies, and numerical solvers. Our analysis reveals fundamental trade-offs between these approaches. In particular, we show that MLP-based INR formulations introduce significant computational overhead and rely on sampling strategies that can degrade smoothness and lead to the increased occurrence of non-diffeomorphic transformations at higher resolutions. Moreover, these approximations do not fully alleviate the computational cost, with some variants exceeding the costs of expensive classical optimization-based methods. In contrast, NODE-based formulations and downsampling strategies consistently provide transformations with more controlled Jacobian extrema while maintaining competitive computational performance. Among the evaluated methods, the original NODE-LDDMM and NODE-PDE-LDDMM formulations achieve the most favorable trade-offs between registration accuracy, geometric consistency, and computational efficiency. These findings provide clear insights into the design of neural representations for continuous transformation modeling, with practical implications for diffeomorphic registration and computational anatomy applications. Full article
(This article belongs to the Section Medical Imaging)
Show Figures

Figure 1

12 pages, 4271 KB  
Article
Clinical Determinants of Dose–Length Product in Pediatric Brain CT: Implications for Age-Specific Imaging Protocols
by Kangmin Lee, Jina Shim and Youngjin Lee
J. Imaging 2026, 12(8), 383; https://doi.org/10.3390/jimaging12080383 - 14 Aug 2026
Viewed by 197
Abstract
Systematic optimization of pediatric brain computed tomography (CT) protocols remains challenging because clinical data are limited for children younger than 5 years. Sixty-nine non-contrast brain CT examinations in children aged 5 years or younger were retrospectively analyzed. Multiple linear regression was used to [...] Read more.
Systematic optimization of pediatric brain computed tomography (CT) protocols remains challenging because clinical data are limited for children younger than 5 years. Sixty-nine non-contrast brain CT examinations in children aged 5 years or younger were retrospectively analyzed. Multiple linear regression was used to assess associations between dose–length product (DLP) and age group, sex, body mass index (BMI), and scanner group. In univariable analyses, DLP differed significantly by age group (p < 0.001) and scanner group (p < 0.001). In the multivariable model, only age group remained significantly associated with DLP; examinations in children aged 1–5 years showed an adjusted DLP increase of 158.71 mGy·cm compared with examinations in children younger than 1 year (p < 0.001). BMI and scanner group were not independently associated with DLP. Model stability was supported by residual normality and absence of multicollinearity. This study found that age group was the only measured variable significantly associated with DLP after adjustment, whereas BMI, sex, and scanner group were not significant. These findings highlight the importance of considering age when interpreting dose variation in pediatric brain CT. Full article
(This article belongs to the Section Medical Imaging)
Show Figures

Figure 1

28 pages, 24738 KB  
Article
GANCIU—Geospatial Analysis with Neural Classification and Image Understanding
by Amedeo Ganciu, Giovannangela Ricci and Margherita Solci
J. Imaging 2026, 12(8), 382; https://doi.org/10.3390/jimaging12080382 - 14 Aug 2026
Viewed by 553
Abstract
Accurate and up-to-date knowledge of land use and land cover represents one of the central challenges in spatial planning and landscape sciences. In this context, the present work introduces GANCIU (Geospatial Analysis with Neural Classification and Image Understanding), an original hybrid pipeline for [...] Read more.
Accurate and up-to-date knowledge of land use and land cover represents one of the central challenges in spatial planning and landscape sciences. In this context, the present work introduces GANCIU (Geospatial Analysis with Neural Classification and Image Understanding), an original hybrid pipeline for the automatic extraction of man-made infrastructure from high-resolution satellite imagery. The primary methodological contribution lies in the sequential integration of four technologically heterogeneous components: a per-pixel Random Forest classifier, a guided image modulation step, edge detection via the Mumford–Shah variational functional solved through the Ambrosio–Tortorelli approximation, and final object delineation via the Segment Anything Model (SAM). Each component does not operate independently but conditions and informs the next: The RF probability map guides the modulation, which in turn directs the sensitivity of the variational step exclusively towards regions of interest; the AT edges provide spatial prompts to SAM, for which its masks are finally filtered by the RF probability in an adaptive manner through a Gaussian Mixture Model. This progressive conditioning scheme constitutes the architectural core of GANCIU and distinguishes it from approaches that combine classification and segmentation in parallel or in purely sequential fashion with each stage conditioning the next but without any reverse correction between them. The Random Forest classifier was trained on 44 manually annotated scenes, geographically disjoint from the twelve independent scenes used for quantitative validation. This validation, based on an instance matching protocol (precision, recall, F1 score, and IoU), confirms the contribution of the full pipeline over a Random-Forest-only baseline: Pooled false positives fall by close to two orders of magnitude (from 8320 to 209), while true positives rise nearly twentyfold (from 5 to 95), with a mean IoU of 0.742 ± 0.060 on correctly matched objects. Notably, the entire pipeline—including SAM-based segmentation—runs end-to-end on a modest, GPU-free consumer laptop (four logical CPU cores, under 16 GB RAM), demonstrating that competitive infrastructure-extraction performance does not require specialised computing hardware. Full article
(This article belongs to the Section Image and Video Processing)
Show Figures

Graphical abstract

24 pages, 15815 KB  
Article
Domain Generalization of Histopathology Foundation Models in Multicenter, Multi-Scanner Cohorts: A Comparative Benchmark
by Hafsa Akebli and Vincenzo Della Mea
J. Imaging 2026, 12(8), 381; https://doi.org/10.3390/jimaging12080381 - 13 Aug 2026
Viewed by 324
Abstract
Histopathology foundation models (FMs) have become widely used as patch-level feature extractors in computational pathology (CPath), where domain shift is a central challenge, yet their generalization ability across acquisition centers and scanning platforms remains insufficiently studied. In this work, we evaluate the domain [...] Read more.
Histopathology foundation models (FMs) have become widely used as patch-level feature extractors in computational pathology (CPath), where domain shift is a central challenge, yet their generalization ability across acquisition centers and scanning platforms remains insufficiently studied. In this work, we evaluate the domain generalization of ten state-of-the-art FMs on two multi-source datasets with different supervision settings: SemiCOL, a colorectal cancer cohort of 499 whole-slide images (WSIs) for weakly labeled slide-level binary tumor classification, and BEETLE, a breast cancer cohort of 583 WSIs for patch-level four-class tissue classification. In an ablation-style setting, FMs are used as patch-level feature extractors, with patch embeddings mean-pooled into slide-level representations for SemiCOL, and a lightweight multi-layer perceptron trained for slide-level and patch-level classification on SemiCOL and BEETLE, respectively. To test FM domain generalization, we use three evaluation protocols: a Baseline source-mixed 5-fold cross-validation (CV) and two leave-source-out CV settings that assess cross-center and cross-scanner performance. On SemiCOL, all FMs achieve near-saturated performance, indicating stable performance under acquisition-source domain shift for slide-level Tumor vs. Benign classification. In contrast, BEETLE reveals clear generalization gaps, with scanner-induced domain shift more challenging than center-induced domain shift, and class-wise results showing that performance losses concentrate in epithelial discrimination. Overall, Virchow2 shows the strongest robustness across all evaluated protocols. These findings show that standard source-mixed CV can overestimate domain generalization across centers and scanners, and that FM choice matters in multi-source cohorts, especially for more challenging CPath tasks, where cross-domain failures are more visible. Full article
(This article belongs to the Section Medical Imaging)
Show Figures

Figure 1

25 pages, 23777 KB  
Article
Medical Textile Stain Detection Based on Chemically Enhanced Visualization and Deep Semantic Segmentation
by Wenjie Min, Junfeng He, Zhenping Wan, Jinde Chen, Zhixiang Zou and Yuandong Mo
J. Imaging 2026, 12(8), 380; https://doi.org/10.3390/jimaging12080380 - 13 Aug 2026
Viewed by 273
Abstract
Pre-wash sorting of medical textiles is essential for hospital infection control, yet accurate stain detection remains challenging because visually apparent stains often have blurred boundaries, whereas dried urine stains lack distinguishable optical features. This study proposes a medical textile stain detection method integrating [...] Read more.
Pre-wash sorting of medical textiles is essential for hospital infection control, yet accurate stain detection remains challenging because visually apparent stains often have blurred boundaries, whereas dried urine stains lack distinguishable optical features. This study proposes a medical textile stain detection method integrating chemically enhanced visualization with deep semantic segmentation. Dimethylaminocinnamaldehyde (DMACA) was used to convert latent urine stains into chemically developed stains with orange–red visual features. Based on the spatial color difference ΔE in the L*a*b* color space, 0.0183 mol/L was selected as the most suitable DMACA concentration among those tested. A dataset of 1974 images was constructed, including blood stains, chemically developed urine stains, medication stains, and uncontaminated textiles. A cascaded preprocessing strategy was applied to enhance stain boundaries and suppress textile texture noise, after which an Enhanced semantic segmentation model incorporating residual feature extraction, multiscale feature fusion, and transfer learning was used for pixel-level recognition. The IoU values for blood stains, chemically developed urine stains, and medication stains were 88.11%, 82.67%, and 89.62%, respectively. The average time required for image preprocessing and network inference was 15.39 ms per image. An input-level ablation comparison showed that DMACA-based color development increased the urine-stain IoU from 3.07% to 86.23%, demonstrating its substantial contribution to latent urine-stain detection. These results support the feasibility of integrating front-end chemical feature enhancement with back-end semantic segmentation for multiclass medical textile stain recognition under the current experimental conditions. Full article
(This article belongs to the Section Image and Video Processing)
Show Figures

Figure 1

18 pages, 3245 KB  
Article
Realistic Ultrasound Simulations of Healthy and Osteoarthritic Cartilage
by Roby Weeteling, Yuexin Qi, Rob P. A. Janssen, Keita Ito, Corrinus C. van Donkelaar, Richard G. P. Lopata and Min Wu
J. Imaging 2026, 12(8), 379; https://doi.org/10.3390/jimaging12080379 - 12 Aug 2026
Viewed by 408
Abstract
Osteoarthritis (OA) causes irreversible cartilage damage, highlighting the need for early and sensitive assessment. Current imaging modalities are limited in detecting early-stage changes. Ultrasound (US) provides a non-invasive and accessible alternative, but its clinical adoption is limited by the lack of standardized protocols [...] Read more.
Osteoarthritis (OA) causes irreversible cartilage damage, highlighting the need for early and sensitive assessment. Current imaging modalities are limited in detecting early-stage changes. Ultrasound (US) provides a non-invasive and accessible alternative, but its clinical adoption is limited by the lack of standardized protocols and reliable cartilage assessment. Simulations can be used to address these challenges by enabling system design, acquisition optimization and validation by providing ground truth when in vivo ground truth is unavailable. The aim of this study is to develop an in silico framework for realistic US imaging of healthy and OA cartilage by combining accurate acoustic wave modeling with a 2D microstructural cartilage phantom. The model was calibrated to healthy cartilage using first-order speckle statistics and extended to simulate degeneration through changes in structural and acoustic properties. As a proof-of-concept study, simulations were evaluated against limited ex vivo US data from healthy and OA cartilage and compared with literature data. The simulations reproduced key OA-related features and trends, including changes in reflection coefficient (R), integrated reflection coefficient (IRC), apparent integrated backscatter (AIB), and gray level distributions. These findings demonstrate the feasibility of using microstructure-based tissue phantoms to model healthy and OA cartilage. The framework provides a platform for systematic investigation of cartilage microstructure and US-derived features and may support future generation of synthetic datasets for data-driven and AI-based OA assessment. Full article
(This article belongs to the Section Medical Imaging)
Show Figures

Figure 1

21 pages, 8413 KB  
Article
A Two-Stage Ensemble Machine Learning Pipeline for Breast Cancer Diagnosis from Digital Mammograms
by Fernando Martín-Rodríguez, Carmen Freire-Bouza, Mónica Fernández-Barciela, Ainhoa Morales-Fernández and María Marante-Boado
J. Imaging 2026, 12(8), 378; https://doi.org/10.3390/jimaging12080378 - 12 Aug 2026
Viewed by 288
Abstract
Breast cancer is the most common cancer among women, and early detection through mammography is essential for reducing mortality. Artificial intelligence can support radiologists by improving diagnostic accuracy. To develop and evaluate a two-stage ensemble machine learning pipeline for breast cancer diagnosis from [...] Read more.
Breast cancer is the most common cancer among women, and early detection through mammography is essential for reducing mortality. Artificial intelligence can support radiologists by improving diagnostic accuracy. To develop and evaluate a two-stage ensemble machine learning pipeline for breast cancer diagnosis from digital mammograms. The proposed framework combines image preprocessing, multiple convolutional neural networks trained under different conditions, and a second-stage classifier that integrates the CNN outputs. Several machine learning models and feature selection techniques were evaluated using publicly available mammography datasets. Results: The ensemble approach consistently outperformed the individual CNN models. The MLP classifier achieved the best overall balance between precision and recall, while the heuristic fusion method provided the highest sensitivity. Feature selection reduced model complexity while maintaining comparable performance, and cross-validation confirmed the robustness of the proposed methodology. Combining complementary information from multiple CNNs with classical machine learning improves diagnostic performance and provides a robust framework for computer-aided breast cancer diagnosis. The proposed two-stage ensemble offers an effective and interpretable approach for mammographic breast cancer classification. A demonstration application incorporating Grad-CAM explainability further supports its potential use as a clinical decision-support tool. Full article
(This article belongs to the Special Issue AI-Driven Medical Image Processing and Analysis)
Show Figures

Graphical abstract

22 pages, 5149 KB  
Article
MoR–Swin: Efficient Vision Transformer Using Mixture of Recursions
by Yongbao Ai, Tianxiang Gao, Zhipeng Lin, Longqi Yang and Qingyu Chang
J. Imaging 2026, 12(8), 377; https://doi.org/10.3390/jimaging12080377 - 12 Aug 2026
Viewed by 242
Abstract
Vision Transformers, especially Swin Transformer, have become default backbones for various vision tasks but suffer from high memory consumption and training costs. This letter proposes MoR–Swin, a novel architecture that integrates Mixture of Recursions (MoR) into Swin Transformer. An adaptive token-level recursion mechanism [...] Read more.
Vision Transformers, especially Swin Transformer, have become default backbones for various vision tasks but suffer from high memory consumption and training costs. This letter proposes MoR–Swin, a novel architecture that integrates Mixture of Recursions (MoR) into Swin Transformer. An adaptive token-level recursion mechanism dynamically allocates computational depth based on semantic complexity. A recursive window attention module and a lightweight router with load balancing loss are introduced. Extensive experiments on ImageNet classification, COCO detection, and ADE20K segmentation show that MoR–Swin reduces parameters by about 50% and accelerates inference up to twofold at a modest accuracy cost (within about 0.5 points of Swin-B on ImageNet-1K). It provides a new technical pathway for optimizing Vision Transformer models, significantly enhancing their applicability in resource-constrained environments. Full article
Show Figures

Figure 1

19 pages, 8375 KB  
Article
Temporal Feature Interaction for Robust Remote Sensing Image Change Detection: A Taxonomy and Cross-Domain Comparative Study
by Mostafa Mosaad, Mahmoud Ahmed, Fawzy Eltohamy, Tarek A. Mahmoud and Mohamed E. Hanafy
J. Imaging 2026, 12(8), 376; https://doi.org/10.3390/jimaging12080376 - 11 Aug 2026
Viewed by 258
Abstract
Remote sensing change detection (RSCD) has advanced through convolutional, attention-based, transformer, and hybrid architectures, yet models are commonly compared as whole architectural families rather than by how their temporal streams interact. This study introduces the Temporal Interaction Taxonomy (TIT), which characterizes temporal feature [...] Read more.
Remote sensing change detection (RSCD) has advanced through convolutional, attention-based, transformer, and hybrid architectures, yet models are commonly compared as whole architectural families rather than by how their temporal streams interact. This study introduces the Temporal Interaction Taxonomy (TIT), which characterizes temporal feature interaction by timing, direction, and operator. Nine representative models, ranging from early-fusion convolutional baselines to hybrid CNN–Transformer designs, were evaluated using faithful Open-CD implementations under a unified protocol on the LEVIR-CD and WHU-CD building change datasets. TIT provided a consistent basis for describing bi-temporal integration across architectures. Cross-dataset performance was model- and direction-dependent: ChangeFormer achieved the highest mIoU in both transfer directions, reaching 59.99% for WHU→LEVIR and 82.58% for LEVIR→WHU. The results suggest an association between interaction design and cross-dataset robustness, but the independent contribution of temporal interaction cannot be separated from other architectural differences. Transfer also showed model-dependent directional asymmetry; its causes could not be isolated because dataset characteristics and model design were not independently controlled. Overall, temporal interaction provides a useful dimension for interpreting model behavior across the two evaluated datasets. These findings are limited to building change detection on LEVIR-CD and WHU-CD and require task-specific validation before extension to other RSCD applications. Full article
(This article belongs to the Topic Intelligent Image Processing Technology, 2nd Edition)
Show Figures

Graphical abstract

18 pages, 2783 KB  
Article
WVM-UNet: A Wavelet–Vision Mamba Framework for Enhanced Medical Image Segmentation
by Yulong Yang, Wen Gao, Zhengguo Wu and Chuanghua Yang
J. Imaging 2026, 12(8), 375; https://doi.org/10.3390/jimaging12080375 - 11 Aug 2026
Viewed by 279
Abstract
Accurate segmentation of skin lesions and gastrointestinal polyps is essential for early diagnosis and treatment planning. Currently, Convolutional Neural Networks (CNNs) are limited by local receptive fields, missing small lesions. While Transformers model global context, their quadratic computational complexity incurs high costs. To [...] Read more.
Accurate segmentation of skin lesions and gastrointestinal polyps is essential for early diagnosis and treatment planning. Currently, Convolutional Neural Networks (CNNs) are limited by local receptive fields, missing small lesions. While Transformers model global context, their quadratic computational complexity incurs high costs. To address these limitations, we propose the Wavelet–Vision Mamba UNet (WVM-UNet), integrating State Space Models (SSMs) for linear-complexity long-range dependencies and wavelet transforms for fine-grained feature extraction. The network employs a Wavelet-based Residual State Space (WRSS) block, combining the multi-scale decomposition of discrete wavelet transforms with Vision Mamba to efficiently capture global features. A Fused Channel–Spatial Attention (FCSA) mechanism is incorporated to adaptively recalibrate feature representations. Additionally, we construct an Encoder–Decoder Semantic Connection (EDSC) to replace traditional skip connections, effectively bridging the semantic gap between cross-level features. Experimental results on multiple public datasets demonstrate the competitive performance of our method. Specifically, on the ISIC 2017 dataset, WVM-UNet achieves an mIoU of 82.94% and a DSC of 90.67%, outperforming the Mamba-based VM-UNet by 2.71% in mIoU. These results indicate our architecture effectively captures discriminative features for precise medical image segmentation. Full article
(This article belongs to the Section Medical Imaging)
Show Figures

Figure 1

20 pages, 2524 KB  
Article
Exploring Prototype Networks for Surgical Vision: Interpretability and Performance in Semantic Segmentation and Surgical Phase Recognition
by Yiping Li, Ronald L. P. D. de Jong, Franco Badaloni, Gino M. Kuiper, Romy C. van Jaarsveld, Jelle P. Ruurda and Marcel Breeuwer
J. Imaging 2026, 12(8), 374; https://doi.org/10.3390/jimaging12080374 - 11 Aug 2026
Viewed by 348
Abstract
Deep learning-based surgical vision systems achieve strong performance in semantic segmentation and phase recognition, but their black-box nature limits traceability in safety-critical clinical settings. Prototype-based networks offer an interpretable alternative by grounding predictions in learned visual exemplars, yet their suitability for surgical video [...] Read more.
Deep learning-based surgical vision systems achieve strong performance in semantic segmentation and phase recognition, but their black-box nature limits traceability in safety-critical clinical settings. Prototype-based networks offer an interpretable alternative by grounding predictions in learned visual exemplars, yet their suitability for surgical video understanding remains insufficiently characterized. We adapted a prototype-based architecture to two surgical datasets, laparoscopic cholecystectomy and robot-assisted minimally invasive esophagectomy (RAMIE), and benchmarked it against conventional baselines. We evaluated a segmentation-only setting, in which prototype size and capacity were ablated, and a multitask setting, in which three strategies for coupling prototype learning to semantic segmentation and surgical phase recognition were compared. Prototype-based models underperformed the conventional baselines across both tasks and datasets. In the segmentation-only setting, the selected prototype configurations achieved Dice scores of 72.23% on Cholecystectomy and 72.07% on RAMIE, compared with 74.35% and 74.02% for the corresponding conventional baselines, and showed weaker boundary agreement. In the multitask setting, the best prototype strategy recovered competitive segmentation performance but remained 6–12 F1 points below the conventional baseline for phase recognition. Qualitatively, prototype activation maps exposed intra-structure decompositions and contextual cues that are not directly available from black-box baselines. Prototype networks provide spatially traceable evidence for surgical scene understanding, but currently trade interpretability for reduced boundary precision and phase-recognition performance. These findings motivate future work on scene-level and temporally aware prototypes for explainable surgical AI. Full article
Show Figures

Figure 1

11 pages, 1263 KB  
Article
Evaluating the Number of Trials for Stable Virtual Reality-Based Subjective Visual Vertical Measurement in Healthy Adults
by Tameto Naoi, Jun Watanabe, Keisuke Hamada and Mitsuya Morita
J. Imaging 2026, 12(8), 373; https://doi.org/10.3390/jimaging12080373 - 11 Aug 2026
Viewed by 213
Abstract
Virtual reality-based subjective visual vertical (VR-SVV) has attracted attention as a potential solution to equipment-related limitations of conventional SVV testing. This study investigated the test–retest reliability and adequate trial number for stable assessment of vertical perception using VR-SVV in healthy adults. Participants performed [...] Read more.
Virtual reality-based subjective visual vertical (VR-SVV) has attracted attention as a potential solution to equipment-related limitations of conventional SVV testing. This study investigated the test–retest reliability and adequate trial number for stable assessment of vertical perception using VR-SVV in healthy adults. Participants performed 10 trials in a VR-SVV test and repeated the assessment after one week. SVV orientation and SVV variability were calculated. Test–retest reliability was evaluated using the intraclass correlation coefficient (ICC [1,2]), standard error of measurement (SEM), and minimal detectable change at the 95% confidence level (MDC95). The minimum number of trials required for stable assessment was examined by comparing results from fewer trials with those obtained from all 10 trials. SVV orientation and variability were −0.14° and 0.79°, respectively. ICC, SEM, and MDC95 were 0.62, 0.64°, and 1.76°, respectively. No adverse events occurred. SVV metrics derived from 6–9 trials differed by less than 10% from those obtained from 10 trials. VR-SVV is feasible and demonstrates moderate test–retest reliability. Six trials may be adequate for practical assessment of vertical perception. Full article
(This article belongs to the Section Computational Imaging and Computational Photography)
Show Figures

Figure 1

24 pages, 2526 KB  
Article
Hybrid PCA–LBP and Wavelet Scattering Framework for Texture Classification in Color Images
by Zahoor M. Aydam, Baidaa Mutasher Rashed and Nidhal K. El Abbadi
J. Imaging 2026, 12(8), 372; https://doi.org/10.3390/jimaging12080372 - 11 Aug 2026
Viewed by 212
Abstract
Color texture classification is an important task in computer vision, with applications in medical imaging, industrial inspection, remote sensing, and material analysis. This paper presents a hybrid framework that integrates Principal Component Analysis (PCA), Local Binary Patterns (LBPs), Wavelet Scattering Transform, and the [...] Read more.
Color texture classification is an important task in computer vision, with applications in medical imaging, industrial inspection, remote sensing, and material analysis. This paper presents a hybrid framework that integrates Principal Component Analysis (PCA), Local Binary Patterns (LBPs), Wavelet Scattering Transform, and the XGBoost classifier for color texture classification. The proposed pipeline first performs image pre-processing, including resizing and denoising, followed by channel-wise feature extraction using LBP and Wavelet Scattering Transform on the Red, Green, and Blue channels independently. Then, the obtained feature vectors were concatenated, and PCA was applied on the fused feature space for dimensionality reduction and redundancy elimination before proceeding to XGBoost classification. This method not only leverages complementary information of Chroma and texture information but also achieves reduced dimensionality and computational burden. The finally optimized features were input into the XGBoost classifier for color texture classification, which is good at fitting non-linear dependency and includes a regularization to generalize better. Our proposed framework was tested on three benchmark color texture datasets: KTH-TIPS, Outex_10, and VisTex. Experimental results have demonstrated that on these three datasets, the average performance reaches 98.0% accuracy, 0.981 precision, 0.981 recall, and 0.979 F1-score, respectively. It demonstrates that the two selected complementary feature extraction methods provide a compact yet effective representation for color texture classification on these datasets. It is expected that the proposed framework serves as an efficient combination of established methods and as a good competitive baseline for color texture analysis. Future works will consider applying it to larger color texture datasets for general verification, enhancing its computational efficiency and automating the parameter selection process. Full article
(This article belongs to the Section Computer Vision and Pattern Recognition)
Show Figures

Figure 1

2 pages, 134 KB  
Correction
Correction: Dash et al. Improving Object Detection in High-Altitude Infrared Thermal Images Using Magnitude-Based Pruning and Non-Maximum Suppression. J. Imaging 2025, 11, 69
by Yajnaseni Dash, Vinayak Gupta, Ajith Abraham and Swati Chandna
J. Imaging 2026, 12(8), 371; https://doi.org/10.3390/jimaging12080371 - 11 Aug 2026
Viewed by 177
Abstract
In the original publication [...] Full article
(This article belongs to the Section Computer Vision and Pattern Recognition)
27 pages, 37755 KB  
Article
Open Long-Tailed Multimodal 3D Model Classification Based on Sample-Enhanced Category-Space Learning
by Yuansa Wang, Xueyao Gao, Chunxiang Zhang and Yongzeng Xue
J. Imaging 2026, 12(8), 370; https://doi.org/10.3390/jimaging12080370 - 10 Aug 2026
Viewed by 222
Abstract
With the rapid development of three-dimensional (3D) sensing technologies, multimodal 3D model classification has achieved significant progress. However, most existing methods are developed under closed and balanced assumptions, which limits their applicability to open long-tailed scenarios with scarce tail classes, ambiguous hard samples, [...] Read more.
With the rapid development of three-dimensional (3D) sensing technologies, multimodal 3D model classification has achieved significant progress. However, most existing methods are developed under closed and balanced assumptions, which limits their applicability to open long-tailed scenarios with scarce tail classes, ambiguous hard samples, and continuously emerging categories. In this work, we propose sample-enhanced category-space learning (SE-CSL) for open long-tailed multimodal 3D model classification. The proposed method first uses dual-branch modality encoders to extract point-cloud structural representations and multi-view semantic representations. Mamba is then introduced to model global dependencies across heterogeneous modalities and generate a unified global category representation. To improve the robustness of category representation, we design a category-space learning strategy that jointly integrates long-tailed learning, few-shot representation stabilization, and hard-sample enhancement. A long-tail balanced loss, a few-shot stabilization loss, and a hard-sample boundary loss are further developed to optimize intra-class compactness, inter-class separability, and boundary discrimination. To handle continually emerging classes, we introduce an incremental category-space expansion mechanism that distinguishes new classes, preserves old-class information, and supports unified classification of old and new categories. Extensive experiments on ModelNet40 and ShapeNet55 demonstrate the effectiveness and robustness of SE-CSL. Full article
(This article belongs to the Section Computer Vision and Pattern Recognition)
Show Figures

Figure 1

24 pages, 3613 KB  
Article
RG-PSR: Reliability-Guided Poisson Surface Reconstruction for Degraded 3D-Imaging Point Clouds
by Na Liu, Fan Zhang, Jiawei Wang, Dan Zhang, Jinliang Wu and Xiaohui Li
J. Imaging 2026, 12(8), 369; https://doi.org/10.3390/jimaging12080369 - 10 Aug 2026
Viewed by 302
Abstract
Three-dimensional (3D) imaging systems, including depth cameras, LiDAR sensors, and multi-view scanning pipelines, often produce point clouds with noisy normals, outliers, sparse sampling, and non-uniform density, which can degrade downstream mesh reconstruction. Poisson surface reconstruction is lightweight and training-free, but its global implicit [...] Read more.
Three-dimensional (3D) imaging systems, including depth cameras, LiDAR sensors, and multi-view scanning pipelines, often produce point clouds with noisy normals, outliers, sparse sampling, and non-uniform density, which can degrade downstream mesh reconstruction. Poisson surface reconstruction is lightweight and training-free, but its global implicit formulation is sensitive to unreliably oriented samples and fixed density-trimming thresholds. This paper presents RG-PSR, a reliability-guided enhancement framework for Poisson-family surface reconstruction from degraded 3D-imaging point clouds. RG-PSR estimates a deterministic per-point reliability score from local density regularity, spacing variation, and normal consistency, and propagates this score through conservative point filtering, reliability-guided normal refinement, adaptive density-reliability trimming, and structure-aware postprocessing. The main pipeline requires no manual labels, neural network training, or ground-truth meshes at inference time. Experiments on three groups of object meshes under five deterministic degradation types show that RG-PSR improves Poisson-family reconstruction under degraded inputs. Compared with fixed density-trimmed Poisson reconstruction, RG-PSR reduces the overall Chamfer-L1 from 0.0218 to 0.0172, improves F0.01 from 0.6618 to 0.6836, and reduces Artifact0.02 from 0.3090 to 0.2632. In the broader classical comparison, local triangulation methods achieve stronger point-wise accuracy, while RG-PSR yields the fewest connected components and the highest largest-component ratio. These results position RG-PSR as a practical reliability layer for coherent Poisson-family reconstruction rather than a universal replacement for all surface-reconstruction methods. Full article
(This article belongs to the Special Issue Advances in 3D Point Cloud Processing)
Show Figures

Figure 1

Previous Issue
Back to TopTop