Topic Editors

State Key Laboratory of Media Convergence and Communication, Communication University of China, Beijing, China
Prof. Dr. Didier El Baz
LAAS-CNRS, Université de Toulouse, CNRS, 31031 Toulouse, France
School of Information Science and Engineering, Hunan Normal University, Changsha, China
Dr. Yujin Zhang
School of Electronic and Electrical Engineering, Shanghai University of Engineering Science, Shanghai 201620, China

Intelligent Image Processing Technology

Abstract submission deadline
closed (28 February 2026)
Manuscript submission deadline
closed (30 April 2026)
Viewed by
29625

Topic Information

Dear Colleagues,

Rapid advancements in intelligent image processing technology have revolutionized various fields by enabling more accurate, efficient, and innovative approaches to image analysis and interpretation. This Topic seeks to explore the multifaceted applications and theoretical developments driving this transformative discipline forward.

Intelligent image processing integrates sophisticated algorithms, machine learning techniques, and computational methodologies to extract meaningful information from digital images. It plays a crucial role in various domains such as medical imaging, remote sensing, industrial automation, and beyond.

Contributions to this Topic will encompass the following topics:

  • Algorithmic Innovations: novel algorithms for image enhancement, segmentation, object detection, and pattern recognition.
  • Applications in Remote Sensing: the utilization of intelligent image processing for analyzing satellite imagery and aerial photography.
  • Medical Image Analysis: advances in diagnostic imaging, image registration, and computer-aided diagnosis.
  • Integration with Sensor Networks: enhancing sensor data interpretation and analysis.
  • Industrial Automation: applications in robotics, quality control, and manufacturing processes.
  • Machine Learning and Reinforcement Learning: the optimization of image processing tasks and system performance.
  • Multimodal Information Processing: integration with natural language processing, speech recognition, and data mining.
  • Autonomous Agents: enhancing decision-making and control capabilities.

Researchers are invited to contribute original research articles, reviews, and methodological studies that explore the frontiers of intelligent image processing across these different areas. This Topic aims to foster interdisciplinary collaboration, showcase cutting-edge developments, and outline future directions in this dynamic field.

By providing a comprehensive platform for researchers and practitioners, this Topic aims to advance the understanding and application of intelligent image processing technology, paving the way for new discoveries and innovations across various disciplines.

Dr. Lei Shi
Prof. Dr. Didier El Baz
Prof. Dr. Jinping Liu
Dr. Yujin Zhang
Topic Editors

Keywords

  • algorithmic innovations
  • remote sensing
  • medical image analysis
  • sensor networks
  • industrial automation
  • machine learning

Participating Journals

Journal Name Impact Factor CiteScore Launched Year First Decision (median) APC
Applied Sciences
applsci
2.9 6.1 2011 15 Days CHF 2400
Automation
automation
2.9 4.5 2020 24.8 Days CHF 1200
Electronics
electronics
2.9 7.0 2012 14.8 Days CHF 2400
Information
information
4.3 8.2 2010 18.7 Days CHF 1800
Journal of Imaging
jimaging
3.8 7.3 2015 21.3 Days CHF 1800
Mathematics
mathematics
2.3 5.4 2013 17.4 Days CHF 2600

Preprints.org is a multidisciplinary platform offering a preprint service designed to facilitate the early sharing of your research. It supports and empowers your research journey from the very beginning.

MDPI Topics is collaborating with Preprints.org and has established a direct connection between MDPI journals and the platform. Authors are encouraged to take advantage of this opportunity by posting their preprints at Preprints.org prior to publication:

  1. Share your research immediately: disseminate your ideas prior to publication and establish priority for your work.
  2. Safeguard your intellectual contribution: Protect your ideas with a time-stamped preprint that serves as proof of your research timeline.
  3. Boost visibility and impact: Increase the reach and influence of your research by making it accessible to a global audience.
  4. Gain early feedback: Receive valuable input and insights from peers before submitting to a journal.
  5. Ensure broad indexing: Web of Science (Preprint Citation Index), Google Scholar, Crossref, SHARE, PrePubMed, Scilit and Europe PMC.

Published Papers (20 papers)

Order results
Result details
Journals
Select all
Export citation of selected articles as:
21 pages, 982 KB  
Article
PaIR: Partition-Based Information Rebalancing for Robust Text-Based Person Search
by Luda Wang, Jiabao Li, Xinpan Yuan and Ningdan Zhang
J. Imaging 2026, 12(9), 400; https://doi.org/10.3390/jimaging12090400 - 25 Aug 2026
Viewed by 214
Abstract
Text-based person search (TPS) suffers from cross-modal informational skewness: pedestrian images are high-dimensional and redundancy-prone, while textual descriptions are sparse, incomplete, and sometimes inaccurate. To address the low alignment accuracy and poor robustness caused by the inherent uneven information distribution of visual and [...] Read more.
Text-based person search (TPS) suffers from cross-modal informational skewness: pedestrian images are high-dimensional and redundancy-prone, while textual descriptions are sparse, incomplete, and sometimes inaccurate. To address the low alignment accuracy and poor robustness caused by the inherent uneven information distribution of visual and textual modalities in TPS, this paper proposes a unified Partition-based Information Rebalancing (PaIR) framework to realize balanced optimization and precise alignment of cross-modal information from both global content and local part dimensions. The framework adopts the CLIP dual-modal encoder for basic feature extraction and constructs a parallel global–local dual representation system to compensate for the lack of fine-grained spatial information in single global features. To eliminate modal redundancy and noise interference, a dual-modal noise suppression module is designed to filter invalid redundant information through visual foreground–background separation and textual token weight screening, while introducing adversarial constraints and orthogonal constraints to purify effective features. On this basis, a part balance alignment module is built to complete human semantic part decomposition and soft matching alignment for dual-modal features. Aiming at the common part semantic missing problem in textual descriptions, a visual part correlation affinity matrix is utilized for semantic associative completion to balance the information density of dual modalities. Finally, a global–local joint alignment strategy integrates hierarchical features and bidirectional cross-modal attention interaction to eliminate global–local semantic discontinuity and enhance fine-grained cross-modal matching capability. Extensive experiments on three public benchmarks demonstrate that PaIR consistently improves multiple baselines. Full article
(This article belongs to the Topic Intelligent Image Processing Technology)
Show Figures

Figure 1

14 pages, 1321 KB  
Article
Road Damage Detection with Direction Awareness and Feature Equalization
by Yutao Wang, Zhengzheng Zhu, Yongqiang Bai, Zhibo Xie and Renwei Tu
Information 2026, 17(7), 683; https://doi.org/10.3390/info17070683 - 14 Jul 2026
Viewed by 312
Abstract
In the field of road damage detection, the accuracy of existing methods still requires further improvement, particularly for elongated cracks, which are crucial for ensuring driving safety and effective road maintenance. To address this limitation, a novel road damage detection algorithm is proposed [...] Read more.
In the field of road damage detection, the accuracy of existing methods still requires further improvement, particularly for elongated cracks, which are crucial for ensuring driving safety and effective road maintenance. To address this limitation, a novel road damage detection algorithm is proposed based on direction awareness and feature equalization. Specifically, a Direction-aware Strip Convolution (DSC) module is constructed to effectively capture the geometric characteristics of elongated cracks and maintain computational efficiency, by integrating asymmetric strip convolution and depthwise separable convolution respectively. In addition, a Multi-level Feature Equalization (MFE) module is designed to address the complex morphology and significant scale variations of road damage during multi-level feature fusion. Specifically, a set of learnable spatial weighting parameters is introduced in this module, whose weighting coefficients are optimized across different network layers and adaptively generated, thereby modulating the contributions of multi-level features and promoting a more balanced multi-level feature representation. Experimental results on the RDD2022-based experimental dataset demonstrate that the proposed method improves mAP@50 by 5.8 percentage points and recall by 5.9 percentage points, while achieving a processing speed of 122 FPS. Notably, the proposed method improves the detection performance of elongated cracks and achieves relatively balanced performance gains across different road damage categories, compared with the baseline model. Full article
(This article belongs to the Topic Intelligent Image Processing Technology)
Show Figures

Figure 1

17 pages, 1853 KB  
Article
GKMANet: Self-Attention Gaussian Kernel Mixture Network for Defocus Deblurring
by Fei Zhang, Yinghui Wang, Shangqian Zhuo and Pengbo Wang
Electronics 2026, 15(12), 2592; https://doi.org/10.3390/electronics15122592 - 12 Jun 2026
Viewed by 278
Abstract
The primary goal of defocus deblurring is to restore details in blurred images, enhancing clarity and usability. Despite recent progress, current methods struggle with uneven blur distribution, particularly in distinguishing multi-scale boundaries, causing inaccurate pixel blur estimation and boundary detail loss. To address [...] Read more.
The primary goal of defocus deblurring is to restore details in blurred images, enhancing clarity and usability. Despite recent progress, current methods struggle with uneven blur distribution, particularly in distinguishing multi-scale boundaries, causing inaccurate pixel blur estimation and boundary detail loss. To address this, we propose an end-to-end deep learning approach incorporating a self-attention mechanism. Our method employs a Gaussian Kernel Mixture (GKM) model for compact blur kernel representation, then builds upon the Gaussian Kernel Mixture Network (GKMNet) to design a novel Attention-based GKM Network (GKMANet) via iterative fixed-point expansion. GKMANet introduces a scale-recursive architecture with self-attention to estimate mixing coefficients for effective deblurring. Experiments demonstrate significant superiority over existing techniques, especially in scenarios with indistinct blurred/clear boundaries, highlighting stronger robustness and higher restoration accuracy. This provides a reliable solution for detail recovery in defocus deblurring. Full article
(This article belongs to the Topic Intelligent Image Processing Technology)
Show Figures

Figure 1

22 pages, 2008 KB  
Article
A Video Frame Prediction Method Based on Latent-Space Autoregressive Modeling
by Congcong Zhang, Jin Tian, Lihua Gong, Yujin Zhang, Fei Wu and Han Pan
Appl. Sci. 2026, 16(9), 4423; https://doi.org/10.3390/app16094423 - 1 May 2026
Viewed by 703
Abstract
Video prediction is a fundamental task in computer vision with broad applications in intelligent robotics, autonomous driving, and related fields. However, existing methods often struggle to simultaneously model long-term temporal dependencies, preserve local details, and alleviate error accumulation during autoregressive prediction. To address [...] Read more.
Video prediction is a fundamental task in computer vision with broad applications in intelligent robotics, autonomous driving, and related fields. However, existing methods often struggle to simultaneously model long-term temporal dependencies, preserve local details, and alleviate error accumulation during autoregressive prediction. To address these issues, this paper proposes a two-stage video prediction framework composed of a HybridResSwin Autoencoder (HRS-AE) and an Enhanced FAR Transformer (EFAR). In the first stage, HRS-AE learns compact and discriminative latent representations from input video frames while preserving essential spatial structures and fine-grained details. In the second stage, EFAR performs autoregressive temporal prediction in the latent space, and the predicted latent representations are then decoded to reconstruct future video frames. Experiments on the KTH, BAIR, and Moving MNIST datasets show that the proposed method achieves competitive performance under the adopted evaluation protocol. Specifically, the proposed framework achieves a PSNR of 30.27 dB and an LPIPS of 0.0722 on KTH, a PSNR of 20.95 dB on BAIR, and an SSIM of 0.961 with an MSE of 22.9 on Moving MNIST. In addition, ablation studies further indicate that the proposed components contribute to latent representation learning and long-horizon prediction stability. These results suggest that the proposed framework provides a promising approach for video prediction with favorable reconstruction quality, perceptual consistency, and temporal coherence. Full article
(This article belongs to the Topic Intelligent Image Processing Technology)
Show Figures

Figure 1

20 pages, 10840 KB  
Article
Infrared Small-Target Segmentation Framework Based on Morphological Attention and Energy Core Loss
by Baoyu Zhu, Qunbo Lv, Yangyang Liu, Haoran Cao and Zheng Tan
J. Imaging 2026, 12(5), 184; https://doi.org/10.3390/jimaging12050184 - 24 Apr 2026
Viewed by 597
Abstract
Infrared small-target segmentation (IRSTS) is crucial for a wide range of applications, including maritime search-and-rescue operations and intelligent traffic surveillance. However, current deep learning methods struggle with dynamic scale variations in infrared small targets, resulting in false detections and missed detections, alongside inadequate [...] Read more.
Infrared small-target segmentation (IRSTS) is crucial for a wide range of applications, including maritime search-and-rescue operations and intelligent traffic surveillance. However, current deep learning methods struggle with dynamic scale variations in infrared small targets, resulting in false detections and missed detections, alongside inadequate core localization accuracy. To address these challenges, we propose an infrared small-target segmentation framework founded on morphological attention and an energy core loss function, IRSTS_Unet. Specifically, we design a Dynamic Shape-adaptive Deformable Attention Module (DSDAM), which achieves parameterized feature extraction via “initial localization–offset deformation–precise sampling”. This approach enables the network to differentially focus on target cores and background cues to suppress clutter. To improve the efficiency of multi-scale feature aggregation, we embed the DSDAM within both the feature extraction and cross-layer fusion stages. Furthermore, we formulate a Core Energy-aware Core-Priority loss (CECP-Loss) function that incorporates the energy prior distribution of small targets, effectively counteracting the “core dilution” phenomenon endemic to conventional loss functions. Through extensive experiments on multiple public datasets, we demonstrate that IRSTS_U-Net outperforms state-of-the-art approaches in terms of both detection accuracy and robustness. Full article
(This article belongs to the Topic Intelligent Image Processing Technology)
Show Figures

Figure 1

17 pages, 144369 KB  
Article
A Portable Multimodal Imaging System for Glossy Ceramic Pattern Acquisition with Highlight Suppression and Acquisition Optimization
by Wenxin Lei, Xiaochuan Ming and Jiyang Gao
Electronics 2026, 15(8), 1573; https://doi.org/10.3390/electronics15081573 - 9 Apr 2026
Viewed by 496
Abstract
Acquiring decorative patterns from glossy ceramic surfaces is challenging because specular reflections often obscure fine details and reduce the reliability of subsequent digital analysis. Although existing highlight-removal methods, including data-driven and single-image enhancement approaches, have improved restoration quality in generic scenes, they are [...] Read more.
Acquiring decorative patterns from glossy ceramic surfaces is challenging because specular reflections often obscure fine details and reduce the reliability of subsequent digital analysis. Although existing highlight-removal methods, including data-driven and single-image enhancement approaches, have improved restoration quality in generic scenes, they are not fully suited to glossy ceramic documentation because they often rely on scene priors, large paired datasets, or post hoc enhancement alone while paying limited attention to acquisition-side optimization for reflective cultural objects. This article presents a portable multimodal imaging system and a processing framework for ceramic pattern acquisition, highlight suppression, and acquisition optimization. Multimodal images captured under different illumination and polarization configurations are first geometrically registered, after which specular regions are localized by jointly exploiting polarization and intensity cues, followed by highlight suppression and perceptual appearance restoration to improve pattern visibility while preserving visual authenticity. Experimental results indicates that warm illumination with 0° polarization is more suitable for warm-toned ceramics or ceramics with large-area patterns, whereas uniform illumination with 45° polarization is more suitable for cool-toned ceramics and ceramics with sparse patterns; additionally, cool illumination with 90° polarization yields the highest average score across the dataset, indicating stronger robustness across diverse samples. The proposed system is portable, supports wireless image transmission, and integrates adjustable illumination with a servo-driven polarizer, thereby providing a practical solution for high-quality digital documentation of glossy ceramic patterns. Full article
(This article belongs to the Topic Intelligent Image Processing Technology)
Show Figures

Figure 1

19 pages, 4570 KB  
Article
Adaptive Deletion of Gaussian Ellipsoids in 3D Gaussian Splatting
by Fei Zhang, Yinghui Wang, Bo Yi and Jiaxin Ma
Mathematics 2026, 14(7), 1197; https://doi.org/10.3390/math14071197 - 3 Apr 2026
Viewed by 918
Abstract
As a leading method for Novel View Synthesis (NVS), 3D Gaussian Splatting (3DGS) faces limitations. Fixed thresholds governing Gaussian scale and opacity lead to over-reconstruction or under-reconstruction, while the linear penalty used for handling outliers during optimization tends to introduce artifacts. Therefore, we [...] Read more.
As a leading method for Novel View Synthesis (NVS), 3D Gaussian Splatting (3DGS) faces limitations. Fixed thresholds governing Gaussian scale and opacity lead to over-reconstruction or under-reconstruction, while the linear penalty used for handling outliers during optimization tends to introduce artifacts. Therefore, we propose Adaptive 3DGS featuring a dynamic deletion mechanism. Specifically, our method calculates coverage for each Gaussian based on its scale during removal. Gaussians with high coverage face stricter scale thresholds to reduce over-reconstruction, while those with lower coverage receive lenient thresholds to preserve details. Simultaneously, transparency-based contribution assessment is applied. Gaussians with low contribution meet stricter transparency thresholds to combat over-reconstruction, while high-contribution ones get lenient thresholds to mitigate under-reconstruction. During optimization, introducing Huber loss promotes quadratic growth for small errors, reducing smoothing to alleviate artifacts and better preserve details. Evaluation on standard datasets shows our method improves peak signal-to-noise ratio (PSNR) by 0.3 dB over 3DGS and 0.5 dB over MS-3DGS at 4× resolution, and it achieves a 0.1 dB gain over Mip-Splatting, confirming its effectiveness and robustness. Full article
(This article belongs to the Topic Intelligent Image Processing Technology)
Show Figures

Figure 1

23 pages, 3274 KB  
Article
Question-Aware Reasoning Framework via Two-Level Cross-Attention
by Junhui Bai, Jun Wu, Mingyu Li, Shichao Yu, Ziming Jiang and Yinghui Wang
Mathematics 2026, 14(5), 857; https://doi.org/10.3390/math14050857 - 3 Mar 2026
Viewed by 888
Abstract
Multimodal chain-of-thought (CoT) reasoning has emerged as a pivotal research direction in artificial intelligence. However, current approaches predominantly adopt a linear CoT structure with a single reasoning module and complex multi-level gated multi-hop cross-attention mechanisms for modality fusion, which exhibit notable limitations. Specifically, [...] Read more.
Multimodal chain-of-thought (CoT) reasoning has emerged as a pivotal research direction in artificial intelligence. However, current approaches predominantly adopt a linear CoT structure with a single reasoning module and complex multi-level gated multi-hop cross-attention mechanisms for modality fusion, which exhibit notable limitations. Specifically, the inability of linear CoT structures to dynamically select appropriate reasoning modules based on problem characteristics often leads to hallucinations during intermediate reasoning. Moreover, tightly coupled gating and cross-attention mechanisms can inadvertently suppress critical information flow during inter-modal interactions, resulting in erroneous predictions. To address these challenges, we propose a novel multimodal reasoning framework, M-TCM, that integrates a two-level cross-attention fusion mechanism with a single-level gating strategy. This design not only reduces the complexity of modality fusion but also effectively preserves information crucial for intermediate reasoning. Furthermore, M-TCM incorporates a novel module selection strategy. We first construct a new dataset, SQ-GPT4, to complement the existing ScienceQA dataset and facilitate the training of two distinct reasoning modules. Subsequently, the model dynamically selects the most appropriate reasoning module for prediction based on the specific skill requirements of each problem. Experimental results on the ScienceQA benchmark demonstrate the superiority of our proposed model, achieving a prediction accuracy of 88.23%. Full article
(This article belongs to the Topic Intelligent Image Processing Technology)
Show Figures

Figure 1

26 pages, 55590 KB  
Article
Adaptive Edge-Aware Detection with Lightweight Multi-Scale Fusion
by Xiyu Pan, Kai Xiong and Jianjun Li
Electronics 2026, 15(2), 449; https://doi.org/10.3390/electronics15020449 - 20 Jan 2026
Cited by 1 | Viewed by 833
Abstract
In object detection, boundary blurring caused by occlusion and background interference often hinders effective feature extraction. To address this challenge, we propose Edge Aware-YOLO, a novel framework designed to enhance edge awareness and efficient feature fusion. Our method integrates three key contributions. First, [...] Read more.
In object detection, boundary blurring caused by occlusion and background interference often hinders effective feature extraction. To address this challenge, we propose Edge Aware-YOLO, a novel framework designed to enhance edge awareness and efficient feature fusion. Our method integrates three key contributions. First, the Variable Sobel Compact Inverted Block (VSCIB) employs convolution kernels with adjustable orientation and size, enabling robust multi-scale edge adaptation. Second, the Spatial Pyramid Shared Convolution (SPSC) replaces standard pooling with shared dilated convolutions, minimizing detail loss during feature reconstruction. Finally, the Efficient Downsampling Convolution (EDC) utilizes a dual-branch architecture to balance channel compression with semantic preservation. Extensive evaluations on public datasets demonstrate that Edge Aware-YOLO significantly outperforms state-of-the-art models. On MS COCO, it achieves 56.3% mAP50 and 40.5% mAP50–95 (gains of 1.5% and 1.0%) with only 2.4M parameters and 5.8 GFLOPs, surpassing advanced models like YOLOv11. Full article
(This article belongs to the Topic Intelligent Image Processing Technology)
Show Figures

Figure 1

14 pages, 9414 KB  
Article
AutoMCA: A Robust Approach for Automatic Measurement of Cranial Angles
by Junjian Chen, Yuqian Wang, Xinyu Shi and Yan Luximon
Automation 2025, 6(4), 88; https://doi.org/10.3390/automation6040088 - 5 Dec 2025
Viewed by 1364
Abstract
Head posture assessment commonly involves measuring cranial angles, with photogrammetry favored for its simplicity over CT scans or goniometers. However, most photo-based measurements remain manual, making them time-consuming and inefficient. Existing automatic measuring approaches often requires specific markers and clean backgrounds, limiting their [...] Read more.
Head posture assessment commonly involves measuring cranial angles, with photogrammetry favored for its simplicity over CT scans or goniometers. However, most photo-based measurements remain manual, making them time-consuming and inefficient. Existing automatic measuring approaches often requires specific markers and clean backgrounds, limiting their usability. We present AutoMCA, a robust automatic measurement system for cranial angles using accessible markers and tolerating typical indoor backgrounds. AutoMCA integrates MediaPipe Pose, a machine-learning solution, for head–neck segmentation and applies color thresholding and morphological operations for marker detection. Validation tests demonstrated Pearson correlation coefficients above 0.98 compared to manual Kinovea measurements for both the craniovertebral angle (CVA) and cranial rotation angle (CRA), confirming high accuracy. Further validation on individuals with neck disorders showed similarly strong correlations, supporting clinical applicability. Speed comparison tests revealed that AutoMCA significantly reduces measurement time compared to traditional photogrammetry. Robustness tests confirmed reliable performance across varied backgrounds and marker types. In conclusion, AutoMCA measures head posture efficiency and lowers the requirements for instruments and space, making the assessment more versatile and applicable. Full article
(This article belongs to the Topic Intelligent Image Processing Technology)
Show Figures

Figure 1

21 pages, 4155 KB  
Article
Integrating Deep Learning and Radiogenomics: A Novel Approach to Glioblastoma Segmentation and MGMT Methylation Prediction
by Nabil M. Abdelaziz, Emad Abdel-Aziz Dawood and Alshaimaa A. Tantawy
J. Imaging 2025, 11(11), 403; https://doi.org/10.3390/jimaging11110403 - 11 Nov 2025
Cited by 1 | Viewed by 1643
Abstract
Radiogenomics, which integrates imaging phenotypes with genomic profiles, enhances diagnosis, prognosis, and treatment planning for glioblastomas. This study specifically establishes a correlation between radiomic features and MGMT promoter methylation status, advancing towards a non-invasive, integrated diagnostic paradigm. Conventional genetic analysis requires invasive biopsies, [...] Read more.
Radiogenomics, which integrates imaging phenotypes with genomic profiles, enhances diagnosis, prognosis, and treatment planning for glioblastomas. This study specifically establishes a correlation between radiomic features and MGMT promoter methylation status, advancing towards a non-invasive, integrated diagnostic paradigm. Conventional genetic analysis requires invasive biopsies, which cause delays in obtaining results and necessitate further surgeries. Our methodology is twofold: First, an enhanced U-Net model segments brain tumor regions with high precision (Dice coefficient: 0.889). Second, a hybrid classifier, leveraging the complementary features of EfficientNetB0 and ResNet50, predicts MGMT promoter methylation status from the segmented volumes. The proposed framework demonstrated superior performance in predicting MGMT promoter methylation status in glioblastoma patients compared to conventional methods, achieving a classification accuracy of 95% and an AUC of 0.96. These results underscore the model’s potential to enhance patient stratification and guide treatment selection. The accurate prediction of MGMT promoter methylation status via non-invasive imaging provides a reliable criterion for anticipating patient responsiveness to alkylating chemotherapy. This capability equips clinicians with a tool to inform personalized treatment strategies, optimizing therapeutic efficacy from the outset. Full article
(This article belongs to the Topic Intelligent Image Processing Technology)
Show Figures

Figure 1

19 pages, 3290 KB  
Article
Multi-Granularity Content-Aware Network with Semantic Integration for Unsupervised Anomaly Detection
by Xinyu Guo, Shihui Zhao, Jianbin Xue, Dongdong Liu, Xinyang Han, Shuai Zhang and Yufeng Zhang
Appl. Sci. 2025, 15(21), 11842; https://doi.org/10.3390/app152111842 - 6 Nov 2025
Cited by 1 | Viewed by 1295
Abstract
Unsupervised anomaly detection has been widely applied to industrial scenarios. Recently, transformer-based methods have also been developed and have produced good performance. Although the global dependencies in anomaly images are considered, the typical patch partition strategy in the vanilla self-attention mechanism ignores the [...] Read more.
Unsupervised anomaly detection has been widely applied to industrial scenarios. Recently, transformer-based methods have also been developed and have produced good performance. Although the global dependencies in anomaly images are considered, the typical patch partition strategy in the vanilla self-attention mechanism ignores the content consistencies in anomaly defects or normal regions. To sufficiently exploit the content consistency in images, we propose the multi-granularity content-aware network with semantic integration (MGCA-Net), in which superpixel segmentation is introduced into feature space to divide images according to their spatial structures. Specifically, we adopt a pre-trained ResNet as the encoder to extract features. Then, we design content-aware attention blocks (CAABs) to capture the global information in features at different granularities. In this block, we impose superpixel segmentation on the features from the encoder and employ the superpixels as tokens for the learning of global relationships. Because the superpixels are divided according to their content consistencies, the spatial structures of objects in anomaly or normal regions are preserved. Meanwhile, the multi-granularity semantic integration block is devised to further integrate the global information of all granularities. Next, we use semantic-guided fusion blocks (SGFBs) to progressively upsample the features with the help of CAABs. Finally, the differences between the outputs of CAABs and SGFBs are calculated and merged to predict the anomaly defects. Thanks to the preservation of content consistency of objects, experimental results on two benchmark datasets demonstrate that our proposed MGCA-Net achieves superior anomaly detection performance over state-of-the-art methods. Full article
(This article belongs to the Topic Intelligent Image Processing Technology)
Show Figures

Figure 1

17 pages, 8813 KB  
Article
A Fast Algorithm for Boundary Point Extraction of Planar Building Components from Point Clouds
by Yongzhong Huang, Ming Chen, Gaoming He and Jianming Liu
Electronics 2025, 14(21), 4313; https://doi.org/10.3390/electronics14214313 - 2 Nov 2025
Viewed by 1376
Abstract
The boundaries of planar building components characterize the structural outline of buildings and are of great importance in applications such as indoor model reconstruction and localization. However, traditional methods for extracting boundary points of planar building components from point clouds are often constrained [...] Read more.
The boundaries of planar building components characterize the structural outline of buildings and are of great importance in applications such as indoor model reconstruction and localization. However, traditional methods for extracting boundary points of planar building components from point clouds are often constrained by high computational complexity, limited efficiency, and insufficient accuracy. To address these challenges, this paper presents a rapid algorithm for the direct extraction of boundary points from point cloud data. The algorithm first performs planar fitting and projects all points onto the fitted plane to mitigate the influence of outlier noise. Next, for each point in the point cloud, a plane is constructed perpendicular to the existing plane points. Coarse boundary points are identified by counting the neighboring points on one side of this plane, which effectively eliminates most of the interior points. Finally, a boundary detection zone is defined for each coarse boundary point and its neighboring points. The precise boundary points are extracted by counting the number of points within this defined region. Experimental validation indicates that our algorithm can extract boundary points of planar building components from point clouds both accurately and efficiently, with notable robustness against noise. Full article
(This article belongs to the Topic Intelligent Image Processing Technology)
Show Figures

Figure 1

26 pages, 21316 KB  
Article
MultS-ORB: Multistage Oriented FAST and Rotated BRIEF
by Shaojie Zhang, Yinghui Wang, Jiaxing Ma, Jinlong Yang, Liangyi Huang and Xiaojuan Ning
Mathematics 2025, 13(13), 2189; https://doi.org/10.3390/math13132189 - 4 Jul 2025
Cited by 1 | Viewed by 1440
Abstract
Feature matching is crucial in image recognition. However, blurring caused by illumination changes often leads to deviations in local appearance-based similarity, resulting in ambiguous or false matches—an enduring challenge in computer vision. To address this issue, this paper proposes a method named MultS-ORB [...] Read more.
Feature matching is crucial in image recognition. However, blurring caused by illumination changes often leads to deviations in local appearance-based similarity, resulting in ambiguous or false matches—an enduring challenge in computer vision. To address this issue, this paper proposes a method named MultS-ORB (Multistage Oriented FAST and Rotated BRIEF). The proposed method preserves all the advantages of the traditional ORB algorithm while significantly improving feature matching accuracy under illumination-induced blurring. Specifically, it first generates initial feature matching pairs using KNN (K-Nearest Neighbors) based on descriptor similarity in the Hamming space. Then, by introducing a local motion smoothness constraint, GMS (Grid-Based Motion Statistics) is applied to filter and optimize the matches, effectively reducing the interference caused by blurring. Afterward, the PROSAC (Progressive Sampling Consensus) algorithm is employed to further eliminate false correspondences resulting from illumination changes. This multistage strategy yields more accurate and reliable feature matches. Experimental results demonstrate that for blurred images affected by illumination changes, the proposed method improves matching accuracy by an average of 75%, reduces average error by 33.06%, and decreases RMSE (Root Mean Square Error) by 35.86% compared to the traditional ORB algorithm. Full article
(This article belongs to the Topic Intelligent Image Processing Technology)
Show Figures

Figure 1

24 pages, 5858 KB  
Article
A YOLO11-Based Method for Segmenting Secondary Phases in Cu-Fe Alloy Microstructures
by Qingxiu Jing, Ruiyang Wu, Zhicong Zhang, Yong Li, Qiqi Chang, Weihui Liu and Xiaodong Huang
Information 2025, 16(7), 570; https://doi.org/10.3390/info16070570 - 3 Jul 2025
Cited by 3 | Viewed by 1873
Abstract
With the development of industrialization, the demand for high-performance metal materials has increased, and copper and its alloys have been widely used. The microstructure of these materials significantly affects their performance. To address the issues of subjectivity, low efficiency, and limited quantitative capability [...] Read more.
With the development of industrialization, the demand for high-performance metal materials has increased, and copper and its alloys have been widely used. The microstructure of these materials significantly affects their performance. To address the issues of subjectivity, low efficiency, and limited quantitative capability in traditional metallographic analysis methods, this paper proposes a deep learning-based approach for segmenting the second phase in Cu-Fe alloys. The method is built upon the YOLO11 framework and incorporates a series of structural enhancements tailored to the characteristics of the secondary-phase microstructure, aiming to improve the model’s detection accuracy and segmentation performance. Specifically, the EIEM module enhances the C3K2 structure to improve edge perception; the CSPSA module is optimized into C2CGA to strengthen multi-scale feature representation; and the RepGFPN and DySample techniques are integrated to construct the GDFPN neck network. Experimental results on the Cu-Fe alloy metallographic image dataset demonstrate that YOLO11 outperforms mainstream semantic segmentation models such as U-Net and DeepLabV3+ in terms of mAP (85.5%), inference speed (208 FPS), and model complexity (10.2 GFLOPs). The improved YOLO11 model achieves an mAP of 89.0%, a precision of 84.6%, and a recall of 81.0% on this dataset, showing significant performance improvements while effectively balancing inference speed and model complexity. Additionally, a quantitative analysis software system for secondary phase uniformity based on this model provides strong technical support for automated metallographic image analysis and demonstrates broad application prospects in materials science research and industrial quality control. Full article
(This article belongs to the Topic Intelligent Image Processing Technology)
Show Figures

Graphical abstract

14 pages, 2210 KB  
Article
AMFFNet: Adaptive Multi-Scale Feature Fusion Network for Urban Image Semantic Segmentation
by Shuting Huang and Haiyan Huang
Electronics 2025, 14(12), 2344; https://doi.org/10.3390/electronics14122344 - 8 Jun 2025
Cited by 4 | Viewed by 2012
Abstract
Urban image semantic segmentation faces challenges including the coexistence of multi-scale objects, blurred semantic relationships between complex structures, and dynamic occlusion interference. Existing methods often struggle to balance global contextual understanding of large scenes and fine-grained details of small objects due to insufficient [...] Read more.
Urban image semantic segmentation faces challenges including the coexistence of multi-scale objects, blurred semantic relationships between complex structures, and dynamic occlusion interference. Existing methods often struggle to balance global contextual understanding of large scenes and fine-grained details of small objects due to insufficient granularity in multi-scale feature extraction and rigid fusion strategies. To address these issues, this paper proposes an Adaptive Multi-scale Feature Fusion Network (AMFFNet). The network primarily consists of four modules: a Multi-scale Feature Extraction Module (MFEM), an Adaptive Fusion Module (AFM), an Efficient Channel Attention (ECA) module, and an auxiliary supervision head. Firstly, the MFEM utilizes multiple depthwise strip convolutions to capture features at various scales, effectively leveraging contextual information. Then, the AFM employs a dynamic weight assignment strategy to harmonize multi-level features, enhancing the network’s ability to model complex urban scene structures. Additionally, the ECA attention mechanism introduces cross-channel interactions and nonlinear transformations to mitigate the issue of small-object segmentation omissions. Finally, the auxiliary supervision head enables shallow features to directly affect the final segmentation results. Experimental evaluations on the CamVid and Cityscapes datasets demonstrate that the proposed network achieves superior mean Intersection over Union (mIoU) scores of 77.8% and 81.9%, respectively, outperforming existing methods. The results confirm that AMFFNet has a stronger ability to understand complex urban scenes. Full article
(This article belongs to the Topic Intelligent Image Processing Technology)
Show Figures

Figure 1

21 pages, 9110 KB  
Article
SwinTCS: A Swin Transformer Approach to Compressive Sensing with Non-Local Denoising
by Xiuying Li, Haoze Li, Hongwei Liao, Zhufeng Suo, Xuesong Chen and Jiameng Han
J. Imaging 2025, 11(5), 139; https://doi.org/10.3390/jimaging11050139 - 29 Apr 2025
Cited by 3 | Viewed by 1781
Abstract
In the era of the Internet of Things (IoT), the rapid growth of interconnected devices has intensified the demand for efficient data acquisition and processing techniques. Compressive Sensing (CS) has emerged as a promising approach for simultaneous signal acquisition and dimensionality reduction, particularly [...] Read more.
In the era of the Internet of Things (IoT), the rapid growth of interconnected devices has intensified the demand for efficient data acquisition and processing techniques. Compressive Sensing (CS) has emerged as a promising approach for simultaneous signal acquisition and dimensionality reduction, particularly in multimedia applications. In response to the challenges presented by traditional CS reconstruction methods, such as boundary artifacts and limited robustness, we propose a novel hierarchical deep learning framework, SwinTCS, for CS-aware image reconstruction. Leveraging the Swin Transformer architecture, SwinTCS integrates a hierarchical feature representation strategy to enhance global contextual modeling while maintaining computational efficiency. Moreover, to better capture local features of images, we introduce an auxiliary convolutional neural network (CNN). Additionally, for suppressing noise and improving reconstruction quality in high-compression scenarios, we incorporate a Non-Local Means Denoising module. The experimental results on multiple public benchmark datasets indicate that SwinTCS surpasses State-of-the-Art (SOTA) methods across various evaluation metrics, thereby confirming its superior performance. Full article
(This article belongs to the Topic Intelligent Image Processing Technology)
Show Figures

Figure 1

20 pages, 8414 KB  
Article
ADCNet: Anomaly-Driven Cross-Modal Contrastive Network for Medical Report Generation
by Yuxue Liu, Junsan Zhang, Kai Liu and Lizhuang Tan
Electronics 2025, 14(3), 532; https://doi.org/10.3390/electronics14030532 - 28 Jan 2025
Cited by 3 | Viewed by 2335
Abstract
Medical report generation has made significant progress in recent years. However, generated reports still suffer from issues such as poor readability, incomplete and inaccurate descriptions of lesions, and challenges in capturing fine-grained abnormalities. The primary obstacles include low image resolution, poor contrast, and [...] Read more.
Medical report generation has made significant progress in recent years. However, generated reports still suffer from issues such as poor readability, incomplete and inaccurate descriptions of lesions, and challenges in capturing fine-grained abnormalities. The primary obstacles include low image resolution, poor contrast, and substantial cross-modal discrepancies between visual and textual features. To address these challenges, we propose an Anomaly-Driven Cross-Modal Contrastive Network (ADCNet), which aims to enhance the quality and accuracy of medical report generation through effective cross-modal feature fusion and alignment. First, we design an anomaly-aware cross-modal feature fusion (ACFF) module that introduces an anomaly embedding vector to guide the extraction and generation of anomaly-related features from visual representations. This process enhances the capability of visual features to capture lesion-related abnormalities and improves the performance of feature fusion. Second, we propose a fine-grained regional feature alignment (FRFA) module, which dynamically filters visual and textual features to suppress irrelevant information and background noise. This module computes cross-modal relevance to align fine-grained regional features, ensuring improved semantic consistency between images and generated reports. The experimental results from the IU X-Ray and MIMIC-CXR datasets demonstrate that the proposed ADCNet method significantly outperforms existing approaches. Specifically, ADCNet achieves notable improvements in natural language generation metrics, as well as the accuracy, completeness, and fluency of medical report generation. Full article
(This article belongs to the Topic Intelligent Image Processing Technology)
Show Figures

Figure 1

22 pages, 28158 KB  
Article
Edge-Aware Dual-Task Image Watermarking Against Social Network Noise
by Hao Jiang, Jiahao Wang, Yuhan Yao, Xingchen Li, Feifei Kou, Xinkun Tang and Limei Qi
Appl. Sci. 2025, 15(1), 57; https://doi.org/10.3390/app15010057 - 25 Dec 2024
Viewed by 2592
Abstract
In the era of widespread digital image sharing on social media platforms, deep-learning-based watermarking has shown great potential in copyright protection. To address the fundamental trade-off between the visual quality of the watermarked image and the robustness of watermark extraction, we explore the [...] Read more.
In the era of widespread digital image sharing on social media platforms, deep-learning-based watermarking has shown great potential in copyright protection. To address the fundamental trade-off between the visual quality of the watermarked image and the robustness of watermark extraction, we explore the role of structural features and propose a novel edge-aware watermarking framework. Our primary innovation lies in the edge-aware secret hiding module (EASHM), which achieves adaptive watermark embedding by aligning watermarks with image structural features. To realize this, the EASHM leverages knowledge distillation from an edge detection teacher and employs a dual-task encoder that simultaneously performs edge detection and watermark embedding through maximal parameter sharing. The framework is further equipped with a social network noise simulator (SNNS) and a secret recovery module (SRM) to enhance robustness against common image noise attacks. Extensive experiments on three public datasets demonstrate that our framework achieves superior watermark imperceptibility, with PSNR and SSIM values exceeding 40.82 dB and 0.9867, respectively, while maintaining an over 99% decoding accuracy under various noise attacks, outperforming existing methods by significant margins. Full article
(This article belongs to the Topic Intelligent Image Processing Technology)
Show Figures

Figure 1

15 pages, 2181 KB  
Article
Micro-Expression Recognition Algorithm Using Regions of Interest and the Weighted ArcFace Loss
by Peiying Zhang, Ruixin Wang, Jia Luo and Lei Shi
Electronics 2025, 14(1), 2; https://doi.org/10.3390/electronics14010002 - 24 Dec 2024
Cited by 3 | Viewed by 3345
Abstract
Micro-expressions often reveal more genuine emotions but are challenging to recognize due to their brief duration and subtle amplitudes. To address these challenges, this paper introduces a micro-expression recognition method leveraging regions of interest (ROIs). Firstly, four specific ROIs are selected based on [...] Read more.
Micro-expressions often reveal more genuine emotions but are challenging to recognize due to their brief duration and subtle amplitudes. To address these challenges, this paper introduces a micro-expression recognition method leveraging regions of interest (ROIs). Firstly, four specific ROIs are selected based on an analysis of the optical flow and relevant action units activated during micro-expressions. Secondly, effective feature extraction is achieved using the optical flow method. Thirdly, a block partition module is integrated into a convolutional neural network to reduce computational complexity, thereby enhancing model accuracy and generalization. The proposed model achieves notable performance, with accuracies of 93.96%, 86.15%, and 81.17% for three-class recognition on the CASME II, SAMM, and SMIC datasets, respectively. For five-class recognition, the model achieves accuracies of 81.63% on the CASME II dataset and 84.31% on the SMIC dataset. Experimental results validate the effectiveness of using ROIs in improving micro-expression recognition accuracy. Full article
(This article belongs to the Topic Intelligent Image Processing Technology)
Show Figures

Figure 1

Back to TopTop