Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (63)

Search Parameters:
Keywords = video denoising

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
15 pages, 22895 KB  
Article
Stable and High-Throughput Single-Cell Sorting of Food Bacteria Using Spatiotemporal Video-Enhanced Raman Tweezers
by Yi Sun, Zhipeng Li, Hua Xia, Kaier Yang, Feng Gao, Yingxiao Peng, Xiangyun Ma and Qifeng Li
Foods 2026, 15(12), 2208; https://doi.org/10.3390/foods15122208 - 18 Jun 2026
Viewed by 308
Abstract
Rapid detection of foodborne pathogenic and spoilage microorganisms is critical for ensuring food safety and quality in liquid matrices. While Raman tweezers spectroscopy (RTS) enables label-free single-cell analysis, its application in high-throughput inline inspection faces a fundamental bottleneck: high flow rates required for [...] Read more.
Rapid detection of foodborne pathogenic and spoilage microorganisms is critical for ensuring food safety and quality in liquid matrices. While Raman tweezers spectroscopy (RTS) enables label-free single-cell analysis, its application in high-throughput inline inspection faces a fundamental bottleneck: high flow rates required for efficiency induce severe motion blur and low signal-to-noise ratios (SNR), which blind automated control systems and destabilize optical trapping. To overcome this, we present a Spatiotemporal Video-Enhanced Raman Tweezers (SVERT) system integrating a deceleration-optimized microfluidic chip with a deep learning-based visual feedback loop. We propose a Local–Global Unified Denoising Network (LGU-Net) tailored to recover high-fidelity bacterial structures from low-SNR video streams, achieving a deterministic processing latency of ~0.49 ms. Experimental results demonstrate that SVERT improves the optical trapping success rate from 21.27% ± 2% to 91.47% ± 1.8% compared to raw video input, enabling a four-fold increase in spectral acquisition efficiency. Leveraging the acquired high-quality dataset, we achieved a classification accuracy of 96.74% across four bacterial species of relevance to food safety and quality. Crucially, we validated the system’s practical robustness by successfully isolating and tracking trace E. coli in an unpurified commercial beverage. This capability to effectively mitigate natural background interference demonstrates the system’s promising potential to be expanded for broader applications in liquid food safety screening. Full article
Show Figures

Graphical abstract

17 pages, 13123 KB  
Article
2s-DAS: Two-Stream Diffusion with Multi-Modal Fusion for Temporal Action Segmentation
by Ce Li, Xuli Guo, Ruijie Wang, Kaipan Zhao, Linlin Yang and Fang Wan
J. Imaging 2026, 12(6), 237; https://doi.org/10.3390/jimaging12060237 - 28 May 2026
Viewed by 426
Abstract
Human temporal action segmentation (TAS) is a fundamental video understanding task aimed at partitioning untrimmed videos into semantically coherent action segments. While temporal convolutional networks and transformers have significantly improved frame representation and temporal modeling, existing methods are still constrained by two critical [...] Read more.
Human temporal action segmentation (TAS) is a fundamental video understanding task aimed at partitioning untrimmed videos into semantically coherent action segments. While temporal convolutional networks and transformers have significantly improved frame representation and temporal modeling, existing methods are still constrained by two critical limitations: the dependence on single-modal inputs and the inefficiency of iterative, frame-wise sequential modeling. To address these gaps, we propose 2s-DAS: a novel two-stream diffusion-based framework for action segmentation characterized by three key contributions. First, we introduce a multi-modal frame representation that integrates optical flow with Br-Prompt RGB features, thereby capturing richer spatial-temporal context and enhancing feature representation. Second, we leverage a diffusion model to perform sequence segmentation, utilizing importance sampling to prioritize key frames for segment-level temporal modeling. Concurrently, a refinement mechanism based on iterative decoding denoising is introduced to ensure fine-grained action prediction. Third, we design a two-stream fusion mechanism that processes the streams of RGB with text and optical flow separately and integrates multi-modal information by a late fusion strategy to explicitly reduce oversegmentation. Evaluation experiments on GTEA, 50Salads, and Breakfast datasets show that our 2s-DAS significantly outperforms state-of-the-art methods, setting new benchmarks while effectively addressing the over-segmentation issue. Full article
Show Figures

Figure 1

21 pages, 66333 KB  
Review
Diffusion Models: Unlocking the “4 Secrets” of High-Quality Image Generation
by Tao Zhou, Zhe Zhang, Mingzhe Zhang, Wenwen Chai, Yong Xia and Fuyuan Hu
Electronics 2026, 15(8), 1755; https://doi.org/10.3390/electronics15081755 - 21 Apr 2026
Viewed by 2523
Abstract
The diffusion model (DM) is a hot topic in deep generative models and is widely applied in image generation. In diffusion models, there are four main “secrets” that affect high-quality image generation: constructing the diffusion model, improving the sampling velocity, designing the diffusion [...] Read more.
The diffusion model (DM) is a hot topic in deep generative models and is widely applied in image generation. In diffusion models, there are four main “secrets” that affect high-quality image generation: constructing the diffusion model, improving the sampling velocity, designing the diffusion process, and guiding diffusion models. How should one construct the diffusion model? How can one improve the sampling velocity? How should one design the diffusion process? How should one guide diffusion models? These questions are critical to enhancing diffusion model performance. However, most existing review papers focus on applications, while discussion of the four key technical aspects remains limited. In response, this paper summarizes four key technologies and six representative application directions. First, the basic principles of diffusion models are reviewed from three perspectives: denoising diffusion probabilistic models, noise conditional score network models, and stochastic differential equation models. Second, key techniques for improving sampling velocity are summarized from three perspectives: non-Markovian sampling, knowledge distillation sampling, and discrete optimization sampling. Third, the diffusion process design is summarized from three perspectives: latent space, Transformer-based diffusion, and non-Euclidean space. Fourth, guidance strategies are summarized from three perspectives: classifier guidance, classifier-free guidance, and multimodal guidance. Fifth, the advantages and applications of diffusion models are discussed in high-quality text-to-image generation, high-quality text-to-video generation, and high-quality image-to-image generation. Finally, this paper discusses the challenges faced by diffusion models in image generation. Overall, this review systematically discusses the four “secrets” of diffusion models for image generation and provides a useful reference for future research in this field. Full article
Show Figures

Graphical abstract

26 pages, 4138 KB  
Article
Self-Supervised Cascade Denoising Auto-Encoder for Accurate Spatial Positioning of Target by Fusing Uncalibrated Video and Low-Cost GNSS
by Xiaofei Zeng, Ruliang He, Songchen Han, Wei Li, Menglong Yang and Binbin Liang
Remote Sens. 2026, 18(8), 1161; https://doi.org/10.3390/rs18081161 - 13 Apr 2026
Cited by 1 | Viewed by 613
Abstract
Accurate measurement of the spatial position of targets in a fixed camera is critical in remote sensing applications. Visual spatial positioning methods that rely solely on images are susceptible to adverse factors such as inaccurate camera calibration, imprecise image target detection, and incorrect [...] Read more.
Accurate measurement of the spatial position of targets in a fixed camera is critical in remote sensing applications. Visual spatial positioning methods that rely solely on images are susceptible to adverse factors such as inaccurate camera calibration, imprecise image target detection, and incorrect feature point selection. Complementary to images, the ubiquitous Global Navigation Satellite System (GNSS) data can provide spatial positions of targets, but most of them are low-cost GNSSs with significant positioning noise. In order to fuse these two valuable but flawed positioning measurements to improve the accuracy and stability of spatial positioning, we propose a deep learning multi-modal spatial positioning method by fusing sequential uncalibrated video images and low-cost GNSSs. Firstly, a self-supervised cascade denoising auto-encoder (SCDAE) architecture is built to endow the auto-encoder with robustness to noise in the raw inputs. Then, based on the SCDAE and Bayesian optimal estimation, a Bayesian self-supervised multi-modal fusion positioning method SCDAE-MFP is presented to achieve accurate and stable spatial positioning by self-supervised manifold learning. Specifically, to provide visual self-supervision to the SCDAE-MFP, a visual position denoising auto-encoder module based on dual unsupervised learning is proposed. Extensive experimental results on public datasets showed that SCDAE-MFP outperformed five other classical and state-of-the-art baseline methods by an average of 56.79% in reducing positioning errors. Full article
(This article belongs to the Special Issue GNSS and Multi-Sensor Integrated Precise Positioning and Applications)
Show Figures

Figure 1

19 pages, 10157 KB  
Article
DiffVP: A Diffusion Model with Explicit Coordinate-Temporal Encoding for Viewport Prediction in 360 Videos
by Huimin Zheng, Lina Du, Xiushan Nie and Fei Dong
Electronics 2026, 15(6), 1326; https://doi.org/10.3390/electronics15061326 - 23 Mar 2026
Viewed by 584
Abstract
Viewport prediction is a key component in tile-based 360° video streaming. Existing viewport prediction models based on Long Short-term Memory Networks (LSTM) or Transformer typically output a single deterministic future trajectory through deterministic mapping, which fails to capture the inherent randomness in viewing [...] Read more.
Viewport prediction is a key component in tile-based 360° video streaming. Existing viewport prediction models based on Long Short-term Memory Networks (LSTM) or Transformer typically output a single deterministic future trajectory through deterministic mapping, which fails to capture the inherent randomness in viewing behavior. Moreover, when encoding trajectory features, such models often map trajectory coordinates directly into a high-dimensional space while neglecting the spatial information inherent in the coordinates themselves. Additionally, they exhibit limitations in capturing cross-modal relationships between visual and trajectory features. To address these issues, this paper proposes DiffVP, a diffusion model for viewport prediction in 360° videos. Under the constraints of viewing historical trajectories and video saliency maps, DiffVP leverages Denoising Diffusion Implicit Models (DDIMs) to model future viewing trajectories in the form of probability distributions, generating diverse and reasonable prediction results. In the denoising network, DiffVP employs Explicit Coordinate-Time Encoding (ECTE) to model the temporal dependencies of trajectories and the spatial relationships among coordinates; moreover, a Coordinate-Aware Saliency Features Fusion (CASF) module is proposed to achieve cross-modal alignment and interactive fusion of saliency and trajectory features. Experimental results on three public datasets demonstrate that DiffVP achieves the best accuracy for 2–5 s viewport prediction without sacrificing the performance of short-term (<1 s) prediction. Full article
Show Figures

Figure 1

45 pages, 2842 KB  
Article
A Taxonomy of Generative Models with a Focus on Diffusion Models and Denoising Techniques
by Aditi Singh, Nikhil Kumar Chatta, Yuvaraj Vagula, Abul Ehtesham, Saket Kumar and Tala Talaei Khoei
Electronics 2026, 15(6), 1293; https://doi.org/10.3390/electronics15061293 - 19 Mar 2026
Viewed by 1883
Abstract
Diffusion models have emerged as a powerful class of generative models, demonstrating impressive results across visual domains such as image and video synthesis. This survey provides a comprehensive taxonomy of generative models, with a particular focus on diffusion models and their applications in [...] Read more.
Diffusion models have emerged as a powerful class of generative models, demonstrating impressive results across visual domains such as image and video synthesis. This survey provides a comprehensive taxonomy of generative models, with a particular focus on diffusion models and their applications in enhancing visual fidelity for text-to-image and text-to-video generation. We discuss the theoretical foundations of diffusion models, including their formulation through stochastic differential equations, and analyze the forward noising and reverse denoising processes that enable stable training and high-quality generation. The survey further categorizes diffusion architectures, including pixel-space and latent-space models, and examines their design choices, training strategies, and trade-offs across different resolution regimes. In addition, we review noise characteristics in real-world imaging domains and discuss their implications for diffusion-based models. Denoising strategies are analyzed by distinguishing between in-model denoising mechanisms and external denoising techniques used in preprocessing and post-processing pipelines. The survey also summarizes commonly used datasets and evaluation metrics for generative modeling, providing a practical perspective on benchmarking and model comparison. Finally, we discuss current challenges, including computational efficiency, scalability, and robustness to diverse noise distributions, and outline potential directions for future research. This survey aims to provide a structured reference for understanding diffusion models and their applications in visual generation tasks. Full article
(This article belongs to the Special Issue Autonomous Intelligence: Concepts and Applications of Agentic AI)
Show Figures

Figure 1

24 pages, 879 KB  
Review
A Survey of Diffusion Models: Methods and Applications
by HaoYu Ma and Hon-Cheng Wong
Appl. Sci. 2026, 16(5), 2482; https://doi.org/10.3390/app16052482 - 4 Mar 2026
Cited by 3 | Viewed by 4762
Abstract
Diffusion models have emerged as the state-of-the-art generative paradigm, surpassing GANs in synthesizing high-fidelity images, videos, and audio. However, their reliance on iterative denoising processes imposes substantial computational burdens and memory overheads, creating a significant barrier to their deployment on resource-constrained edge devices. [...] Read more.
Diffusion models have emerged as the state-of-the-art generative paradigm, surpassing GANs in synthesizing high-fidelity images, videos, and audio. However, their reliance on iterative denoising processes imposes substantial computational burdens and memory overheads, creating a significant barrier to their deployment on resource-constrained edge devices. Unlike existing surveys that broadly cover general methodologies, this paper provides a focused review with a specific emphasis on efficient and lightweight diffusion models. We systematically analyze the trade-offs between generation quality and computational cost, categorizing acceleration techniques into sampling optimization, architectural compression, and knowledge distillation. Furthermore, we explore the integration of diffusion models with emerging architectures (e.g., Mamba) and their evolution towards general-purpose world simulators. This survey aims to provide a roadmap for “Green AI,” bridging the gap between high-end academic research and practical, real-world applications. Full article
(This article belongs to the Section Computing and Artificial Intelligence)
Show Figures

Figure 1

26 pages, 30049 KB  
Article
HVIFormer: A Dual-Stage Low-Light Image Enhancement Method Based on HVI Representation
by Yimei Li, Liuhong Luo and Hongjun Li
Appl. Sci. 2026, 16(5), 2450; https://doi.org/10.3390/app16052450 - 3 Mar 2026
Viewed by 1034
Abstract
Low-light image enhancement improves the quality of video surveillance and image analysis and, as a result, has long been a hot topic in image processing. However, current research on this topic faces a difficult challenge—effectively suppressing noise while improving brightness and maintaining color [...] Read more.
Low-light image enhancement improves the quality of video surveillance and image analysis and, as a result, has long been a hot topic in image processing. However, current research on this topic faces a difficult challenge—effectively suppressing noise while improving brightness and maintaining color consistency, especially in extremely dark scenes, where dark noise amplification, uneven exposure, and color shifts often interact, leading to detail loss and color distortion. To address the issue, we propose a dual-stage low-light enhancement framework based on the HVI (Horizontal/Vertical-Intensity) color space. The low-light image is first mapped to the HVI space, obtaining the intensity component I and the HVI-based feature map, with I being explicitly extracted as an intensity prior. A Transformer-based pre-recovery module is introduced for global dependency modeling, guided by the intensity prior I through an Intensity-Conditioned Block (ICB) for conditional feature interaction. Subsequently, a dual-branch enhancement network utilizes lightweight Complementary Cross-Attention (CCA) blocks for brightness refinement and color denoising. Finally, the enhanced image is remapped to the sRGB color space. The proposed framework decouples global brightness recovery and feature preprocessing from detail enhancement and color refinement, improving stability in extremely dark and high-noise scenarios. Through 18 quantitative and qualitative experiments, we demonstrate that our proposed method achieves superior performance in dark noise suppression and color restoration across multiple low-light datasets. Full article
Show Figures

Figure 1

19 pages, 3156 KB  
Article
Detecting Escherichia coli on Conventional Food Processing Surfaces Using UV-C Fluorescence Imaging and Deep Learning
by Zafar Iqbal, Thomas F. Burks, Snehit Vaddi, Pappu Kumar Yadav, Quentin Frederick, Satya Aakash Chowdary Obellaneni, Jianwei Qin, Moon Kim, Mark A. Ritenour, Jiuxu Zhang and Fartash Vasefi
Appl. Sci. 2026, 16(2), 968; https://doi.org/10.3390/app16020968 - 17 Jan 2026
Viewed by 957
Abstract
Detecting Escherichia coli on food preparation and processing surfaces is critical for ensuring food safety and preventing foodborne illness. This study focuses on detecting E. coli contamination on common food processing surfaces using UV-C fluorescence imaging and deep learning. Four concentrations of E. [...] Read more.
Detecting Escherichia coli on food preparation and processing surfaces is critical for ensuring food safety and preventing foodborne illness. This study focuses on detecting E. coli contamination on common food processing surfaces using UV-C fluorescence imaging and deep learning. Four concentrations of E. coli (0, 105, 107, and 108 colony forming units (CFU)/mL) and two egg solutions (white and yolk) were applied to stainless steel and white rubber to simulate realistic contamination with organic interference. For each concentration level, 256 droplets were inoculated in 16 groups, and fluorescence videos were captured. Droplet regions were extracted from the video frames, subdivided into quadrants, and augmented to generate a robust dataset, ensuring 3–4 droplets per sample. Wavelet-based denoising further improved image quality, with Haar wavelets producing the highest Peak Signal-to-Noise Ratio (PSNR) values, up to 51.0 dB on white rubber and 48.2 dB on stainless steel. Using this dataset, multiple deep learning (DL) models, including ConvNeXtBase, EfficientNetV2L, and five YOLO11-cls variants, were trained to classify E. coli concentration levels. Additionally, Eigen-CAM heatmaps were used to visualize model attention to bacterial fluorescence regions. Across four dataset groupings, YOLO11-cls models achieved consistently high performance, with peak test accuracies of 100% on white rubber and 99.60% on stainless steel, even in the presence of egg substances. YOLO11s-cls provided the best balance of accuracy (up to 98.88%) and inference speed (4–5 ms) whilst having a compact size (11 MB), outperforming larger models such as EfficientNetV2L. Classical machine learning models lagged significantly behind, with Random Forest reaching 89.65% accuracy and SVM only 67.62%. Overall, the results highlight the potential of combining UV-C fluorescence imaging with deep learning for rapid and reliable detection of E. coli on stainless steel and rubber conveyor belt surfaces. Additionally, this approach could support the design of effective interventions to remove E. coli from food processing environments. Full article
Show Figures

Figure 1

25 pages, 8224 KB  
Article
QWR-Dec-Net: A Quaternion-Wavelet Retinex Framework for Low-Light Image Enhancement with Applications to Remote Sensing
by Vladimir Frants, Sos Agaian, Karen Panetta and Artyom Grigoryan
Information 2026, 17(1), 89; https://doi.org/10.3390/info17010089 - 14 Jan 2026
Cited by 1 | Viewed by 1194
Abstract
Computer vision and deep learning are essential in diverse fields such as autonomous driving, medical imaging, face recognition, and object detection. However, enhancing low-light remote sensing images remains challenging for both research and real-world applications. Low illumination degrades image quality due to sensor [...] Read more.
Computer vision and deep learning are essential in diverse fields such as autonomous driving, medical imaging, face recognition, and object detection. However, enhancing low-light remote sensing images remains challenging for both research and real-world applications. Low illumination degrades image quality due to sensor limitations and environmental factors, weakening visual fidelity and reducing performance in vision tasks. Common issues such as insufficient lighting, backlighting, and limited exposure create low contrast, heavy shadows, and poor visibility, particularly at night. We propose QWR-Dec-Net, a quaternion-based Retinex decomposition network tailored for low-light image enhancement. QWR-Dec-Net consists of two key modules: a decomposition module that separates illumination and reflectance, and a denoising module that fuses a quaternion holistic color representation with wavelet multi-frequency information. This structure jointly improves color constancy and noise suppression. Experiments on low-light remote sensing datasets (LSCIDMR and UCMerced) show that QWR-Dec-Net outperforms current methods in PSNR, SSIM, LPIPS, and classification accuracy. The model’s accurate illumination estimation and stable reflectance make it well-suited for remote sensing tasks such as object detection, video surveillance, precision agriculture, and autonomous navigation. Full article
(This article belongs to the Section Artificial Intelligence)
Show Figures

Figure 1

15 pages, 3599 KB  
Article
High-Fidelity rPPG Waveform Reconstruction from Palm Videos Using GANs
by Tao Li and Yuliang Liu
Sensors 2026, 26(2), 563; https://doi.org/10.3390/s26020563 - 14 Jan 2026
Cited by 1 | Viewed by 1362
Abstract
Remote photoplethysmography (rPPG) enables non-contact acquisition of human physiological parameters using ordinary cameras, and has been widely applied in medical monitoring, human–computer interaction, and health management. However, most existing studies focus on estimating specific physiological metrics, such as heart rate and heart rate [...] Read more.
Remote photoplethysmography (rPPG) enables non-contact acquisition of human physiological parameters using ordinary cameras, and has been widely applied in medical monitoring, human–computer interaction, and health management. However, most existing studies focus on estimating specific physiological metrics, such as heart rate and heart rate variability, while paying insufficient attention to reconstructing the underlying rPPG waveform. In addition, publicly available datasets typically record facial videos accompanied by fingertip PPG signals as reference labels. Since fingertip PPG waveforms differ substantially from the true photoplethysmography (PPG) signals obtained from the face, deep learning models trained on such datasets often struggle to recover high-quality rPPG waveforms. To address this issue, we collected a new dataset consisting of palm-region videos paired with wrist-based PPG signals as reference labels, and experimentally validated its effectiveness for training neural network models aimed at rPPG waveform reconstruction. Furthermore, we propose a generative adversarial network (GAN)-based pulse-wave synthesis framework that produces high-quality rPPG waveforms by denoising the mean green-channel signal. By incorporating time-domain peak-aware loss, frequency-domain loss, and adversarial loss, our method achieves promising performance, with an RMSE (Root Mean Square Error) of 0.102, an MAPE (Mean Absolute Percentage Error) of 0.028, a Pearson correlation of 0.987, and a cosine similarity of 0.989. These results demonstrate the capability of the proposed approach to reconstruct high-fidelity rPPG waveforms with improved morphological accuracy compared to noisy raw rPPG signals, rather than directly validating health monitoring performance. This study presents a high-quality rPPG waveform reconstruction approach from both data and model perspectives, providing a reliable foundation for subsequent physiological signal analysis, waveform-based studies, and potential health-related applications. Full article
(This article belongs to the Special Issue Systems for Contactless Monitoring of Vital Signs)
Show Figures

Figure 1

17 pages, 558 KB  
Article
FPGA-Accelerated Multi-Resolution Spline Reconstruction for Real-Time Multimedia Signal Processing
by Manuel J. C. S. Reis
Electronics 2026, 15(1), 173; https://doi.org/10.3390/electronics15010173 - 30 Dec 2025
Cited by 1 | Viewed by 1547
Abstract
This paper presents an FPGA-based architecture for real-time spline-based signal reconstruction, targeted at multimedia signal processing applications. Leveraging the multi-resolution properties of B-splines, the proposed design enables efficient upsampling, denoising, and feature preservation for image and video signals. Implemented on a mid-range FPGA, [...] Read more.
This paper presents an FPGA-based architecture for real-time spline-based signal reconstruction, targeted at multimedia signal processing applications. Leveraging the multi-resolution properties of B-splines, the proposed design enables efficient upsampling, denoising, and feature preservation for image and video signals. Implemented on a mid-range FPGA, the system supports parallel processing of multiple channels, with low-latency memory access and pipelined arithmetic units. The proposed pipeline achieves a throughput of up to 33.1 megasmples per second for 1D signals and 19.4 megapixels per second for 2D images, while maintaining average power consumption below 250 mW. Compared to CPU and embedded GPU implementations, the design delivers >15× improvement in energy efficiency and deterministic low-latency performance (8–12 clock cycles). A key novelty lies in combining multi-resolution B-spline reconstruction with fixed-point arithmetic and streaming-friendly pipelining, making the architecture modular, compact, and robust to varying input rates. Benchmarking results on synthetic and real multimedia datasets show significant improvements in throughput and energy efficiency compared to conventional CPU and GPU implementations. The architecture supports flexible resolution scaling, making it suitable for edge-computing scenarios in multimedia environments. Full article
(This article belongs to the Special Issue Digital Signal and Image Processing for Multimedia Technology)
Show Figures

Figure 1

19 pages, 5393 KB  
Article
Mine Water Hazard Video Recognition Based on Residual Preprocessing and Temporal–Spatial Descriptors
by Shuai Zhang, Haining Wang, Yuanze Du, Xinrui Li, Hongrui Luo and Yingwang Zhao
Appl. Sci. 2026, 16(1), 265; https://doi.org/10.3390/app16010265 - 26 Dec 2025
Viewed by 532
Abstract
Traditional water hazard monitoring often relies on manual inspection and water level sensors, typically lacking in accuracy and real-time capabilities. However, the method of using video surveillance for monitoring water hazard characteristics can compensate for these shortcomings. Therefore, this study proposes a method [...] Read more.
Traditional water hazard monitoring often relies on manual inspection and water level sensors, typically lacking in accuracy and real-time capabilities. However, the method of using video surveillance for monitoring water hazard characteristics can compensate for these shortcomings. Therefore, this study proposes a method to detect water hazards in mines using video recognition technology, combining temporal and spatial descriptors to enhance recognition accuracy. This study employs residual preprocessing technology to effectively eliminate complex underground static backgrounds, focusing solely on dynamic water flow features, thereby addressing the issue of the absence of water inrush samples. The method involves analyzing dynamic water flow pixels and applying an iterative denoising algorithm to successfully remove discrete noise points while preserving connected water flow areas. Experimental results show that this method achieves a detection accuracy of 90.68% for gushing water, significantly surpassing methods that rely solely on temporal or spatial descriptors. Moreover, this method not only focuses on the temporal characteristics of water flow but also addresses the challenge of detection difficulties due to the lack of historical gushing water samples. This research provides an effective technical solution and new insights for future water gushing monitoring in mines. Full article
Show Figures

Figure 1

23 pages, 3710 KB  
Article
Multi-Domain Intelligent State Estimation Network for Highly Maneuvering Target Tracking with Non-Gaussian Noise
by Zhenzhen Ma, Xueying Wang, Yuan Huang, Qingyu Xu, Wei An and Weidong Sheng
Remote Sens. 2025, 17(24), 4016; https://doi.org/10.3390/rs17244016 - 12 Dec 2025
Viewed by 914
Abstract
In the field of remote sensing, tracking highly maneuvering targets is challenging due to its rapidly changing patterns and uncertainties, particularly under non-Gaussian noise conditions. In this paper, we consider the problem of tracking highly maneuvering targets without using preset parameters in non-Gaussian [...] Read more.
In the field of remote sensing, tracking highly maneuvering targets is challenging due to its rapidly changing patterns and uncertainties, particularly under non-Gaussian noise conditions. In this paper, we consider the problem of tracking highly maneuvering targets without using preset parameters in non-Gaussian noise. We propose a multi-domain intelligent state estimation network (MIENet). It consists of two main models to estimate the key parameter for the Unscented Kalman Filter, enabling robust tracking of highly maneuvering targets under various intensities and distributions of observation noise. The first model, called a fusion denoising model (FDM), is designed to eliminate observation noise by enhancing multi-domain feature fusion. The second model, called a parameter estimation model (PEM), is designed to estimate key parameters of target motion by learning both global and local motion information. Additionally, we design a physically constrained loss function (PCLoss) that incorporates physics-informed constraints and prior knowledge. We evaluate our method on radar trajectory simulation and real remote sensing video datasets. Simulation results on the LAST dataset demonstrate that the proposed FDM can reduce the root mean square error (RMSE) of observation noise by more than 60%. Moreover, the proposed MIENet consistently outperforms the state-of-the-art state estimation algorithms across various highly maneuvering scenes, achieving this performance without requiring adjustment of noise parameters under non-Gaussian noise. Furthermore, experiments conducted on the real-world SV248S dataset confirm that MIENet effectively generalizes to satellite video object tracking tasks. Full article
(This article belongs to the Section AI Remote Sensing)
Show Figures

Figure 1

25 pages, 8373 KB  
Article
Performance Improvement of Vehicle and Human Localization and Classification by YOLO Family Networks in Noisy UAV Images
by Viktor Makarichev, Rostyslav Tsekhmystro, Vladimir Lukin and Dmytro Krytskyi
Information 2025, 16(12), 1087; https://doi.org/10.3390/info16121087 - 7 Dec 2025
Cited by 1 | Viewed by 877
Abstract
Many important tasks in smart city development and management are solved by systems of monitoring and control installed on-board of unmanned aerial vehicles (UAVs). UAV sensors can be imperfect or they can operate in unfavorable conditions, which can then result in obtaining images [...] Read more.
Many important tasks in smart city development and management are solved by systems of monitoring and control installed on-board of unmanned aerial vehicles (UAVs). UAV sensors can be imperfect or they can operate in unfavorable conditions, which can then result in obtaining images or video sequences that are noisy. Noise can degrade the performance of methods of vehicle and human localization and classification. Therefore, specific techniques to improve performance have to be applied. In this paper, we consider YOLO family neural networks as tools for solving the aforementioned tasks. This family of networks is rapidly developing; however, the input data may still require pre-processing. One option is to apply denoising before object localization and classification. In addition, approaches based on augmentation and training can be used as well. We consider the performance of these approaches for various noise intensities. We identify the noise levels at which network performance starts to degrade and analyze possibilities of performance improvement for two filters–BM3D and DRUNet. Both improve such performance criteria as the F1 score, the Intersection over Union and the mean Average Precision. Datasets of urban areas are used in the network training and verification. Full article
(This article belongs to the Special Issue Artificial Intelligence and Data Science for Smart Cities)
Show Figures

Graphical abstract

Back to TopTop