1. Introduction
With the development of space programs worldwide, human beings have progressively expanded their activities into the field of space. The increasing number of satellites and spacecraft launched into orbit has exacerbated the problem of space debris. These uncontrolled space debris [
1] in deep space pose a threat to the normal operation of satellites in orbit, making the cataloging and detection of such debris critically important.
The use of optical telescopes to observe space objects offers advantages such as low operating costs, long detection range, high measurement accuracy, and strong concealment [
2]. The operational modes of optical telescopes can be categorized into two types: the first is the ‘stellar-gazing mode’, in which the telescope maintains a fixed gaze on a stationary star; the second is the ‘target-tracking mode’, in which the telescope requires continual readjustment to sustain its gaze on a spatial target. The specific differences between these two modes are illustrated in
Figure 1:
When optical telescopes employ long exposure times across different observation modes, they accumulate signals from star points. However, relative motion between these star points and the detector’s imaging plane causes elongated “trailing” effects in the star images [
3]. This disperses the star point energy across more pixels, resulting in blurred or fragmented point representations. Consequently, the accuracy of star centroid extraction is compromised, hindering precise star image recognition in subsequent processes.
Image deblurring can be generally categorized into traditional methods and deep learning-based methods. Traditional methods typically infer the relationship between the blur kernel and the sharp image in an inverse manner via deconvolution algorithms. Depending on whether the blur kernel is known, they can be further divided into non-blind and blind image deblurring methods. Non-blind image deblurring methods [
4,
5,
6] solve for the underlying sharp image using deconvolution techniques when the blur kernel is known. In contrast, blind image deblurring methods [
7,
8,
9,
10,
11] first estimate the blur kernel through prior modeling approaches when it is unknown, then derive the solution. Common methods for blurred image restoration include inverse filtering [
12], Wiener filtering [
13], regularized filtering [
14], and Lucy-Richardson filtering (L-R) [
15,
16]. Deep learning-based methods [
17,
18,
19,
20,
21,
22] do not require manual preconfiguration of various image priors and do not need to estimate the blur kernel. Leveraging the powerful feature extraction capability of convolutional neural networks, they establish a mapping from blurred images to sharp images, enabling an end-to-end image deblurring process.
Under long exposures or in high-dynamic-range conditions, star points in images exhibit trailing phenomena, i.e., motion blur [
23]. Numerous scholars have investigated both traditional methods and deep learning approaches. Jiang et al. [
24] proposed an improved Radon transform combined with a Z-function and a dual-threshold mask. They estimated the blur kernel from a single image and achieved restoration using an improved RL acceleration algorithm. However, the Radon transform is sensitive to noise and image resolution, which can easily lead to errors in blur kernel angle estimation. Hou et al. [
25] quickly estimated the motion blur angle via the bispectral transform and PCA, estimated the blur length by combining the Radon transform with adjustable weights, and finally reconstructed the star points using a regional filtering method based on the super-Laplacian prior restoration algorithm. Nevertheless, this method relies on empirically set thresholds, resulting in poor stability. Additionally, some scholars have estimated the blur kernel using gyroscope data [
26,
27,
28], but this approach depends on assumptions about the motion model. Zhang et al. [
29] proposed a simple and effective
regularization deblurring method based on star image brightness, which is applicable only to uniform blur scenarios in star images.
Chen et al. [
30] proposed a deep learning method for star image motion blur, estimating blur kernel parameters essential for star deblurring, including blur angle and blur length, using sparse representation, a super-Laplacian prior, and an integrated neural network. To improve star sensor performance in high-dynamic-range environments, Liu et al. [
31] proposed a Richardson-Lucy (RL) algorithm based on a radial basis function neural network (RBFNN). However, both methods require auxiliary processing and do not achieve end-to-end direct restoration. Zhang et al. [
32] proposed an improved network based on MIMO-UNet, replacing the ResBlock module with the Res FFT-Conv module. This module learns image features from both the spatial and frequency domains, thereby enhancing the network’s deblurring performance. Wang et al. [
33] incorporated motion flow maps and sharp star points as supervision signals, conducting end-to-end training via conditional generative adversarial networks (cGANs) to achieve direct restoration of motion-blurred star maps. However, both methods require large-scale datasets for network training and impose high demands on dataset quality.
Traditional methods for star image deblurring often require prior knowledge of the blur kernel, which is difficult to obtain in practical applications. Additionally, their restoration performance depends on the quality of the original image, making them prone to generating ringing artifacts [
34]. When dealing with star images with complex backgrounds and non-uniform blur, traditional methods often exhibit limitations that restrict their effectiveness under variable observation conditions. However, the datasets used in current deep learning methods are often common ones, such as the GOPRO dataset [
35] and RealBlur dataset [
36], with no dedicated datasets for star images. On the other hand, the starfield detection images employed in this study present challenges in reconstructing overlapping regions where motion blur occurs between neighboring stars and space debris. Furthermore, the inconsistent motion directions of space debris relative to stars within the detection images result in multi-directional, multi-scale non-uniform blurring. This phenomenon complicates the simultaneous accurate reconstruction of blur kernels across different regions.
To address the issues above, this paper proposes a star image motion blur restoration network based on a variable attention mechanism. It adopts the Multi-scale Dynamic Strip Pooling Module (MDSPM) to extract image features and capture contextual blur information at different scales. Furthermore, the Multi-scale Feature Fusion Module (MFFM) fuses the extracted feature maps. Ultimately, the blurred star images are deblurred into sharp images, laying a foundation for subsequent star point detection.
2. Materials and Methods
2.1. Image Degradation Model
When capturing images with optical telescopes, spatial degradation occurs in the resulting space exploration images due to diffraction within the optical system, lens distortion of the sensor, relative motion between stars and the camera, and noise effects during imaging and transmission process. In modeling the star image degradation process, the optical imaging system is usually regarded as a linear time-invariant system [
37]. When the blurring in an image is uniform, the blurring process can be represented by Equation (
1), which is a simple convolution:
where
denotes the original sharp image,
h represents the spatially uniform blurring kernel, ∗ is the convolution operation,
stands for the blurred image, and
represents noise. When the blurring kernel ceases to be uniform, the mathematical model may be expressed as a linear transformation [
38]:
where
H denotes the blur matrix, with each row corresponding to the blur kernel modeling at a specific pixel location.
For motion blur, its blur kernel function is given by Equation (
3), where
L denotes the motion blur length,
represents the blur angle relative to the x-axis, and non-uniformly blurred images contain multiple
:
The application of the frequency domain transform to
in Equation (
3) results in the following frequency domain representation:
2.2. Network Architecture of MSDeblurNet
The network architecture of MSDeblurNet (Multi-Scale DeblurNet) is illustrated in
Figure 2. The network takes 16-bit single-channel star detection images of size
as input. These images exhibit blurring with overlapping motion blur regions and non-uniform motion blur (i.e., containing multiple blur kernels with varying angles and lengths). The network primarily comprises an encoder, a multi-scale feature fusion module, and a decoder.
The network first extracts features from the input images via the encoder. Each encoder layer comprises convolutional layers, a Multi-scale Dynamic Strip Pooling Module (MDSPM), and eight residual blocks. Specifically, the MDSPM is designed to extract multi-scale contextual information from the image, enabling the network to capture motion blur characteristics of diverse sizes and directions and thereby mitigating blurring artifacts in the input. Subsequently, the Multi-scale Feature Fusion Module fuses the multi-resolution feature maps extracted by the network, providing rich feature representations that facilitate the restoration of overlapping, occluded blur and non-uniform motion blur. The decoder employs eight residual blocks to reconstruct the output features, ultimately restoring motion-blurred star detection images and yielding high-quality, sharp detection images.
2.3. Multi-Scale Dynamic Strip Pooling Module
During long exposure of the telescope, relative motion between stars and space debris induces streak trails on the image plane. Debris moving in diverse directions gives rise to inconsistent blur kernels, which require separate estimation and thereby exacerbate the difficulty of restoration. Furthermore, when streaked stars overlap with each other, the edges of adjacent stars become blurred and merged, resulting in ambiguous edge information.
To address the above issues, this paper designs a Multi-scale Dynamic Strip Attention Mechanism Module. The Strip Attention Mechanism captures horizontal and vertical information in the image, adjusting spatial pixel weights to emphasize key features. Given the characteristics of non-uniform motion blur, square pooling at different scales and the Multi-scale Dynamic Strip Attention Mechanism are designed in the feature extraction stage. While preserving the key spatial information of the input feature maps, this module captures blur features at various scales in the image, followed by feature extraction and enhancement. The specific structure of the module is illustrated in
Figure 3:
Let the module input be
X. Through two parallel
convolutions, intermediate features
and
are generated, respectively, achieving dimensionality reduction across channels and preliminary feature extraction.
Equation (
5) represents the multi-scale square pooling branch, corresponding to the upper half of
Figure 3. Three parallel processing streams extract local blurred features at different scales. The three feature streams are element-wise summed to yield
, thereby achieving adaptive perception capabilities for blurring at varying scales.
Equation (
6) represents the dynamic strip attention branch corresponding to the lower half of
Figure 3. This branch achieves adaptive enhancement of directional features for strip-shaped blurs with inconsistent motion directions. The variable strip attention output
is obtained by multiplying horizontal and vertical strip features by their respective weights and subsequently summing them, thereby resolving the reconstruction challenge posed by inconsistent motion directions among fragments.
The combined action of multi-scale local features and directional strip features, coupled with residual connections preserving original information, ultimately yields the enhanced feature map . This module possesses adaptive perception and enhancement capabilities for multi-scale, multi-directional non-uniform blurring.
2.4. Multi-Scale Feature Fusion Module
Dilated convolution achieves the dual objective of expanding the receptive field while preserving the resolution of the output feature map by introducing fixed gaps (dilation) into standard convolution. This approach significantly reduces the loss of image detail caused by downsampling in standard convolution, which is used to increase the receptive field. Dilated convolution introduces the core parameter ’dilation rate’, which defines the pixel spacing between elements within the convolutional kernel. As illustrated in
Figure 4: (a) represents standard convolution, which is equivalent to dilated convolution with a dilation rate of 1 and a receptive field of
; (b) depicts dilated convolution with a dilation rate of 2 and a receptive field of
; and (c) depicts dilated convolution with a dilation rate of 3, which yields a receptive field of
. Clearly, a
dilated convolution achieves the same result as a
or
convolution. This approach preserves greater spatial detail by avoiding information loss from direct downsampling and effectively integrates multi-scale feature information.
During the feature fusion stage, the Multi-scale Feature Fusion Module is designed using dilated convolutions and strip pooling, with its structure shown in
Figure 5. This module integrates multi-scale feature maps extracted by the encoder. By employing dilated convolutions, it expands the receptive field while preserving the size of intermediate feature maps, thereby minimizing the loss of spatial detail information. The fused feature maps contain more comprehensive image information, alleviating the limitations of relying on a single feature set. This enhances the robustness of the network’s output, empowering the network to handle images in a wider variety of scenarios.
To effectively restore long-strip motion blur in spatial images, the module employs parallel multi-branch processing to integrate multi-directional and multi-scale contextual information. Adaptive average pooling compresses features along the height and width dimensions, reconstructing feature maps with receptive fields of distinct aspect ratios to capture the structural features of elongated blur in images. Meanwhile, dilated convolutions learn spatial features at various scales. Finally, fusing outputs from multi-directional pooling branches and multi-scale dilated convolution branches significantly enhances the model’s capability to restore elongated blur in spatial images.
The multi-scale features output by the encoder are denoted as
,
and
, which also serve as inputs to the modules depicted in
Figure 5.
Equation (
8) performs feature fusion on the feature maps from different layers of the encoder, providing a unified input feature for subsequent multi-branch processing.
Equations (
9)–(
11) address the diversity in direction and scale of strip-shaped motion blur within images by designing three parallel strip pooling branches. Through adaptive average pooling, these branches form receptive fields of varying scales along both the height and width dimensions, thereby capturing the structural features of strip-shaped blurring.
Equation (
12) designs a multi-scale dilated convolution kernel, achieving the fusion of multi-scale spatial features and addressing the shortcomings of single-scale convolution in perceiving strip-shaped blur patterns.
,
, and
represent the results of different pooling operations, respectively, and
is the output of convolutions with different dilation rates applied to the feature map. Splicing these together and then incorporating the initial fusion feature map Y via residual connections yields the module output
. This module can assign higher weights to blurred regions in the image through multi-scale strip pooling and dilated convolution operations, enabling the network to better learn key information in the image and improve the efficiency of feature utilization.
5. Conclusions
This paper proposes an image restoration algorithm for star images with elongated motion blur induced by long exposure times, enabling end-to-end image deblurring. The algorithm adopts an encoder–decoder architecture and employs a multi-scale dynamic strip attention mechanism in the feature extraction stage. Through the design of dynamically weighted strip convolution, it performs differentiated processing on blur patterns, aiding the network in adapting to images with varying blur characteristics. Additionally, addressing the issues that conventional convolutional pooling in the feature fusion stage tends to cause detail loss and insufficient complementarity of cross-level features, this paper innovatively proposes a multi-scale feature fusion module integrating dilated convolution and strip pooling. On one hand, it extracts contextual features via the multi-scale receptive fields of dilated convolution; on the other hand, it captures prominent strip structural information in horizontal and vertical directions through strip pooling. Comparative experiments demonstrate that MSDeblurNet achieves the best performance across all quantitative metrics, including PSNR (84.08 dB), SSIM (0.9928), as well as astronomy-specific star extraction rate (80.37%) and centroid error (0.4589 pixels), with its restoration capability significantly superior to other methods. Further validation on measured data shows that images restored by MSDeblurNet achieve remarkable improvements in the number of extracted stars and centroid positioning accuracy, verifying the proposed algorithm’s strong adaptability to real-world scenarios, along with excellent robustness and effectiveness.