Next Article in Journal
A Hybrid Smoothing Model for Historical Reconstruction of Route-Level Airline Passenger Demand
Previous Article in Journal
Wolffia globosa-Fortified Hydrogels for Extrusion-Based 3D Food Printing: Effects of Particle Microstructure and Process Parameters on Dimensional Fidelity
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Spatial-Aware Modulation for Implicit Neural Representations

1
The 10th Research Institute of China Electronics Technology Group Corporation, Chengdu 610036, China
2
School of Computer Science and Engineering, University of Electronic Science and Technology of China, Chengdu 611731, China
3
School of Robotics and Advanced Manufacture, Harbin Institute of Technology, Shenzhen 518055, China
*
Author to whom correspondence should be addressed.
Appl. Sci. 2026, 16(17), 8870; https://doi.org/10.3390/app16178870
Submission received: 10 July 2026 / Revised: 30 August 2026 / Accepted: 2 September 2026 / Published: 7 September 2026

Abstract

Implicit Neural Representations (INRs) provide a flexible and resolution-independent formulation for continuous signal representation. Despite their strong representation ability, standard INRs usually predict each queried coordinate independently, making it difficult to explicitly exploit the local coherence widely observed in natural signals. For images and volumetric data, neighboring locations often share correlated responses in smooth regions, while sharp variations mainly appear around spatial transitions. Ignoring such local dependency may reduce learning efficiency and weaken the reconstruction of spatially consistent details. To address this limitation, we propose Spatial-Aware Implicit Neural Representation (SA-INR), which enhances INRs by introducing local feature aggregation into the hidden representation space. Motivated by local feature coherence, SA-INR aggregates neighboring coordinate features through a learnable spatial-aware local operator. The aggregation weights are initialized as a uniform mean filter, providing a smooth local bias during early optimization. As training proceeds, the aggregation weights are updated by reconstruction supervision and become adaptive to spatial content. To preserve coordinate-specific information and avoid over-smoothing, the aggregated feature is further integrated with the original feature through a residual connection. Extensive experiments on image representation, CT reconstruction, and image denoising demonstrate that SA-INR consistently improves reconstruction fidelity across different INR backbones and reconstruction tasks. These results suggest that explicitly modeling local feature interaction is an effective way to enhance continuous signal representation.

1. Introduction

Implicit Neural Representations (INRs) model signals as continuous functions of spatial coordinates, providing a resolution-independent alternative to discrete representations such as pixel grids or voxel arrays. INRs have been successfully applied to tasks including image fitting [1,2,3,4,5], super-resolution [6,7,8], and general signal modeling [9,10,11,12]. Recent advances, such as periodic activations, WIRE [13], Gaussian-based activations [14], and SL2A [15], improve the ability of INRs to capture high-frequency details and complex signal variations.
Despite these improvements, standard INRs predict each coordinate independently, without explicitly modeling the strong local correlations present in natural signals. For images and volumetric data, neighboring locations often share similar responses in smooth regions, with large variations mainly occurring around sharp transitions. Ignoring such local dependency can reduce learning efficiency and limit the reconstruction of fine spatial details [16,17,18,19].
Prior studies have attempted to alleviate the limitations of standard INRs in modeling complex natural signals from different perspectives. Multi-scale and band-limited methods, such as BACON [20], MINER [21], and BANF [22], mainly improve the reconstruction of high-frequency details by enhancing frequency representation. However, these methods primarily focus on frequency-domain modeling and lack explicit exploitation of local correlations in spatial neighborhoods. Localized approaches, such as Point-NeRF [23] and LINR [24], introduce local context through neighborhood-based computations, but they often rely on specific data structures or additional computational procedures. Other methods incorporate spatial priors using convolutional encoders, spatial partitions, or hash-grid-based discretized representations [25,26,27,28,29]. Although these strategies improve representation capacity, they may compromise the continuous and compact coordinate-based representation advantage of INRs. In addition, global mechanisms such as self-attention and kernel transformations can capture long-range dependencies [30,31,32]; nevertheless, they typically incur high computational complexity and are not specifically designed to model the local coherence inherent in natural signals.
In this work, we propose Spatial-Aware Implicit Neural Representation (SA-INR), which enhances INRs by explicitly modeling local feature interactions. SA-INR aggregates hidden features over spatially neighboring coordinates using a learnable spatial-aware operator. The aggregation weights are initialized as a uniform mean filter, providing a smooth local bias consistent with natural signal coherence. To preserve coordinate-specific information and prevent over-smoothing, the aggregated features are integrated with the original features through a residual connection, allowing the network to exploit local context while maintaining fine-grained details. Although SA-INR introduces local spatial feature interaction, it still follows the coordinate-based implicit function formulation of INR models. The proposed local aggregation operates in the hidden feature space to provide additional spatial context without introducing discrete feature grids, thereby preserving the continuous and resolution-independent representation capability of INR models.
Our main contributions are summarized as follows:
  • We propose Spatial-Aware Implicit Neural Representation (SA-INR), which introduces local feature interactions into coordinate-based implicit neural representations.
  • We design a learnable spatial-aware local operator with channel-independent aggregation, mean-filter initialization, and residual feature integration, enabling neighboring coordinate features to be incorporated while preserving pointwise representations.
  • We conduct extensive experiments on different datasets, demonstrating consistent performance improvements across different INR backbones and reconstruction tasks.

2. Related Work

2.1. Implicit Neural Representations

Implicit Neural Representations (INRs), also referred to as coordinate-based networks, model continuous signals as neural functions typically implemented using Multi-Layer Perceptrons (MLPs). In image representation, INRs learn a mapping from spatial coordinates to pixel intensities, providing a continuous alternative to discrete grid-based representations. This formulation enables resolution-independent modeling and has been widely adopted in tasks such as image fitting, compression [1,2,3], and super-resolution [6,7,8], with extensions to other continuous image transformation tasks such as registration [33]. Moreover, INR-based formulations have also been extended to more general continuous signal modeling settings, such as dynamic scene representation and neural fields, where dense spatial queries are required [34,35,36]. While highly expressive, standard MLP-based INRs equipped with ReLU activations exhibit spectral bias [16], an inherent tendency to fit low-frequency components faster than high-frequency ones. This phenomenon, analyzed from theoretical and Fourier perspectives [17,18,19], helps explain the difficulty of reconstructing fine details. To mitigate this limitation, a substantial body of work has focused on improving frequency representations. Fourier Features [10] project input coordinates into a higher-dimensional sinusoidal space to broaden spectral coverage, while SIREN [9] employs periodic activations to better model high-frequency signals. Attention-based INR methods such as TransINR [31] introduce global feature interactions through self-attention mechanisms. Recent work LINR [24] explores local implicit representation learning for visual object tracking by refining target-specific local regions with implicit neural representations. Subsequent efforts further improve frequency modeling and training stability, including Fourier reparameterization [37], frequency tuning strategies [38], and architectural variants such as Multiplicative Filter Networks [39], SL2A [15], FINER [40], and WIRE [13]. Despite these advances in frequency modeling, most INR formulations process each coordinate independently and do not explicitly model spatial relationships between neighboring locations. This limitation motivates the development of approaches that incorporate spatial dependencies into implicit representations, which we review next.

2.2. Spatial Dependency Modeling

To address the lack of explicit spatial modeling in coordinate-based networks, a growing body of work introduces spatial dependencies into implicit representations. One common direction augments coordinate-based models with auxiliary encoders that provide spatial context. Methods such as ConvONet [25], PixelNeRF [27], and DeepLS [26] extract localized features on discrete grids or latent codes to condition the implicit function, while subsequent extensions incorporate prior embeddings or hybrid representations to improve structural consistency [41,42]. These approaches effectively inject spatial information but rely on external feature extractors and discrete representations. Another line of work introduces spatial structure through explicit decomposition or discretization. Approaches such as KiloNeRF [28] and DIVeR [29] partition the spatial domain into multiple sub-networks, following a divide-and-conquer strategy. Related ideas are further explored in partition-based learning frameworks [43] and ensemble methods such as Levels-of-Experts [44]. To improve scalability, structured spatial representations have been adopted, including sparse octrees [45,46] and multi-resolution hash encodings [47]. In addition, recent studies introduce spatially aware modules or point-based representations to guide reconstruction [23,24], enabling more flexible local modeling. Beyond explicit locality, several works attempt to capture spatial dependencies through global interaction mechanisms. For example, TransINR [31] applies self-attention to model long-range correlations, while subsequent studies explore attention-based or kernel-based formulations for implicit representations [32,48]. Related developments in spatially-aware INR variants further demonstrate the effectiveness of incorporating local attention or structured aggregation for improving reconstruction fidelity [49]. Overall, existing methods introduce spatial dependencies into implicit representations through auxiliary feature encoders, explicit spatial decomposition or discretization, and global interaction mechanisms. From the perspective of modern representation learning, these approaches incorporate spatial information by augmenting coordinate representations with external features, structured spatial representations, or feature interaction mechanisms. In contrast, SA-INR directly introduces learnable local interactions among neighboring coordinate-conditioned latent representations in the hidden feature space. This design enables spatial information exchange within the INR itself, without relying on additional feature encoders or discrete spatial representations. Therefore, SA-INR provides a lightweight way to incorporate local spatial coherence while preserving the continuous formulation of coordinate-based networks.The overall architecture of the proposed SA-INR is illustrated in Figure 1.

3. Method

3.1. Local Coherence Propagation

Implicit neural representations (INRs) represent continuous signals through coordinate-based neural networks. Given a spatial coordinate x i R d , an INR learns a mapping f θ : x i y i , where y i denotes the signal value associated with coordinate x i . Let Φ θ denote the hidden feature mapping induced by the preceding layers of the INR network. The intermediate feature of coordinate x i can then be written as z i = Φ θ ( x i ) R C , where C is the number of feature channels.
Standard INRs usually predict the signal value at each coordinate mainly from its own hidden response. Such a coordinate-wise modeling paradigm implicitly treats different queried coordinates as independent samples during feature transformation. However, continuous natural signals, such as images and volumetric data, generally exhibit strong local correlations. Neighboring locations tend to have similar signal values in smooth regions, whereas significant variations are mainly concentrated around edges, textures, or structural boundaries. Since the hidden representation is expected to preserve the underlying spatial structure required for signal reconstruction, neighboring coordinates are also expected to produce correlated latent responses before the final prediction layer. Therefore, besides the output signal itself, local coherence should naturally extend to the hidden representation space, where neighboring features provide complementary spatial context for each queried coordinate.
To provide a theoretical interpretation of such local coherence, we consider an ideal undirected graph constructed over the queried coordinates. Each node in the graph corresponds to a hidden feature z i , and each edge encodes the local relationship between neighboring coordinates. Let S R N × N denote the affinity matrix of this ideal local graph, where S i j measures the local coherence between coordinates x i and x j . Here, S is introduced only for theoretical motivation and is assumed to be symmetric and non-negative: S i j = S j i , S i j 0 . Accordingly, the degree matrix D s is defined as
D s , i i = j = 1 N S i j ,
and the corresponding graph Laplacian is given by L s = D s S .
By stacking all hidden features into a matrix Z = [ z 1 , z 2 , , z N ] R N × C , the local coherence energy in the hidden representation space can be formulated as
E loc = 1 2 Tr Z L s Z .
Since S is symmetric and non-negative, the above energy is equivalently expressed as
E loc = 1 4 i = 1 N j = 1 N S i j z i z j 2 2 .
The factor 1 / 4 is used to compensate for the double counting of undirected edges in the double summation. This energy encourages neighboring coordinates with high affinity to produce similar feature responses in the hidden space.
Since L s is a symmetric graph Laplacian, the gradient of E loc with respect to Z is
E loc Z = L s Z .
Therefore, one gradient descent step on the local coherence energy yields.
Z + = Z λ L s Z , where λ denotes the propagation step size. Substituting L s = D s S into the above equation gives Z + = I λ D s + λ S Z . This update provides an interpretation of how local coherence can be encouraged through interactions among neighboring hidden representations. Specifically, under the ideal graph formulation, neighboring features with high affinity tend to become more consistent during the corresponding optimization process. It is worth emphasizing that the above local coherence energy is introduced only as a theoretical motivation for local feature interaction. Inspired by this graph-based local coherence perspective, we design a learnable spatial-aware aggregation operator to introduce interactions among neighboring coordinate representations. The proposed operator does not explicitly optimize the graph Laplacian energy or enforce the constraints of the ideal affinity matrix. Instead, its aggregation weights are learned directly from the reconstruction objective.

3.2. Spatial-Aware Local Operator

Based on the above local-propagation motivation, we parameterize neighborhood feature interaction using a learnable spatial-aware aggregation operator. It should be noted that the practical aggregation weights and the ideal affinity matrix S in the theoretical analysis serve different roles. The theoretical matrix S is assumed to be symmetric and non-negative to establish the standard graph-Laplacian smoothness interpretation. In contrast, the practical aggregation weights are learned end-to-end from reconstruction supervision without imposing symmetry, non-negativity, or other graph-affinity constraints. Consequently, the learned weights should not be interpreted as direct estimates of the ideal matrix S, nor are they guaranteed to define a strict graph-diffusion process. This relaxation is intentional, as it allows different spatial offsets and feature channels to learn different local responses, thereby providing a more flexible mechanism for feature interaction than a fixed graph diffusion. Meanwhile, this flexibility also means that the practical operator does not inherit the strict Laplacian properties of the ideal graph model. Therefore, the relationship between the graph formulation and the implemented spatial-aware operator should be understood as conceptual inspiration based on the principle of local propagation, rather than mathematical equivalence of the optimization processes.
For the hidden feature z i associated with coordinate x i , let Δ denote the set of spatial offsets within a local neighborhood. For image representation, Δ corresponds to the relative offsets within a local window centered at the queried coordinate. For the c-th feature channel, the aggregated feature is computed as
z ˜ i , c = δ Δ k δ , c z i + δ , c ,
where z i + δ , c denotes the neighboring hidden feature at offset δ , and k δ , c is the learnable aggregation weight for the corresponding spatial offset and feature channel. In vector form, the operator is written as
z ˜ i = A θ ( z i ) = δ Δ k δ z i + δ ,
where k δ R C denotes the channel-wise aggregation weights, ⊙ represents element-wise multiplication, and A θ denotes the proposed spatial-aware aggregation operator.
Unlike dense cross-channel interaction, the proposed operator performs independent aggregation for each feature channel while introducing only local spatial interaction. This design preserves the original channel-wise representation and avoids unnecessary interference across feature channels. Moreover, since the aggregation is restricted to a local neighborhood, the proposed module remains computationally lightweight and does not require an additional image encoder or discrete feature grid.
To introduce a stable local smoothness prior during the early stage of optimization, the aggregation weights are uniformly initialized as
k δ , c ( 0 ) = 1 | Δ | , δ Δ ,
where | Δ | denotes the number of spatial offsets in the local neighborhood. With this initialization, the proposed operator initially behaves as a local mean filter in the hidden feature space, providing a smooth optimization bias consistent with the local coherence prior.
During training, the aggregation weights are optimized together with the INR backbone through the reconstruction objective. Consequently, the operator gradually evolves from uniform neighborhood averaging into a content-adaptive local aggregation mechanism, enabling effective information sharing in smooth regions while preserving structural details around edges and texture transitions.

3.3. Residual Feature Integration

Directly replacing the original feature inevitably weakens the coordinate-specific response that constitutes the basis of INR representations. This issue is particularly critical in regions containing high-frequency details or sharp structural transitions, where overly strong local aggregation may lead to over-smoothed feature responses. To incorporate local feature interaction while preserving the original coordinate-wise information, we integrate the aggregated feature through a residual formulation.
For the c-th feature channel, the residual integration is defined as z ^ i , c = z i , c + γ c z ˜ i , c , where γ c denotes the channel-wise scaling coefficient that controls the contribution of the aggregated feature to the updated hidden representation. When a learnable gating mechanism is adopted, γ c can be optimized as a trainable parameter. Otherwise, setting γ c = 1 reduces the formulation to a standard residual connection. In vector form, the above integration can be written as z ^ i = z i + γ A θ ( z i ) , where γ R C is a channel-wise scaling vector and ⊙ denotes element-wise multiplication. This residual design serves two purposes. First, it preserves the original coordinate-wise feature response, preventing the local aggregation from completely overriding coordinate-specific information. Second, it allows neighborhood features to participate in the subsequent nonlinear transformation, thereby enhancing local coherence in the hidden representation space.
For SIREN-based INRs, the integrated feature is fed into a sinusoidal activation function. For the c-th channel, this process is given by h i , c = sin ω 0 z ^ i , c , where ω 0 controls the frequency scale of the sinusoidal activation. Substituting the residual integration into the above equation yields
h i , c = sin ω 0 z i , c + γ c z ˜ i , c .
Using the trigonometric identity for the sine of a sum, the activation response can be decomposed as
h i , c = sin ω 0 z i , c cos ω 0 γ c z ˜ i , c + cos ω 0 z i , c sin ω 0 γ c z ˜ i , c .
This decomposition provides an intuitive interpretation of how neighborhood information participates in the nonlinear activation. Instead, it modulates the nonlinear sinusoidal response through coupled sine–cosine interactions. In this way, neighborhood information can influence the activation dynamics, while the original feature z i is still preserved through the residual path. Consequently, the coordinate-specific expressiveness of the original INR is maintained while local feature interaction is introduced.
Overall, the aggregated feature serves as a complementary source of local spatial context rather than a replacement for the original coordinate feature. Combined with the residual pathway, the proposed design preserves the expressive power of conventional INRs while introducing adaptive neighborhood interactions for more accurate signal reconstruction.

4. Experiments

To improve reproducibility, we provide additional implementation details. All experiments are independently trained using three random seeds (42, 43, and 44). The batch size is set to 1, and all models are optimized using the Adam optimizer with an initial learning rate of 1 × 10 4 for 500 optimization steps. The input coordinates are normalized to [ 1 , 1 ] 2 before training. During evaluation, the model corresponding to the best PSNR obtained during optimization is selected for reporting.
To comprehensively evaluate the proposed method, we conduct experiments on three representative tasks: image representation, CT reconstruction, and image denoising. Detailed descriptions of the datasets and evaluation protocols are provided in the corresponding subsections. Unless otherwise specified, all models adopt a standard Multi-Layer Perceptron (MLP) backbone with three hidden layers of width 256. To ensure fair comparisons, we strictly follow the original configurations of all baseline methods. In particular, sinusoidal activations are used for SIREN-based models, while the native architectural designs and activation functions are preserved for FINER and FR-INR. All models are trained for 500 optimization steps under identical settings.
The proposed SA-INR is designed as a general, plug-and-play enhancement for coordinate-based networks. By augmenting standard linear transformations with a spatial-aware residual branch, our method explicitly introduces local structural dependencies into the inherently point-wise optimization process, without modifying the core architecture of the underlying INR backbone. For all experiments, optimization is performed using the Adam optimizer with an initial learning rate of 1 × 10 4 . Reconstruction quality is evaluated using two standard metrics: Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index (SSIM), which jointly measure pixel-wise accuracy and perceptual structural fidelity.

4.1. Image Representation

To assess the core representational capacity of SA-INR, we first conduct experiments on the task of image representation using the Kodak24 dataset [50], with all images resized to 256 × 256 resolution. As reported in Table 1, incorporating the spatial-aware module generally improves performance across all baseline architectures. For example, SA-SIREN increases the PSNR from 34.62 dB to 37.12 dB. More importantly, SA-FR-INR achieves the best overall performance, reaching 45.25 dB in PSNR and 0.9900 in SSIM. These consistent gains across different backbones indicate that explicitly modeling local spatial dependencies effectively enhances the representational capacity of coordinate-based networks, which otherwise rely solely on point-wise mappings.
Figure 2 presents qualitative comparisons. Standard INR baselines tend to produce artifacts in homogeneous regions and exhibit degraded reconstruction of high-frequency details. In contrast, SA-INR yields sharper structural boundaries and smoother flat regions. This improvement can be attributed to the spatial-aware residual branch, which introduces local neighborhood interactions into the inherently independent coordinate-wise predictions. As a result, the network achieves a better balance between preserving fine details and maintaining spatial coherence, leading to reconstructions that are both quantitatively superior and visually more faithful to the ground truth.
Evaluation on High-Resolution Images. To further evaluate the scalability and resolution-independent capability of the proposed method, we extend our experiments to the Flickr2K dataset [51], which contains high-resolution images with rich high-frequency details. Quantitative comparisons are reported in Table 2. As shown in Table 2, SA-INR consistently improves different INR backbones under high-resolution image representation settings. For example, SA-FR-INR achieves 30.74 dB PSNR and 0.8771 SSIM, outperforming FR-INR by 1.08 dB and 0.0478, respectively. The qualitative comparisons in Figure 3 further demonstrate that standard coordinate-based baselines tend to produce over-smoothed regions or local artifacts when reconstructing complex high-resolution signals. In contrast, SA-INR achieves sharper structural boundaries and better preserves fine-grained textures through spatial-aware feature aggregation. These results demonstrate that SA-INR remains effective for high-resolution continuous signal representation. Since the reconstruction is still performed through coordinate queries without introducing discrete feature grids, the proposed spatial-aware module preserves the resolution-independent property of INR models.

4.2. CT Reconstruction

We further evaluate SA-INR on sparse-view CT reconstruction, a highly ill-posed inverse problem that requires strong structural priors. For this task, we use 20 randomly selected CT images from the TCIA dataset [52], each with a resolution of 512 × 512 . To simulate a realistic sparse-view setting, we generate measurements from 150 projection angles. The sparse-view CT reconstruction experiments follow the experimental protocol of WIRE, including CT data preprocessing, projection generation, and reconstruction evaluation settings.
Quantitative results are reported in Table 3. Incorporating the proposed spatial-aware module consistently improves performance across all baseline methods. For instance, SA-SIREN increases the PSNR from 29.23 dB to 32.44 dB. These results indicate that SA-INR significantly enhances the ability of coordinate-based models to recover accurate signals under severely limited measurements. Figure 4 shows qualitative comparisons. Due to the underdetermined nature of sparse-view reconstruction, standard INR baselines suffer from noticeable streak artifacts and structural distortions. In contrast, SA-INR produces cleaner reconstructions with improved anatomical consistency. This improvement can be attributed to the spatial aggregation mechanism, which acts as an implicit regularizer by enforcing local consistency across neighboring regions. As a result, SA-INR effectively reduces ambiguity caused by missing projections and better preserves fine structural details.

4.3. Image Denoising

To evaluate the performance of the proposed SA-INR under noisy conditions, we conduct image denoising experiments on the Flickr2K dataset [51]. We randomly select 20 high-resolution images and downsample them by a factor of 0.5 to balance computational efficiency and structural complexity. Noisy observations are generated by modeling a combination of photon and readout noise, where each pixel is corrupted using independent Poisson random variables. In this setup, the mean photon count ( τ ) and the readout noise level ( r o ) are set to 50 and 2, respectively, resulting in a stochastic observation setting that poses significant challenges for accurate signal recovery.
Quantitative results are reported in Table 4. The proposed spatial-aware module consistently improves performance across all baseline methods. In particular, SA-FR-INR achieves the best results, reaching 27.59 dB in PSNR and 0.7720 in SSIM. These results indicate that incorporating spatial awareness improves the reconstruction of clean signals from stochastic corruption.Specifically, all SA-INR variants achieve improvements in PSNR, while SA-FINER exhibits a slight decrease in SSIM compared with the original FINER model. This indicates that the effect of local feature aggregation may vary across different evaluation metrics, reflecting a potential trade-off between pixel-wise reconstruction fidelity and structural similarity in this setting. Figure 5 shows qualitative comparisons. Standard INR baselines often exhibit a trade-off between noise suppression and detail preservation, resulting in either residual noise or over-smoothed structures. In contrast, SA-INR produces cleaner reconstructions while preserving important structural details and fine textures. This improvement can be attributed to the spatial-aware aggregation mechanism, which leverages local neighborhood consistency to distinguish noise from underlying signal structures. As a result, SA-INR effectively suppresses noise while maintaining high-frequency details.

5. Ablation Study

To evaluate the contribution of each component in SA-INR, we conduct ablation studies on the Kodak24 image fitting task using SIREN as the backbone. The ablation follows the design order of Section 3: the local feature aggregation operator, the residual point-wise pathway, and the mean-filter initialization. We report PSNR and SSIM to evaluate the reconstruction quality. Although the main experiments demonstrate the effectiveness of SA-INR across different INR backbones, including FINER and FR-INR, comprehensive component-level ablation studies across different architectures remain an interesting direction for future research.
Local aggregation design. We first examine how local feature interaction should be introduced into the hidden representation space. As shown in Table 5, directly applying dense cross-channel aggregation severely degrades the performance of SIREN. This result suggests that unconstrained cross-channel mixing is unsuitable for sinusoidal coordinate networks, as it may disturb the channel-wise frequency responses. In contrast, the proposed channel-independent local aggregation improves the baseline, indicating that local spatial interaction is beneficial when it is introduced in a lightweight and channel-preserving manner.
Residual point-wise pathway. We next study the effect of preserving the original coordinate-wise feature response. As reported in Table 6, removing the residual connection leads to a clear performance drop. This confirms that the aggregated neighborhood context should not replace the original INR feature entirely. Instead, combining the aggregated feature with the original point-wise feature enables the model to exploit local context while retaining coordinate-specific details, which is consistent with the residual spectral modulation described in Section 3.3.
Mean-filter initialization. We then evaluate the initialization strategy of the aggregation weights. Table 7 shows that random initialization substantially reduces the reconstruction quality. In contrast, initializing the aggregation weights as a uniform mean filter provides a smooth local prior at the beginning of optimization and leads to more stable training. This result supports our motivation that initialization is an important part of the local operator design rather than a minor implementation detail.
Neighborhood size. We further investigate the effect of the local neighborhood size used in the aggregation operator. As shown in Table 8, the setting without spatial neighborhood interaction performs only point-wise feature modulation, without information exchange among neighboring coordinates. Compared with this point-wise variant, the 3 × 3 neighborhood incorporates features from adjacent coordinates and achieves the best performance, demonstrating the benefit of explicitly modeling local spatial interactions. Specifically, the 3 × 3 setting improves the PSNR by 1.35 dB and the SSIM by 0.0197 compared with the variant without spatial neighborhood interaction. As the neighborhood size increases to 5 × 5 and 7 × 7 , the performance gradually decreases. This suggests that excessively large neighborhoods may introduce redundant spatial information and weaken coordinate-specific responses. Therefore, the 3 × 3 neighborhood provides an effective balance between local spatial interaction and coordinate-specific representation preservation.

6. Additional Results

Computational Cost. To evaluate the additional computational cost introduced by SA-INR, we compare SIREN and SA-SIREN under the same network architecture, input resolution, training configuration, and hardware environment. We report the number of parameters, training time, inference time, and peak GPU memory consumption, as summarized in Table 9. Compared with SIREN, SA-SIREN introduces additional parameters mainly from the learnable local aggregation operator. Since the proposed operator performs channel-wise aggregation only within a small spatial neighborhood and does not introduce dense cross-channel interactions, the increase in model complexity remains limited. As shown in Table 9, SA-SIREN achieves improved reconstruction performance while introducing only a moderate increase in computational cost and memory consumption. These results demonstrate that the proposed spatial interaction mechanism provides an effective trade-off between representation quality and computational efficiency.

7. Discussion

Unlike conventional coordinate-based INRs that independently map each coordinate to a signal value, SA-INR introduces explicit local spatial interaction into the hidden representation space. This design provides a complementary perspective to existing INR improvement strategies. Existing methods mainly enhance representation capability through frequency-aware encoding, auxiliary feature extraction, spatial discretization, or global feature interaction, whereas SA-INR focuses on exploiting the local coherence naturally existing in continuous signals by enabling information exchange among neighboring coordinate representations while preserving the continuous formulation of INRs. The effectiveness of SA-INR can be attributed to the introduction of a spatial inductive bias into coordinate-based neural representations. By aggregating neighboring latent features and integrating them with original coordinate-specific features through residual learning, SA-INR provides additional spatial context without replacing the coordinate-based representation mechanism. Compared with approaches relying on external encoders or discrete spatial representations, SA-INR directly models spatial dependency within the implicit representation itself, maintaining the flexibility and resolution-independent property of INRs while improving spatial consistency and reconstruction fidelity. Moreover, the proposed local interaction mechanism is not limited to 2D image representation. Since Neural Radiance Fields (NeRF) and spatiotemporal neural fields also rely on coordinate-based implicit representations, SA-INR can potentially be extended to these scenarios by introducing spatial or spatiotemporal neighboring interactions. Despite these advantages, the current framework defines local interactions using a predefined neighborhood configuration, and adaptive neighborhood construction remains an interesting direction for future investigation.

8. Conclusions

In this paper, we presented Spatial-Aware Implicit Neural Representation (SA-INR), a simple yet effective enhancement to coordinate-based networks that explicitly incorporates local spatial dependencies into the implicit representation framework. By augmenting standard linear transformations with a spatial-aware aggregation mechanism and a residual modulation pathway, the proposed method relaxes the point-wise independence assumption of conventional INRs while preserving their continuous formulation and architectural simplicity. Extensive experiments on image representation, CT reconstruction, and image denoising demonstrate that SA-INR consistently improves reconstruction quality across a range of backbone architectures. The results show that explicitly modeling local spatial coherence not only enhances representational capacity but also provides an implicit regularization effect, leading to improved reconstruction performance under ill-posed and noisy conditions. Overall, this work highlights the importance of explicitly modeling local spatial interactions within implicit neural representations. We hope that the proposed design offers a useful perspective for bridging coordinate-based formulations with the intrinsic spatial structure of natural signals. Future work will further investigate comprehensive module-level analyses of SA-INR across more INR architectures, including different network structures and activation mechanisms, to better understand the general applicability of spatial feature interaction in implicit neural representations.

Author Contributions

Conceptualization, C.Q. (Chen Qing), D.H. and C.Q. (Caiyan Qin); methodology, C.Q. (Chen Qing) and D.H.; software, W.Z., H.W. and D.H.; validation, W.Z., H.W. and D.H.; formal analysis, C.Q. (Chen Qing), D.H. and M.Z.; investigation, C.Q. (Chen Qing), W.Z. and D.H.; resources, D.H., M.Z. and C.Q. (Caiyan Qin); data curation, W.Z. and H.W.; writing—original draft preparation, C.Q. (Chen Qing) and D.H.; writing—review and editing, C.Q. (Chen Qing), D.H., M.Z. and C.Q. (Caiyan Qin); visualization, W.Z., H.W. and D.H.; supervision, M.Z. and C.Q. (Caiyan Qin); project administration, C.Q. (Caiyan Qin); funding acquisition, C.Q. (Caiyan Qin). All authors have read and agreed to the published version of the manuscript.

Funding

This work is supported by National Key Research and Development Program of China (2024YFE0212200).

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The datasets used in this study are publicly available. The Kodak24 dataset is available at https://r0k.us/graphics/kodak/, accessed on 10 October 2025. The Flickr2K dataset is available at https://github.com/LimBee/NTIRE2017, accessed on 11 October 2025. The TCIA (The Cancer Imaging Archive) dataset is available at https://www.cancerimagingarchive.net/, accessed on 12 October 2025.

Acknowledgments

We would also like to thank the handling Associate Editor and the anonymous reviewers for their valuable comments and suggestions for this paper.

Conflicts of Interest

Author Chen Qing was employed by the 10th Research Institute of China Electronics Technology Group Corporation. The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

References

  1. Dupont, E.; Goliński, A.; Alizadeh, M.; Teh, Y.W.; Doucet, A. Coin: Compression with implicit neural representations. arXiv 2021, arXiv:2103.03123. [Google Scholar]
  2. Strümpler, Y.; Postels, J.; Yang, R.; Gool, L.V.; Tombari, F. Implicit neural representations for image compression. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2022; pp. 74–91. [Google Scholar]
  3. Han, J.; Zheng, H.; Bi, C. Kd-inr: Time-varying volumetric data compression via knowledge distillation-based implicit neural representation. IEEE Trans. Vis. Comput. Graph. 2023, 30, 6826–6838. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  4. Wu, Q.; Li, Y.; Xu, L.; Feng, R.; Wei, H.; Yang, Q.; Yu, B.; Liu, X.; Yu, J.; Zhang, Y. IREM: High-resolution magnetic resonance image reconstruction via implicit neural representation. In Proceedings of the Medical Image Computing and Computer Assisted Intervention–MICCAI 2021: 24th International Conference, Strasbourg, France, 27 September–1 October 2021; Proceedings, Part VI 24; Springer: Berlin/Heidelberg, Germany, 2021; pp. 65–74. [Google Scholar]
  5. Molaei, A.; Aminimehr, A.; Tavakoli, A.; Kazerouni, A.; Azad, B.; Azad, R.; Merhof, D. Implicit neural representation in medical imaging: A comparative survey. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 2381–2391. [Google Scholar]
  6. Chen, Y.; Liu, S.; Wang, X. Learning continuous image representation with local implicit image function. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 19–25 June 2021; pp. 8628–8638. [Google Scholar]
  7. Lee, J.; Jin, K.H. Local texture estimator for implicit representation function. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 1929–1938. [Google Scholar]
  8. Lu, Y.; Wang, Z.; Liu, M.; Wang, H.; Wang, L. Learning spatial-temporal implicit neural representations for event-guided video super-resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 1557–1567. [Google Scholar]
  9. Sitzmann, V.; Martel, J.; Bergman, A.; Lindell, D.; Wetzstein, G. Implicit neural representations with periodic activation functions. Adv. Neural Inf. Process. Syst. 2020, 33, 7462–7473. [Google Scholar]
  10. Tancik, M.; Srinivasan, P.; Mildenhall, B.; Fridovich-Keil, S.; Raghavan, N.; Singhal, U.; Ramamoorthi, R.; Barron, J.; Ng, R. Fourier features let networks learn high frequency functions in low dimensional domains. Adv. Neural Inf. Process. Syst. 2020, 33, 7537–7547. [Google Scholar]
  11. Su, K.; Chen, M.; Shlizerman, E. INRAS: Implicit Neural Representation for Audio Scenes. Adv. Neural Inf. Process. Syst. 2022, 35, 8144–8158. [Google Scholar] [CrossRef] [Scilit]
  12. Szatkowski, F.; Piczak, K.J.; Spurek, P.; Tabor, J.; Trzciński, T. Hypernetworks build implicit neural representations of sounds. In Proceedings of the Joint European Conference on Machine Learning and Knowledge Discovery in Databases; Springer: Berlin/Heidelberg, Germany, 2023; pp. 661–676. [Google Scholar]
  13. Saragadam, V.; LeJeune, D.; Tan, J.; Balakrishnan, G.; Veeraraghavan, A.; Baraniuk, R.G. Wire: Wavelet implicit neural representations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 18507–18516. [Google Scholar]
  14. Ramasinghe, S.; Lucey, S. Beyond periodicity: Towards a unifying framework for activations in coordinate-mlps. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2022; pp. 142–158. [Google Scholar]
  15. Rezaeian, R.; Heidari, M.; Azad, R.; Merhof, D.; Soltanian-Zadeh, H.; Hacihaliloglu, I. SL2A-INR: Single-Layer Learnable Activation for Implicit Neural Representation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Honolulu, HI, USA, 19–25 October 2025; pp. 26065–26074. [Google Scholar]
  16. Rahaman, N.; Baratin, A.; Arpit, D.; Draxler, F.; Lin, M.; Hamprecht, F.; Bengio, Y.; Courville, A. On the spectral bias of neural networks. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2019; pp. 5301–5310. [Google Scholar]
  17. Benbarka, N.; Höfer, T.; ul-moqeet Riaz, H.; Zell, A. Seeing implicit neural representations as fourier series. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA, 4–8 January 2022; pp. 2041–2050. [Google Scholar]
  18. Cao, Y.; Fang, Z.; Wu, Y.; Zhou, D.X.; Gu, Q. Towards understanding the spectral bias of deep learning. arXiv 2019, arXiv:1912.01198. [Google Scholar]
  19. Fridovich-Keil, S.; Gontijo Lopes, R.; Roelofs, R. Spectral bias in practice: The role of function frequency in generalization. Adv. Neural Inf. Process. Syst. 2022, 35, 7368–7382. [Google Scholar] [CrossRef] [Scilit]
  20. Lindell, D.B.; Van Veen, D.; Park, J.J.; Wetzstein, G. Bacon: Band-limited coordinate networks for multiscale scene representation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 16252–16262. [Google Scholar]
  21. Saragadam, V.; Tan, J.; Balakrishnan, G.; Baraniuk, R.G.; Veeraraghavan, A. Miner: Multiscale implicit neural representation. In Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel, 23–27 October 2022; pp. 318–333. [Google Scholar]
  22. Shabanov, A.; Govindarajan, S.; Reading, C.; Goli, L.; Rebain, D.; Yi, K.M.; Tagliasacchi, A. BANF: Band-limited Neural Fields for Levels of Detail Reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 20571–20580. [Google Scholar]
  23. Xu, Q.; Xu, Z.; Philip, J.; Bi, S.; Shu, Z.; Sunkavalli, K.; Neumann, U. Point-nerf: Point-based neural radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 5438–5448. [Google Scholar]
  24. Chen, Y.; Jia, G.; Zha, Y.; Zhang, P.; Zhang, Y. LINR: A Plug-and-Play Local Implicit Neural Representation Module for Visual Object Tracking. IEEE Trans. Circuits Syst. Video Technol. 2025, 35, 12652–12665. [Google Scholar] [CrossRef] [Scilit]
  25. Peng, S.; Niemeyer, M.; Mescheder, L.; Pollefeys, M.; Geiger, A. Convolutional occupancy networks. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2020; pp. 523–540. [Google Scholar]
  26. Chabra, R.; Lenssen, J.E.; Ilg, E.; Schmidt, T.; Straub, J.; Lovegrove, S.; Newcombe, R. Deep local shapes: Learning local sdf priors for detailed 3d reconstruction. In Proceedings of the Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, 23–28 August 2020; Proceedings, Part XXIX 16; Springer: Berlin/Heidelberg, Germany, 2020; pp. 608–625. [Google Scholar]
  27. Yu, A.; Ye, V.; Tancik, M.; Kanazawa, A. Pixelnerf: Neural radiance fields from one or few images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 19–25 June 2021; pp. 4578–4587. [Google Scholar]
  28. Reiser, C.; Peng, S.; Liao, Y.; Geiger, A. Kilonerf: Speeding up neural radiance fields with thousands of tiny mlps. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 19–25 June 2021; pp. 14335–14345. [Google Scholar]
  29. Wu, L.; Lee, J.Y.; Bhattad, A.; Wang, Y.X.; Forsyth, D. Diver: Real-time and accurate neural radiance fields with deterministic integration for volume rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 16200–16209. [Google Scholar]
  30. Anokhin, I.; Demochkin, K.; Khakhulin, T.; Sterkin, G.; Lempitsky, V.; Korzhenkov, D. Image generators with conditionally-independent pixel synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 19–25 June 2021; pp. 14278–14287. [Google Scholar]
  31. Chen, Y.; Wang, X. Transformers as meta-learners for implicit neural representations. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2022; pp. 170–187. [Google Scholar]
  32. Zhang, S.; Liu, K.; Gu, J.; Cai, X.; Wang, Z.; Bu, J.; Wang, H. Attention beats linear for fast implicit neural representation generation. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2024; pp. 1–18. [Google Scholar]
  33. Byra, M.; Poon, C.; Rachmadi, M.F.; Schlachter, M.; Skibbe, H. Exploring the performance of implicit neural representations for brain image registration. Sci. Rep. 2023, 13, 17334. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  34. Park, K.; Sinha, U.; Barron, J.T.; Bouaziz, S.; Goldman, D.B.; Seitz, S.M.; Martin-Brualla, R. Nerfies: Deformable Neural Radiance Fields. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 10–17 October 2020; pp. 5845–5854. [Google Scholar]
  35. Li, Z.; Niklaus, S.; Snavely, N.; Wang, O. Neural Scene Flow Fields for Space-Time View Synthesis of Dynamic Scenes. In Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021; pp. 6494–6504. [Google Scholar]
  36. Mildenhall, B.; Srinivasan, P.P.; Tancik, M.; Barron, J.T.; Ramamoorthi, R.; Ng, R. Nerf: Representing scenes as neural radiance fields for view synthesis. Commun. ACM 2021, 65, 99–106. [Google Scholar]
  37. Shi, K.; Zhou, X.; Gu, S. Improved Implicit Neural Representation with Fourier Reparameterized Training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 25985–25994. [Google Scholar]
  38. Novello, T.; Aldana, D.; Araujo, A.; Velho, L. Tuning the Frequencies: Robust Training for Sinusoidal Neural Networks. In Proceedings of the Computer Vision and Pattern Recognition Conference, Nashville, TN, USA, 11–15 June 2025; pp. 3071–3080. [Google Scholar]
  39. Fathony, R.; Sahu, A.K.; Willmott, D.; Kolter, J.Z. Multiplicative filter networks. In Proceedings of the International Conference on Learning Representations, Virtual, 26 April–1 May 2020. [Google Scholar]
  40. Liu, Z.; Zhu, H.; Zhang, Q.; Fu, J.; Deng, W.; Ma, Z.; Guo, Y.; Cao, X. Finer: Flexible spectral-bias tuning in implicit neural representation by variable-periodic activation functions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 2713–2722. [Google Scholar]
  41. Kazerouni, A.; Azad, R.; Hosseini, A.; Merhof, D.; Bagci, U. INCODE: Implicit Neural Conditioning with Prior Knowledge Embeddings. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA, 3–8 January 2024; pp. 1298–1307. [Google Scholar]
  42. Xu, Q.; Wang, W.; Ceylan, D.; Mech, R.; Neumann, U. Disn: Deep implicit surface network for high-quality single-view 3d reconstruction. Adv. Neural Inf. Process. Syst. 2019, 32. Available online: https://proceedings.neurips.cc/paper/2019/hash/39059724f73a9969845dfe4146c5660e-Abstract.html (accessed on 1 September 2026).
  43. Liu, K.; Liu, F.; Wang, H.; Ma, N.; Bu, J.; Han, B. Partition Speeds Up Learning Implicit Neural Representations Based on Exponential-Increase Hypothesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 5474–5483. [Google Scholar]
  44. Hao, Z.; Mallya, A.; Belongie, S.; Liu, M.Y. Implicit neural representations with levels-of-experts. Adv. Neural Inf. Process. Syst. 2022, 35, 2564–2576. [Google Scholar] [CrossRef] [Scilit]
  45. Martel, J.N.P.; Lindell, D.B.; Lin, C.Z.; Chan, E.R.; Monteiro, M.; Wetzstein, G. ACORN: Adaptive coordinate networks for neural scene representation. ACM Trans. Graph. (SIGGRAPH) 2021, 40, 58. [Google Scholar]
  46. Takikawa, T.; Litalien, J.; Yin, K.; Kreis, K.; Loop, C.; Nowrouzezahrai, D.; Jacobson, A.; McGuire, M.; Fidler, S. Neural geometric level of detail: Real-time rendering with implicit 3d shapes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 19–25 June 2021; pp. 11358–11367. [Google Scholar]
  47. Müller, T.; Evans, A.; Schied, C.; Keller, A. Instant neural graphics primitives with a multiresolution hash encoding. ACM Trans. Graph. (TOG) 2022, 41, 1–15. [Google Scholar] [CrossRef] [Scilit]
  48. Zheng, S.; Zhang, C.; Han, D.; Puspitasari, F.D.; Hao, X.; Yang, Y.; Shen, H.T. Exploring Kernel Transformations for Implicit Neural Representations. arXiv 2025, arXiv:2504.04728. [Google Scholar]
  49. Li, J.C.L.; Liu, C.; Huang, B.; Wong, N. Learning spatially collaged fourier bases for implicit neural representation. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI Press: Washington, DC, USA, 2024; Volume 38, pp. 13492–13499. [Google Scholar]
  50. Eastman Kodak Company. Kodak Lossless True Color Image Suite. 1999. Available online: http://r0k.us/graphics/kodak/ (accessed on 13 January 2026).
  51. Lim, B.; Son, S.; Kim, H.; Nah, S.; Mu Lee, K. Enhanced deep residual networks for single image super-resolution. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Honolulu, HI, USA, 21–26 July 2017; pp. 136–144. [Google Scholar]
  52. Clark, K.; Vendt, B.; Smith, K.; Freymann, J.; Kirby, J.; Koppel, P.; Moore, S.; Phillips, S.; Maffitt, D.; Pringle, M.; et al. The Cancer Imaging Archive (TCIA): Maintaining and operating a public information repository. J. Digit. Imaging 2013, 26, 1045–1057. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Figure 1. Overview of the proposed SA-INR framework. Given input coordinates, a coordinate-based MLP first maps each query location to a latent feature. The proposed structure-aware implicit operator introduces local interactions in the coordinate domain by aggregating neighboring latent features with a learnable neighborhood operator. The aggregated context is then combined with the original point-wise feature through a residual pathway and further modulates the sinusoidal activation for signal reconstruction.
Figure 1. Overview of the proposed SA-INR framework. Given input coordinates, a coordinate-based MLP first maps each query location to a latent feature. The proposed structure-aware implicit operator introduces local interactions in the coordinate domain by aggregating neighboring latent features with a learnable neighborhood operator. The aggregated context is then combined with the original point-wise feature through a residual pathway and further modulates the sinusoidal activation for signal reconstruction.
Applsci 16 08870 g001
Figure 2. Qualitative comparison of image representation. Compared with the baselines, SA-INR better preserves high-frequency structural details and reduces ringing artifacts in smooth regions.
Figure 2. Qualitative comparison of image representation. Compared with the baselines, SA-INR better preserves high-frequency structural details and reduces ringing artifacts in smooth regions.
Applsci 16 08870 g002
Figure 3. Qualitative comparison on high-resolution images. SA-INR demonstrates strong scalability, reconstructing complex textures while avoiding over-smoothing.
Figure 3. Qualitative comparison on high-resolution images. SA-INR demonstrates strong scalability, reconstructing complex textures while avoiding over-smoothing.
Applsci 16 08870 g003
Figure 4. Qualitative comparison of sparse-view CT reconstruction with 150 angles. Standard INRs exhibit noticeable streak artifacts due to the underdetermined nature of the problem, while SA-INR produces reconstructions with reduced artifacts and improved anatomical consistency.
Figure 4. Qualitative comparison of sparse-view CT reconstruction with 150 angles. Standard INRs exhibit noticeable streak artifacts due to the underdetermined nature of the problem, while SA-INR produces reconstructions with reduced artifacts and improved anatomical consistency.
Applsci 16 08870 g004
Figure 5. Qualitative comparison of image denoising. SA-INR leverages local neighborhood consistency to suppress noise while preserving fine-grained structures.
Figure 5. Qualitative comparison of image denoising. SA-INR leverages local neighborhood consistency to suppress noise while preserving fine-grained structures.
Applsci 16 08870 g005aApplsci 16 08870 g005b
Table 1. Quantitative comparison of image representation. SA-INR improves image fitting performance across multiple MLP backbones.
Table 1. Quantitative comparison of image representation. SA-INR improves image fitting performance across multiple MLP backbones.
MethodsSIRENSA-SIRENFINERSA-FINERFR-INRSA-FR-INR
PSNR34.6237.1237.9140.8840.4445.25
SSIM0.93020.95830.96390.97880.97650.9900
Table 2. Quantitative comparison of large-scale image representation.
Table 2. Quantitative comparison of large-scale image representation.
MethodsSIRENSA-SIRENFINERSA-FINERFR-INRSA-FR-INR
PSNR26.3727.7928.1728.9129.6630.74
SSIM0.74300.78900.80270.82070.82930.8771
Table 3. Quantitative comparison of sparse-view CT reconstruction. Incorporating the spatial-aware module improves reconstruction performance under severely limited measurements.
Table 3. Quantitative comparison of sparse-view CT reconstruction. Incorporating the spatial-aware module improves reconstruction performance under severely limited measurements.
MethodsSIRENSA-SIRENFINERSA-FINERFR-INRSA-FR-INR
PSNR29.2332.4430.7831.8830.8432.78
SSIM0.86080.90750.89130.89710.88870.9222
Table 4. Quantitative comparison for image denoising. The results indicate that modeling spatial coherence improves the ability to recover clean signals from noisy observations.
Table 4. Quantitative comparison for image denoising. The results indicate that modeling spatial coherence improves the ability to recover clean signals from noisy observations.
MethodsSIRENSA-SIRENFINERSA-FINERFR-INRSA-FR-INR
PSNR25.5327.0726.7527.0126.8727.59
SSIM0.72460.73830.75860.74770.76770.7720
Table 5. Impact of local aggregation designs on image fitting.
Table 5. Impact of local aggregation designs on image fitting.
Model VariantPSNRSSIM
Vanilla SIREN34.620.9302
+ Dense cross-channel aggregation26.240.7002
+ Channel-independent local aggregation35.300.9394
Table 6. Effect of the residual point-wise pathway.
Table 6. Effect of the residual point-wise pathway.
Model VariantPSNRSSIM
Without residual connection35.300.9394
With residual connection (SA-INR)37.120.9583
Table 7. Effect of mean-filter initialization.
Table 7. Effect of mean-filter initialization.
Model VariantPSNRSSIM
Random initialization36.600.9404
Mean-filter initialization (SA-INR)37.120.9583
Table 8. Effect of local neighborhood size in SA-INR.
Table 8. Effect of local neighborhood size in SA-INR.
Model VariantPSNRSSIM
w/o Spatial35.770.9386
3 × 3 37.120.9583
5 × 5 36.740.9480
7 × 7 36.400.9401
Table 9. Computational cost comparison between SIREN and SA-SIREN under the same experimental setting.
Table 9. Computational cost comparison between SIREN and SA-SIREN under the same experimental setting.
MethodParams. (M)Train. Time (s)Infer. Time (s)Memory (GB)
SIREN0.19913.110.004680.19426
SA-SIREN0.20115.120.006030.19430
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Qing, C.; Zhang, W.; Wang, H.; Han, D.; Zhang, M.; Qin, C. Spatial-Aware Modulation for Implicit Neural Representations. Appl. Sci. 2026, 16, 8870. https://doi.org/10.3390/app16178870

AMA Style

Qing C, Zhang W, Wang H, Han D, Zhang M, Qin C. Spatial-Aware Modulation for Implicit Neural Representations. Applied Sciences. 2026; 16(17):8870. https://doi.org/10.3390/app16178870

Chicago/Turabian Style

Qing, Chen, Wenxin Zhang, Haoyu Wang, Dongshen Han, Mingming Zhang, and Caiyan Qin. 2026. "Spatial-Aware Modulation for Implicit Neural Representations" Applied Sciences 16, no. 17: 8870. https://doi.org/10.3390/app16178870

APA Style

Qing, C., Zhang, W., Wang, H., Han, D., Zhang, M., & Qin, C. (2026). Spatial-Aware Modulation for Implicit Neural Representations. Applied Sciences, 16(17), 8870. https://doi.org/10.3390/app16178870

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop