Lightweight SNR-Adaptive Receiver-Side Enhancement for DeepJSCC-Based Wireless Image Transmission
Abstract
1. Introduction
- We propose the first receiver-only plug-and-play perceptual enhancer for deployed DeepJSCC systems. The proposed approach maintains compatibility with existing infrastructure by keeping both the encoder and decoder frozen, while introducing only 0.29 million trainable parameters (<1% of the backbone).
- We design an SNR-adaptive residual enhancement mechanism based on FiLM modulation [34], enabling a single lightweight model to dynamically adjust restoration strength according to channel quality. Under low-SNR conditions, the mechanism applies conservative enhancement to reduce the risk of noise amplification, whereas under high-SNR conditions it enables stronger high-frequency recovery.
- We introduce a radially weighted FFT magnitude loss that provides spectral-domain supervision complementary to spatial perceptual losses. This frequency-targeted guidance directly addresses the high-frequency suppression commonly observed in MSE-trained DeepJSCC models.
- We conduct comprehensive experiments on the Kodak24 and DIV2K datasets under AWGN channels, demonstrating consistent LPIPS improvements with 34.4–37.5% relative reduction, while incurring minimal computational overhead, including 29.8 ms enhancer-only latency, 216 GFLOPs, and modest receiver-side deployment cost.
2. Related Works
2.1. Deep Learning for Joint Source-Channel Coding
2.2. Perceptual Quality Optimization in Image Transmission
2.3. Adaptive Neural Networks for Wireless Communications
3. Methods
3.1. System Model
3.2. Network Architecture
3.3. Training Objective
3.4. Training Protocol
| Algorithm 1 Training and Inference Procedure of the Proposed SNR-Adaptive Residual Enhancement Framework |
Input: Pretrained and frozen DeepJSCC encoder–decoder ; training set ; SNR set dB; loss weights Output: Trained enhancer and enhanced image
|
4. Experimental Results
4.1. Experimental Setup
- PSNR (dB): Peak signal-to-noise ratio, which measures pixel-level distortion. Higher values indicate lower distortion.
- LPIPS [21]: Learned perceptual image patch similarity, which measures perceptual discrepancy. Lower values indicate better perceptual quality.
- MS-SSIM [58]: Multi-scale structural similarity, which evaluates structural fidelity. Higher values indicate better structural preservation.
- Cross-backbone reference: DiffJSCC [29], which relies on its own independently trained backbone and diffusion-based refinement. It is included as a reference rather than a strictly controlled baseline.
4.2. Main Results
4.3. Computational Efficiency
4.4. Frequency Analysis
- Band 6: 0.04% → 39.8% (995× increase);
- Band 7: 0.06% → 50.6% (843× increase);
- Band 8: 0.08% → 41.7% (521× increase).
4.5. Visual Comparison
4.6. Ablation Study
4.7. Generalization to DIV2K Valid HR
4.8. Compatibility Across CBR Configurations
5. Discussion
5.1. Implications for Deployed Wireless Systems
5.2. Robustness to Rayleigh Fading Mismatch
5.3. SNR Estimation Robustness
5.4. Perception–Distortion Tradeoff
5.5. Comparison with Alternative Approaches
5.6. Latency and Deployment Analysis
5.7. Limitations and Future Work
- Channel model. The proposed enhancer has been evaluated under AWGN, Rayleigh fading, and Rician fading (K = 3, 10) channels, demonstrating robust performance across these conditions. However, more complex fading scenarios, including frequency-selective fading, mobility-induced time variation, and multi-path channels requiring OFDM-based models, remain to be systematically evaluated. These scenarios may require modifications to the single-carrier DeepJSCC framework and are left as future work.
- CBR generalization. The current enhancer is trained for specific CBR configurations (C2, C4, C8), and separate enhancers are trained for each CBR. Developing a CBR-agnostic enhancer through explicit CBR conditioning—e.g., incorporating CBR as an additional input to the FiLM modulation network—is an important direction for future work.
- Video extension. The measured latency suggests potential for near-real-time video applications. However, temporal consistency mechanisms are needed to prevent flickering artifacts across frames.
- Backbone diversity. The proposed method is evaluated on DeepJSCC backbones with different CBR configurations of the same architecture family. Compatibility with structurally different JSCC backbones, such as transformer-based or GAN-based architectures, remains to be further investigated.
- SNR estimation. The current framework assumes that SNR information is available at the receiver. We have evaluated robustness to SNR estimation errors up to dB at multiple true SNR levels (1, 7, 13 dB), showing graceful degradation. In practical fast-fading systems, SNR may vary rapidly within a frame; integrating dedicated channel estimation and prediction modules into the enhancement pipeline is an important direction for future work.
- NPU deployment. We have demonstrated real-time enhancer inference on a Huawei Ascend 910B NPU (21.7 ms, 46.2 FPS). However, evaluation on additional edge NPU platforms (e.g., mobile-grade NPUs in smartphones), as well as INT8 quantization and hardware-specific operator optimization for further latency reduction, remain as future work.
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AWGN | Additive White Gaussian Noise |
| CBR | Channel Bandwidth Ratio |
| DCT | Discrete Cosine Transform |
| DeepJSCC | Deep Joint Source-Channel Coding |
| FFT | Fast Fourier Transform |
| FiLM | Feature-wise Linear Modulation |
| GAN | Generative Adversarial Network |
| LPIPS | Learned Perceptual Image Patch Similarity |
| MLP | Multi-Layer Perceptron |
| MSE | Mean Squared Error |
| MS-SSIM | Multi-Scale Structural Similarity |
| PSNR | Peak Signal-to-Noise Ratio |
| SNR | Signal-to-Noise Ratio |
References
- Shannon, C.E. A mathematical theory of communication. Bell Syst. Tech. J. 1948, 27, 379–423. [Google Scholar] [CrossRef] [Scilit]
- Dai, J.; Zhang, P.; Niu, K.; Wang, S.; Si, Z.; Qin, X. Communication beyond transmitting bits: Semantics-guided source and channel coding. IEEE Wirel. Commun. 2023, 30, 170–177. [Google Scholar] [CrossRef] [Scilit]
- Yang, W.; Du, H.; Liew, Z.Q.; Lim, W.Y.B.; Xiong, Z.; Niyato, D.; Chi, X.; Shen, X.; Miao, C. Semantic communications for future Internet: Fundamentals, applications, and challenges. IEEE Commun. Surv. Tutor. 2023, 25, 213–250. [Google Scholar] [CrossRef] [Scilit]
- Strinati, E.C.; Barbarossa, S.; Gonzalez-Jimenez, J.L.; Ktenas, D.; Cassiau, N.; Mokdad, L.; Visoz, R. 6G: The next frontier: From holographic messaging to artificial intelligence using subterahertz and visible light communication. IEEE Veh. Technol. Mag. 2019, 14, 42–50. [Google Scholar] [CrossRef] [Scilit]
- Getu, T.M.; Kaddoum, G.; Bennis, M. Semantic communication: A survey on research landscape, challenges, and future directions. Proc. IEEE 2024, 112, 1649–1685. [Google Scholar] [CrossRef] [Scilit]
- Chowdhury, M.Z.; Shahjalal, M.; Ahmed, S.; Jang, Y.M. 6G wireless communication systems: Applications, requirements, technologies, challenges, and research directions. IEEE Open J. Commun. Soc. 2020, 1, 957–975. [Google Scholar] [CrossRef] [Scilit]
- Bourtsoulatze, E.; Burth Kurka, D.; Gündüz, D. Deep joint source-channel coding for wireless image transmission. IEEE Trans. Cogn. Commun. Netw. 2019, 5, 567–579. [Google Scholar] [CrossRef] [Scilit]
- Xie, H.; Qin, Z.; Tao, X.; Li, K. Deep learning enabled semantic communication systems. IEEE Trans. Signal Process. 2021, 69, 2663–2675. [Google Scholar] [CrossRef] [Scilit]
- Jin, Z.; Song, T.; Jia, W.-K.; Zou, W.; Song, X. Task-oriented semantic communication with adaptive semantic reconstruction network. IEEE Internet Things J. 2025, 12, 35784–35798. [Google Scholar] [CrossRef] [Scilit]
- O’Shea, T.J.; Hoydis, J. An introduction to deep learning for the physical layer. IEEE Trans. Cogn. Commun. Netw. 2017, 3, 563–575. [Google Scholar] [CrossRef] [Scilit]
- Dörner, S.; Cammerer, S.; Hoydis, J.; ten Brink, S. Deep learning based communication over the air. IEEE J. Sel. Top. Signal Process. 2018, 12, 132–143. [Google Scholar] [CrossRef] [Scilit]
- Dai, L.; Jiao, R.; Adachi, F.; Poor, H.V.; Hanzo, L. Deep learning for wireless communications: An emerging interdisciplinary paradigm. IEEE Wirel. Commun. 2020, 27, 133–139. [Google Scholar] [CrossRef] [Scilit]
- Xu, J.; Ai, B.; Chen, W.; Yang, A.; Sun, P.; Rodrigues, M. Wireless image transmission using deep source channel coding with attention modules. IEEE Trans. Circuits Syst. Video Technol. 2022, 32, 2315–2328. [Google Scholar] [CrossRef] [Scilit]
- Yang, K.; Wang, S.; Dai, J.; Tan, K.; Niu, K.; Zhang, P. WITT: A wireless image transmission transformer for semantic communications. In Proceedings of the ICASSP 2023–2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, 4–10 June 2023; IEEE: Piscataway, NJ, USA, 2023; pp. 1–5. [Google Scholar] [CrossRef] [Scilit]
- Yang, K.; Wang, S.; Dai, J.; Qin, X.; Niu, K.; Zhang, P. SwinJSCC: Taming Swin transformer for deep joint source-channel coding. IEEE Trans. Cogn. Commun. Netw. 2025, 11, 90–104. [Google Scholar] [CrossRef] [Scilit]
- Blau, Y.; Michaeli, T. The perception-distortion tradeoff. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; IEEE: Piscataway, NJ, USA, 2018; pp. 6228–6237. [Google Scholar] [CrossRef] [Scilit]
- Mao, Q.; Hu, F.; Hao, Q. Deep learning for intelligent wireless networks: A comprehensive survey. IEEE Commun. Surv. Tutor. 2018, 20, 2595–2621. [Google Scholar] [CrossRef] [Scilit]
- Ledig, C.; Theis, L.; Huszár, F.; Caballero, J.; Cunningham, A.; Acosta, A.; Aitken, A.; Tejani, A.; Totz, J.; Wang, Z.; et al. Photo-realistic single image super-resolution using a generative adversarial network. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; IEEE: Piscataway, NJ, USA, 2017; pp. 105–114. [Google Scholar] [CrossRef] [Scilit]
- Wang, X.; Yu, K.; Wu, S.; Gu, J.; Liu, Y.; Dong, C.; Loy, C.C.; Qiao, Y.; Tang, X. ESRGAN: Enhanced super-resolution generative adversarial networks. arXiv 2018, arXiv:1809.00219. [Google Scholar] [CrossRef] [Scilit]
- Lim, B.; Son, S.; Kim, H.; Nah, S.; Lee, K.M. Enhanced deep residual networks for single image super-resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW); IEEE: Piscataway, NJ, USA, 2017; pp. 136–144. [Google Scholar] [CrossRef] [Scilit]
- Zhang, R.; Isola, P.; Efros, A.A.; Shechtman, E.; Wang, O. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; IEEE: Piscataway, NJ, USA, 2018; pp. 586–595. [Google Scholar] [CrossRef] [Scilit]
- Johnson, J.; Alahi, A.; Fei-Fei, L. Perceptual losses for real-time style transfer and super-resolution. In Proceedings of the European Conference on Computer Vision (ECCV); Springer: Cham, Switzerland, 2016; pp. 694–711. [Google Scholar] [CrossRef] [Scilit]
- Ding, K.; Ma, K.; Wang, S.; Simoncelli, E.P. Image quality assessment: Unifying structure and texture similarity. IEEE Trans. Pattern Anal. Mach. Intell. 2020, 44, 2567–2581. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, J.; Wang, S.; Dai, J.; Si, Z.; Zhou, D.; Niu, K. Perceptual learned source-channel coding for high-fidelity image semantic transmission. In Proceedings of the GLOBECOM 2022-2022 IEEE Global Communications Conference, Rio de Janeiro, Brazil, 4–8 December 2022; pp. 3959–3964. [Google Scholar] [CrossRef] [Scilit]
- Ho, J.; Jain, A.; Abbeel, P. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems (NeurIPS); Neural Information Processing Systems Foundation: San Diego, CA, USA, 2020; Volume 33, pp. 6840–6851. Available online: https://proceedings.neurips.cc/paper/2020/hash/4c5bcfec8584af0d967f1ab10179ca4b-Abstract.html (accessed on 1 June 2026).
- Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2022; pp. 10684–10695. [Google Scholar] [CrossRef] [Scilit]
- Saharia, C.; Ho, J.; Chan, W.; Salimans, T.; Fleet, D.J.; Norouzi, M. Image super-resolution via iterative refinement. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 4713–4726. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Dhariwal, P.; Nichol, A. Diffusion models beat GANs on image synthesis. In Advances in Neural Information Processing Systems (NeurIPS); Neural Information Processing Systems Foundation: San Diego, CA, USA, 2021; Volume 34, pp. 8780–8794. Available online: https://proceedings.neurips.cc/paper/2021/hash/49ad23d1ec9fa4bd8d77d02681df5cfa-Abstract.html (accessed on 1 June 2026).
- Yang, M.; Liu, B.; Wang, B.; Kim, H.-S. Diffusion-aided joint source channel coding for high realism wireless image transmission. IEEE Trans. Mach. Learn. Commun. Netw. 2025, 3, 1227–1243. [Google Scholar] [CrossRef] [Scilit]
- Zhang, K.; Zuo, W.; Gu, S.; Zhang, L. Learning deep CNN denoiser prior for image restoration. arXiv 2017, arXiv:1704.03264. [Google Scholar] [CrossRef] [Scilit]
- Chen, L.; Chu, X.; Zhang, X.; Sun, J. NAFNet: Simple baselines for image restoration. In Proceedings of the European Conference on Computer Vision Workshops (ECCVW); Springer: Cham, Switzerland, 2022; pp. 17–33. [Google Scholar] [CrossRef] [Scilit]
- Liang, J.; Cao, J.; Sun, G.; Zhang, K.; Van Gool, L.; Timofte, R. SwinIR: Image restoration using Swin Transformer. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW); IEEE: Piscataway, NJ, USA, 2021; pp. 1833–1844. [Google Scholar] [CrossRef] [Scilit]
- Mei, Y.; Fan, Y.; Zhou, Y.; Huang, L.; Huang, T.S.; Shi, H. Image super-resolution with cross-scale non-local attention and exhaustive self-exemplars mining. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2020; pp. 5689–5698. [Google Scholar] [CrossRef] [Scilit]
- Perez, E.; Strub, F.; de Vries, H.; Dumoulin, V.; Courville, A. FiLM: Visual reasoning with a general conditioning layer. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI: Washington, DC, USA, 2018; Volume 32, pp. 3942–3951. [Google Scholar] [CrossRef] [Scilit]
- O’Shea, T.; Erpek, T.; Clancy, T.C. Deep learning based MIMO communications. arXiv 2017, arXiv:1707.07980. [Google Scholar] [CrossRef] [Scilit]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, L.; Polosukhin, I. Attention is all you need. In Advances in Neural Information Processing Systems (NeurIPS); Neural Information Processing Systems Foundation: San Diego, CA, USA, 2017; Volume 30, pp. 5998–6008. Available online: https://proceedings.neurips.cc/paper/2017/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html (accessed on 1 June 2026).
- Tung, T.-Y.; Kurka, D.B.; Jankowski, M.; Gündüz, D. DeepJSCC-Q: Constellation constrained deep joint source-channel coding. IEEE J. Sel. Areas Inf. Theory 2022, 3, 720–731. [Google Scholar] [CrossRef] [Scilit]
- Huang, Y.; Qin, Z. Wireless video transmission with joint semantic-channel coding. In Proceedings of the IEEE Globecom Workshops (GC Wkshps); IEEE: Piscataway, NJ, USA, 2024; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
- Liu, S.; Gao, Z.; Chen, G.; Su, Y.; Peng, L. Transformer-based joint source channel coding for textual semantic communication. In Proceedings of the IEEE/CIC International Conference on Communications in China (ICCC); IEEE: Piscataway, NJ, USA, 2023; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
- Xie, H.; Qin, Z.; Tao, X.; Letaief, K.B. Task-oriented multi-user semantic communications. IEEE J. Sel. Areas Commun. 2022, 40, 2584–2597. [Google Scholar] [CrossRef] [Scilit]
- Shao, J.; Mao, Y.; Zhang, J. Task-oriented communication for multidevice cooperative edge inference. IEEE Trans. Wirel. Commun. 2023, 22, 73–87. [Google Scholar] [CrossRef] [Scilit]
- Xiong, L.; Yu, P.; Wu, Y. MADPHash: Manipulation-aware deep perceptual hashing using feature consistency. In Proceedings of the 33rd ACM International Conference on Multimedia; ACM: New York, NY, USA, 2025; pp. 8038–8047. [Google Scholar] [CrossRef] [Scilit]
- Woo, S.; Park, J.; Lee, J.-Y.; Kweon, I.S. CBAM: Convolutional block attention module. In Proceedings of the European Conference on Computer Vision (ECCV); Springer: Cham, Switzerland, 2018; pp. 3–19. [Google Scholar] [CrossRef] [Scilit]
- Zamir, S.W.; Arora, A.; Khan, S.; Hayat, M.; Khan, F.S.; Yang, M.-H. Restormer: Efficient transformer for high-resolution image restoration. arXiv 2022, arXiv:2111.09881. [Google Scholar] [CrossRef] [Scilit]
- Liang, J.; Cao, J.; Fan, Y.; Zhang, K.; Ranjan, R.; Li, Y.; Timofte, R.; Van Gool, L. VRT: A video restoration transformer. IEEE Trans. Image Process. 2024, 33, 2171–2182. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Cui, Y.; Ren, W.; Cao, X.; Knoll, A. Revitalizing convolutional network for image restoration. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 46, 9423–9438. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kawar, B.; Elad, M.; Ermon, S.; Song, J. Denoising diffusion restoration models. arXiv 2022, arXiv:2201.11793. [Google Scholar] [CrossRef] [Scilit]
- Choi, J.; Kim, S.; Jeong, Y.; Gwon, Y.; Yoon, S. ILVR: Conditioning method for denoising diffusion probabilistic models. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: Piscataway, NJ, USA, 2021; pp. 14347–14356. [Google Scholar] [CrossRef] [Scilit]
- Meng, C.; He, Y.; Song, Y.; Song, J.; Wu, J.; Zhu, J.-Y.; Ermon, S. SDEdit: Guided image synthesis and editing with stochastic differential equations. arXiv 2022, arXiv:2108.01073. [Google Scholar] [CrossRef] [Scilit]
- Ren, M.; Qiao, L.; Yang, L.; Gao, Z.; Chen, J.; Mashhadi, M.B.; Xiao, P.; Tafazolli, R.; Bennis, M. Generative semantic communication via textual prompts: Latency performance tradeoffs. IEEE Trans. Veh. Technol. 2025, 74, 14843–14848. [Google Scholar] [CrossRef] [Scilit]
- Hello, N.; Di Lorenzo, P.; Calvanese Strinati, E. Semantic communication enhanced by knowledge graph representation learning. arXiv 2024, arXiv:2407.19338. [Google Scholar] [CrossRef] [Scilit]
- Dumoulin, V.; Perez, E.; Schucher, N.; Strub, F.; de Vries, H.; Courville, A.; Bengio, Y. Feature-wise transformations. Distill 2018, 3, e11. [Google Scholar] [CrossRef] [Scilit]
- Ba, J.L.; Kiros, J.R.; Hinton, G.E. Layer normalization. arXiv 2016, arXiv:1607.06450. [Google Scholar] [CrossRef] [Scilit]
- Lin, T.-Y.; Dollár, P.; Girshick, R.; He, K.; Hariharan, B.; Belongie, S. Feature pyramid networks for object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2017; pp. 2117–2125. [Google Scholar] [CrossRef] [Scilit]
- Zhu, G.; Lyu, Z.; Jiao, X.; Liu, P.; Chen, M.; Xu, J.; Cui, S.; Zhang, P. Pushing AI to wireless network edge: An overview on integrated sensing, communication, and computation towards 6G. Sci. China Inf. Sci. 2023, 66, 130301. [Google Scholar] [CrossRef] [Scilit]
- Hu, J.; Shen, L.; Sun, G. Squeeze-and-excitation networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2018; pp. 7132–7141. [Google Scholar] [CrossRef] [Scilit]
- Kingma, D.P.; Ba, J. Adam: A method for stochastic optimization. arXiv 2015, arXiv:1412.6980. [Google Scholar] [CrossRef] [Scilit]
- Wang, Z.; Bovik, A.C.; Sheikh, H.R.; Simoncelli, E.P. Image quality assessment: From error visibility to structural similarity. IEEE Trans. Image Process. 2004, 13, 600–612. [Google Scholar] [CrossRef] [Scilit] [PubMed]





| Component | Parameters | Percentage |
|---|---|---|
| Feature extraction head | 4416 | 1.5% |
| Residual blocks (6 blocks) | 256,608 | 88.5% |
| FiLM MLP | 4256 | 1.5% |
| FiLM linear heads | 24,576 | 8.5% |
| Reconstruction tail | 1315 | 0.5% |
| Total | 291,171 | 100% |
| Method | 1 dB | 4 dB | 7 dB | 10 dB | 13 dB |
|---|---|---|---|---|---|
| DeepJSCC | 22.13 | 22.86 | 23.39 | 23.75 | 23.98 |
| DiffJSCC † | 21.26 | 21.98 | 22.41 | 22.74 | 23.09 |
| DnCNN * | 21.80 | 22.37 | 22.76 | 23.02 | 23.18 |
| Ours * | 21.66 | 22.31 | 22.82 | 23.20 | 23.43 |
| Method | 1 dB | 4 dB | 7 dB | 10 dB | 13 dB |
|---|---|---|---|---|---|
| DeepJSCC | 0.566 | 0.523 | 0.496 | 0.481 | 0.472 |
| DiffJSCC † | 0.449 | 0.412 | 0.372 | 0.354 | 0.335 |
| DnCNN * | 0.478 | 0.426 | 0.397 | 0.378 | 0.369 |
| Ours * | 0.371 | 0.340 | 0.317 | 0.302 | 0.295 |
| Method | 1 dB | 4 dB | 7 dB | 10 dB | 13 dB |
|---|---|---|---|---|---|
| DeepJSCC | 0.718 | 0.766 | 0.798 | 0.819 | 0.831 |
| DiffJSCC † | 0.677 | 0.726 | 0.757 | 0.780 | 0.797 |
| DnCNN * | 0.705 | 0.747 | 0.775 | 0.794 | 0.805 |
| Ours * | 0.687 | 0.734 | 0.769 | 0.791 | 0.805 |
| Method | Parameters | FLOPs | Steps | LPIPS | ||
|---|---|---|---|---|---|---|
| Total | Add. | Total | Add. | |||
| DeepJSCC | 31.2 M | – | 165 G | – | 1 | 0.496 |
| DiffJSCC † | 1734 M | 1705 M | N/A | N/A | 40 | 0.372 |
| DnCNN * | 31.4 M | 0.21 M | 331 G | 166 G | 1 | 0.397 |
| Ours * | 31.5 M | 0.29 M | 381 G | 216 G | 1 | 0.317 |
| Configuration | 1 dB | 7 dB | 13 dB | Avg. |
|---|---|---|---|---|
| Full Model | 0.372 | 0.316 | 0.295 | baseline |
| w/o LPIPS Loss | +45.1% | +48.9% | +51.5% | +48.5% |
| w/o Residual | +6.9% | +8.0% | +9.4% | +8.1% |
| w/o SNR Cond. | +2.3% | −0.8% | −0.8% | +0.3% |
| w/o Freq. Loss | +0.6% | +0.5% | +1.2% | +0.7% |
| w/o Attention | +0.3% | −0.4% | +0.1% | +0.0% |
| C | N | a | LPIPS | PSNR (dB) | Params |
|---|---|---|---|---|---|
| 24 | 3 | 0.5 | 0.317 | 22.98 | 0.08 M |
| 24 | 6 | 0.5 | 0.310 | 22.97 | 0.15 M |
| 24 | 9 | 0.5 | 0.310 | 22.84 | 0.23 M |
| 48 | 3 | 0.5 | 0.311 | 22.92 | 0.15 M |
| 48 | 6 | 0.5 | 0.306 | 22.93 | 0.29 M |
| 48 | 9 | 0.5 | 0.303 | 22.86 | 0.44 M |
| 96 | 3 | 0.5 | 0.307 | 22.94 | 0.59 M |
| 96 | 6 | 0.5 | 0.304 | 22.86 | 1.17 M |
| 96 | 9 | 0.5 | 0.304 | 22.80 | 1.75 M |
| Method | 1 dB | 4 dB | 7 dB | 10 dB | 13 dB |
|---|---|---|---|---|---|
| DeepJSCC | 0.509 | 0.462 | 0.432 | 0.415 | 0.405 |
| DiffJSCC † | 0.638 | 0.600 | 0.578 | 0.556 | 0.552 |
| DnCNN * | 0.368 | 0.326 | 0.300 | 0.286 | 0.279 |
| Ours * | 0.384 | 0.342 | 0.312 | 0.294 | 0.282 |
| Method | 1 dB | 7 dB | 13 dB |
|---|---|---|---|
| LPIPS | |||
| DeepJSCC (C8) | 0.655 | 0.595 | 0.573 |
| +Ours | 0.410 | 0.373 | 0.364 |
| Method | 1 dB | 7 dB | 13 dB |
|---|---|---|---|
| PSNR (dB) | |||
| DeepJSCC (C8) | 21.12 | 22.06 | 21.99 |
| +Ours | 20.48 | 21.38 | 21.27 |
| CBR | Method | 1 dB | 7 dB | 13 dB |
|---|---|---|---|---|
| 1/96 (C2) | DeepJSCC | 0.735 | 0.650 | 0.611 |
| + Ours | 0.415 | 0.381 | 0.370 | |
| 1/48 (C4) | DeepJSCC | 0.566 | 0.497 | 0.471 |
| +Ours | 0.342 | 0.292 | 0.271 | |
| 1/24 (C8) | DeepJSCC | 0.654 | 0.595 | 0.573 |
| +Ours | 0.358 | 0.327 | 0.318 | |
| Relative LPIPS Improvement | ||||
| 1/96 (C2) | Ours | 43.5% | 41.4% | 39.4% |
| 1/48 (C4) | Ours | 39.6% | 41.2% | 42.5% |
| 1/24 (C8) | Ours | 45.3% | 45.0% | 44.5% |
| Method | 1 dB | 4 dB | 7 dB | 10 dB | 13 dB |
|---|---|---|---|---|---|
| LPIPS | |||||
| DeepJSCC | 0.649 | 0.606 | 0.550 | 0.535 | 0.514 |
| DnCNN * | 0.654 | 0.612 | 0.556 | 0.543 | 0.521 |
| Ours * | 0.451 | 0.419 | 0.372 | 0.363 | 0.343 |
| Relative LPIPS Improvement | |||||
| Ours vs. DeepJSCC | 30.5% | 30.9% | 32.4% | 32.1% | 33.3% |
| K | Method | 1 dB | 4 dB | 7 dB | 10 dB | 13 dB |
|---|---|---|---|---|---|---|
| LPIPS (mean ± std) | ||||||
| 3 | DeepJSCC | 0.620 | 0.578 | 0.530 | 0.503 | 0.487 |
| Ours * | 0.404 | 0.376 | 0.336 | 0.312 | 0.293 | |
| 10 | DeepJSCC | 0.588 | 0.545 | 0.505 | 0.489 | 0.475 |
| Ours * | 0.370 | 0.342 | 0.308 | 0.293 | 0.275 | |
| Relative LPIPS Improvement (Ours vs. DeepJSCC) | ||||||
| 3 | Ours | 34.8% | 34.9% | 36.6% | 38.0% | 39.8% |
| 10 | Ours | 37.1% | 37.2% | 39.0% | 40.1% | 42.1% |
| Component | GPU (ms) | NPU (ms) | CPU (ms) | Memory (MB) | Params |
|---|---|---|---|---|---|
| DeepJSCC decoder | 9.4 ± 0.5 | — | 504.1 ± 35.9 | 429.6 | 31.2 M |
| Enhancer only | 28.0 ± 0.1 | 21.7 ± 0.05 | 1005.5 ± 74.1 | 293.7 | 0.29 M |
| Full pipeline | 38.3 ± 1.5 | — | 1512.0 ± 130.7 | 431.5 | 31.5 M |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Hou, S.; Zhao, P.; Chen, N. Lightweight SNR-Adaptive Receiver-Side Enhancement for DeepJSCC-Based Wireless Image Transmission. Sensors 2026, 26, 5134. https://doi.org/10.3390/s26165134
Hou S, Zhao P, Chen N. Lightweight SNR-Adaptive Receiver-Side Enhancement for DeepJSCC-Based Wireless Image Transmission. Sensors. 2026; 26(16):5134. https://doi.org/10.3390/s26165134
Chicago/Turabian StyleHou, Shouquan, Peng Zhao, and Nuo Chen. 2026. "Lightweight SNR-Adaptive Receiver-Side Enhancement for DeepJSCC-Based Wireless Image Transmission" Sensors 26, no. 16: 5134. https://doi.org/10.3390/s26165134
APA StyleHou, S., Zhao, P., & Chen, N. (2026). Lightweight SNR-Adaptive Receiver-Side Enhancement for DeepJSCC-Based Wireless Image Transmission. Sensors, 26(16), 5134. https://doi.org/10.3390/s26165134

