Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (22)

Search Parameters:
Keywords = decoding-complexity–rate–distortion

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
25 pages, 3230 KB  
Article
Lightweight State-Space Model-Based Video Quality Enhancement for Quadruped Robot Dog Decoded Streams
by Wentao Feng, Yuanchun Huang and Zhenglong Yang
Electronics 2026, 15(6), 1151; https://doi.org/10.3390/electronics15061151 - 10 Mar 2026
Viewed by 620
Abstract
In the field of intelligent inspection, high-definition video data collected by quadruped robot dogs face severe transmission and storage constraints. Although existing advanced lossy video coding standards can significantly improve compression efficiency, they inevitably introduce severe compression artifacts in low-bit-rate scenarios. To address [...] Read more.
In the field of intelligent inspection, high-definition video data collected by quadruped robot dogs face severe transmission and storage constraints. Although existing advanced lossy video coding standards can significantly improve compression efficiency, they inevitably introduce severe compression artifacts in low-bit-rate scenarios. To address this issue, this paper proposes a video decoding quality enhancement network named Video Quality Restoration Network (VQRNet), based on a dual-stream architecture. Specifically, the Local Feature Extraction component incorporates a Progressive Feature Fusion Module (PFFM) with a four-stage progressive structure. By integrating reparameterized convolution and attention mechanisms, PFFM focuses on capturing high-frequency texture details to repair small-scale distortions. Simultaneously, the Multi-Scale Lightweight Spatial Attention Module (MLSA) performs spatial feature recalibration, leveraging multi-scale convolution to adaptively identify and enhance key spatial regions, specifically addressing multi-scale distortion. In the Global Feature Extraction component, the State-Space Attention Module (SSAM) combines State-Space Models (SSMs) with attention mechanisms to capture long-range dependencies and contextual information, for large-scale distortions caused by high-intensity compression. To verify the performance of the proposed algorithm, a dedicated dataset comprising 20 real-world video sequences captured by quadruped robot dogs (partitioned into 15 training and 5 testing sequences) was constructed, and the VTM 23.4 reference software was employed to simulate compression degradation using four quantization parameters (QP 30, 35, 40, and 45). Experimental results demonstrate that VQRNet outperforms state-of-the-art quality enhancement methods in terms of core metrics, including PSNR and SSIM, specifically including MIRNet, NAFNet, TRRHA, and CTNet. In the QP = 30 scenario, VQRNet achieves an average PSNR of 40.33 dB, a significant improvement of 3.32 dB over the VTM 23.4 baseline (37.01 dB), while demonstrating significant advantages in computational complexity and parameter efficiency—requiring only 5.27 G FLOPs and 1.40 M parameters, with an average inference latency of only 11.82 ms per 128 × 128 patch. This work provides robust technical support for the efficient video perception of quadruped robot dogs. Full article
Show Figures

Figure 1

27 pages, 827 KB  
Article
Deep Learning-Enabled LoRa-JSCC for Efficient and Reliable Multivariate Sensor Data Transmission in IoT Environments
by Fatimah Alghamdi and Fuad Bajaber
Electronics 2026, 15(5), 1040; https://doi.org/10.3390/electronics15051040 - 2 Mar 2026
Viewed by 738
Abstract
Integrating Joint Source–Channel Coding (JSCC) with the LoRa Chirp Spread Spectrum (CSS) physical layer (PHY) presents a significant challenge due to the complexity of joint optimization, which remains underexplored despite the known advantages of JSCC. Traditional LoRa systems rely on decoupled source and [...] Read more.
Integrating Joint Source–Channel Coding (JSCC) with the LoRa Chirp Spread Spectrum (CSS) physical layer (PHY) presents a significant challenge due to the complexity of joint optimization, which remains underexplored despite the known advantages of JSCC. Traditional LoRa systems rely on decoupled source and channel coding, resulting in redundant overhead and limited adaptability under dynamic Wireless Body Area Network (WBAN) conditions. To address these limitations, we propose a novel LoRa–JSCC framework: a fully learned, end-to-end differentiable architecture that jointly optimizes source compression and channel redundancy. The proposed system integrates a Denoising Autoencoder (DAE) for non-linear source compression with learned neural channel encoder and decoder modules, trained via backpropagation to minimize reconstruction distortion under noisy channel conditions. Rigorous Monte Carlo simulations conducted under unified and reproducible channel conditions demonstrate consistent performance improvements across LoRa configurations. The proposed approach achieves an average 25–30% improvement in goodput across moderate-to-high SNR regimes, with gains exceeding 100% under noise-limited conditions. It further reduces Time on Air (ToA) by approximately 30–35%, enhancing spectral efficiency and lowering effective energy cost per delivered bit. In the transitional Bit Error Rate (BER) region, the proposed LoRa–JSCC framework exhibits an effective SNR gain of approximately 18–20 dB relative to conventional LoRa, corresponding to multiple orders-of-magnitude reduction in BER. These results indicate substantial improvements in reliability, coverage robustness, and energy efficiency for WBAN and IoT deployments. Full article
Show Figures

Figure 1

5 pages, 398 KB  
Proceeding Paper
A Lightweight Deep Learning Framework for Robust Video Watermarking in Adversarial Environments
by Antonio Cedillo-Hernandez, Lydia Velazquez-Garcia and Manuel Cedillo-Hernandez
Eng. Proc. 2026, 123(1), 25; https://doi.org/10.3390/engproc2026123025 - 5 Feb 2026
Cited by 1 | Viewed by 714
Abstract
The widespread distribution of digital videos in social networks, streaming services, and surveillance systems has increased the risk of manipulation, unauthorized redistribution, and adversarial tampering. This paper presents a lightweight deep learning framework for robust and imperceptible video watermarking designed specifically for cybersecurity [...] Read more.
The widespread distribution of digital videos in social networks, streaming services, and surveillance systems has increased the risk of manipulation, unauthorized redistribution, and adversarial tampering. This paper presents a lightweight deep learning framework for robust and imperceptible video watermarking designed specifically for cybersecurity environments. Unlike heavy architectures that rely on multi-scale feature extractors or complex adversarial networks, our model introduces a compact encoder–decoder pipeline optimized for real-time watermark embedding and recovery under adversarial attacks. The proposed system leverages spatial attention and temporal redundancy to ensure robustness against distortions such as compression, additive noise, and adversarial perturbations generated via Fast Gradient Sign Method (FGSM) or recompression attacks from generative models. Experimental simulations using a reduced Kinetics-600 subset demonstrate promising results, achieving an average PSNR of 38.9 dB, SSIM of 0.967, and Bit Error Rate (BER) below 3% even under FGSM attacks. These results suggest that the proposed lightweight framework achieves a favorable trade-off between resilience, imperceptibility, and computational efficiency, making it suitable for deployment in video forensics, authentication, and secure content distribution systems. Full article
(This article belongs to the Proceedings of First Summer School on Artificial Intelligence in Cybersecurity)
Show Figures

Figure 1

26 pages, 10192 KB  
Article
Multi-Robot Task Allocation with Spatiotemporal Constraints via Edge-Enhanced Attention Networks
by Yixiang Hu, Daxue Liu, Jinhong Li, Junxiang Li and Tao Wu
Appl. Sci. 2026, 16(2), 904; https://doi.org/10.3390/app16020904 - 15 Jan 2026
Cited by 2 | Viewed by 1359
Abstract
Multi-Robot Task Allocation (MRTA) with spatiotemporal constraints presents significant challenges in environmental adaptability. Existing learning-based methods often overlook environmental spatial constraints, leading to spatial information distortion. To address this, we formulate the problem as an asynchronous Markov Decision Process over a directed heterogeneous [...] Read more.
Multi-Robot Task Allocation (MRTA) with spatiotemporal constraints presents significant challenges in environmental adaptability. Existing learning-based methods often overlook environmental spatial constraints, leading to spatial information distortion. To address this, we formulate the problem as an asynchronous Markov Decision Process over a directed heterogeneous graph and propose a novel heterogeneous graph neural network named the Edge-Enhanced Attention Network (E2AN). This network integrates a specialized encoder, the Edge-Enhanced Heterogeneous Graph Attention Network (E2HGAT), with an attention-based decoder. By incorporating edge attributes to effectively characterize path costs under spatial constraints, E2HGAT corrects spatial distortion. Furthermore, our approach supports flexible extension to diverse payload scenarios via node attribute adaptation. Extensive experiments conducted in simulated environments with obstructed maps demonstrate that the proposed method outperforms baseline algorithms in task success rate. Remarkably, the model maintains its advantages in generalization tests on unseen maps as well as in scalability tests across varying problem sizes. Ablation studies further validate the critical role of the proposed encoder in capturing spatiotemporal dependencies. Additionally, real-time performance analysis confirms the method’s feasibility for online deployment. Overall, this study offers an effective solution for MRTA problems with complex constraints. Full article
(This article belongs to the Special Issue Motion Control for Robots and Automation)
Show Figures

Figure 1

16 pages, 10696 KB  
Article
A Framework for Symmetric-Quality S3D Video Streaming Services
by Juhyeon Lee, Seungjun Lee, Sunghoon Kim and Dongwook Kang
Appl. Sci. 2024, 14(23), 11011; https://doi.org/10.3390/app142311011 - 27 Nov 2024
Cited by 1 | Viewed by 1380
Abstract
This paper proposes an efficient encoding framework based on Scalable High Efficiency Video Coding (SHVC) technology, which supports both low- and high-resolution 2D videos as well as stereo 3D (S3D) video simultaneously. Previous studies have introduced Cross-View SHVC, which encodes two videos with [...] Read more.
This paper proposes an efficient encoding framework based on Scalable High Efficiency Video Coding (SHVC) technology, which supports both low- and high-resolution 2D videos as well as stereo 3D (S3D) video simultaneously. Previous studies have introduced Cross-View SHVC, which encodes two videos with different viewpoints and resolutions using a Cross-View SHVC encoder, where the low-resolution video is encoded as the base layer and the other video as the enhancement layer. This encoder provides resolution diversity and allows the decoder to combine the two videos, enabling 3D video services. Even when 3D videos are composed of left and right videos with different resolutions, the viewer tends to perceive the quality based on the higher-resolution video due to the binocular suppression effect, where the brain prioritizes the high-quality image and suppresses the lower-quality one. However, recent experiments have shown that when the disparity between resolutions exceeds a certain threshold, it can lead to a subjective degradation of the perceived 3D video quality. To address this issue, a conditional replenishment algorithm has been studied, which replaces some blocks of the video using a disparity-compensated left-view image based on rate–distortion cost. This conditional replenishment algorithm (also known as VEI technology) effectively reduces the quality difference between the base layer and enhancement layer videos. However, the algorithm alone cannot fully compensate for the quality difference between the left and right videos. In this paper, we propose a novel encoding framework to solve the asymmetry issue between the left and right videos in 3D video services and achieve symmetrical video quality. The proposed framework focuses on improving the quality of the right-view video by combining the conditional replenishment algorithm with Cross-View SHVC. Specifically, the framework leverages the non-HEVC option of the SHVC encoder, using a VEI (Video Enhancement Information) restored image as the base layer to provide higher-quality prediction signals and reduce encoding complexity. Experimental results using animation and live-action UHD sequences show that the proposed method achieves BD-RATE reductions of 57.78% and 45.10% compared with HEVC and SHVC codecs, respectively. Full article
Show Figures

Figure 1

21 pages, 6596 KB  
Article
MRACNN: Multi-Path Residual Asymmetric Convolution and Enhanced Local Attention Mechanism for Industrial Image Compression
by Zikang Yan, Peishun Liu, Xuefang Wang, Haojie Gao, Xiaolong Ma and Xintong Hu
Symmetry 2024, 16(10), 1342; https://doi.org/10.3390/sym16101342 - 10 Oct 2024
Cited by 1 | Viewed by 2249
Abstract
The rich information and complex background of industrial images make it a challenging task to improve the high compression rate of images. Current learning-based image compression methods mostly use customized convolutional neural networks (CNNs), which find it difficult to cope with the complex [...] Read more.
The rich information and complex background of industrial images make it a challenging task to improve the high compression rate of images. Current learning-based image compression methods mostly use customized convolutional neural networks (CNNs), which find it difficult to cope with the complex production background of industrial images. This causes useful information to be lost in the abundance of irrelevant data, making it difficult to accurately extract important features during the feature extraction stage. To address this, a Multi-path Residual Asymmetric Convolutional Compression Network (MRACNN) is proposed. Firstly, a Multi-path Residual Asymmetric Convolution Block (MRACB) is introduced, which includes the Multi-path Residual Asymmetric Convolution Down-sampling Module for down-sampling in the encoder to extract key features, and the Mult-path Residual Asymmetric Convolution Up-sampling Module for up-sampling in the decoder to recover details and reconstruct the image. This feature transfer and information flow enables the better capture of image details and important information, thereby improving the quality and efficiency of image compression and decompression. Furthermore, a two-branch enhanced local attention mechanisms, and a channel-squeezing entropy model based on the compression-based enhanced local attention module is proposed to enhance the performance of the modeled compression. Extensive experimental evaluations demonstrate that the proposed method outperforms state-of-the-art techniques, achieves superior Rate–Distortion Performance, and excels in preserving local details. Full article
(This article belongs to the Special Issue Symmetry/Asymmetry in Neural Networks and Applications)
Show Figures

Figure 1

22 pages, 13134 KB  
Article
Syntax-Guided Content-Adaptive Transform for Image Compression
by Yunhui Shi, Liping Ye, Jin Wang, Lilong Wang, Hui Hu, Baocai Yin and Nam Ling
Sensors 2024, 24(16), 5439; https://doi.org/10.3390/s24165439 - 22 Aug 2024
Cited by 4 | Viewed by 2379
Abstract
The surge in image data has significantly increased the pressure on storage and transmission, posing new challenges for image compression technology. The structural texture of an image implies its statistical characteristics, which is effective for image encoding and decoding. Consequently, content-adaptive compression methods [...] Read more.
The surge in image data has significantly increased the pressure on storage and transmission, posing new challenges for image compression technology. The structural texture of an image implies its statistical characteristics, which is effective for image encoding and decoding. Consequently, content-adaptive compression methods based on learning can better capture the content attributes of images, thereby enhancing encoding performance. However, learned image compression methods do not comprehensively account for both the global and local correlations among the pixels within an image. Moreover, they are constrained by rate-distortion optimization, which prevents the attainment of a compact representation of image attributes. To address these issues, we propose a syntax-guided content-adaptive transform framework that efficiently captures image attributes and enhances encoding efficiency. Firstly, we propose a syntax-refined side information module that fully leverages syntax and side information to guide the adaptive transformation of image attributes. Moreover, to more thoroughly exploit the global and local correlations in image space, we designed global–local modules, local–global modules, and upsampling/downsampling modules in codecs, further eliminating local and global redundancies. The experimental findings indicate that our proposed syntax-guided content-adaptive image compression model successfully adapts to the diverse complexities of different images, which enhances the efficiency of image compression. Concurrently, the method proposed has demonstrated outstanding performance across three benchmark datasets. Full article
(This article belongs to the Section Sensing and Imaging)
Show Figures

Figure 1

17 pages, 5104 KB  
Article
Hierarchical Vector-Quantized Variational Autoencoder and Vector Credibility Mechanism for High-Quality Image Inpainting
by Cheng Li, Dan Xu and Kuai Chen
Electronics 2024, 13(10), 1852; https://doi.org/10.3390/electronics13101852 - 9 May 2024
Cited by 7 | Viewed by 4719
Abstract
Image inpainting infers the missing areas of a corrupted image according to the information of the undamaged part. Many existing image inpainting methods can generate plausible inpainted results from damaged images with the fast-developed deep-learning technology. However, they still suffer from over-smoothed textures [...] Read more.
Image inpainting infers the missing areas of a corrupted image according to the information of the undamaged part. Many existing image inpainting methods can generate plausible inpainted results from damaged images with the fast-developed deep-learning technology. However, they still suffer from over-smoothed textures or textural distortion in the cases of complex textural details or large damaged areas. To restore textures at a fine-grained level, we propose an image inpainting method based on a hierarchical VQ-VAE with a vector credibility mechanism. It first trains the hierarchical VQ-VAE with ground truth images to update two codebooks and to obtain two corresponding vector collections containing information on ground truth images. The two vector collections are fed to a decoder to generate the corresponding high-fidelity outputs. An encoder then is trained with the corresponding damaged image. It generates vector collections approximating the ground truth by the help of the prior knowledge provided by the codebooks. After that, the two vector collections pass through the decoder from the hierarchical VQ-VAE to produce the inpainted results. In addition, we apply a vector credibility mechanism to promote vector collections from damaged images and approximate vector collections from ground truth images. To further improve the inpainting result, we apply a refinement network, which uses residual blocks with different dilation rates to acquire both global information and local textural details. Extensive experiments conducted on several datasets demonstrate that our method outperforms the state-of-the-art ones. Full article
(This article belongs to the Special Issue Applications of Artificial Intelligence in Image and Video Processing)
Show Figures

Figure 1

20 pages, 4604 KB  
Article
Full-Process Adaptive Encoding and Decoding Framework for Remote Sensing Images Based on Compression Sensing
by Huiling Hu, Chunyu Liu, Shuai Liu, Shipeng Ying, Chen Wang and Yi Ding
Remote Sens. 2024, 16(9), 1529; https://doi.org/10.3390/rs16091529 - 26 Apr 2024
Cited by 2 | Viewed by 2394
Abstract
Faced with the problem of incompatibility between traditional information acquisition mode and spaceborne earth observation tasks, starting from the general mathematical model of compressed sensing, a theoretical model of block compressed sensing was established, and a full-process adaptive coding and decoding compressed sensing [...] Read more.
Faced with the problem of incompatibility between traditional information acquisition mode and spaceborne earth observation tasks, starting from the general mathematical model of compressed sensing, a theoretical model of block compressed sensing was established, and a full-process adaptive coding and decoding compressed sensing framework for remote sensing images was proposed, which includes five parts: mode selection, feature factor extraction, adaptive shape segmentation, adaptive sampling rate allocation and image reconstruction. Unlike previous semi-adaptive or local adaptive methods, the advantages of the adaptive encoding and decoding method proposed in this paper are mainly reflected in four aspects: (1) Ability to select encoding modes based on image content, and maximizing the use of the richness of the image to select appropriate sampling methods; (2) Capable of utilizing image texture details for adaptive segmentation, effectively separating complex and smooth regions; (3) Being able to detect the sparsity of encoding blocks and adaptively allocate sampling rates to fully explore the compressibility of images; (4) The reconstruction matrix can be adaptively selected based on the size of the encoding block to alleviate block artifacts caused by non-stationary characteristics of the image. Experimental results show that the method proposed in this article has good stability for remote sensing images with complex edge textures, with the peak signal-to-noise ratio and structural similarity remaining above 35 dB and 0.8. Moreover, especially for ocean images with relatively simple image content, when the sampling rate is 0.26, the peak signal-to-noise ratio reaches 50.8 dB, and the structural similarity is 0.99. In addition, the recovered images have the smallest BRISQUE value, with better clarity and less distortion. In the subjective aspect, the reconstructed image has clear edge details and good reconstruction effect, while the block effect is effectively suppressed. The framework designed in this paper is superior to similar algorithms in both subjective visual and objective evaluation indexes, which is of great significance for alleviating the incompatibility between traditional information acquisition methods and satellite-borne earth observation missions. Full article
Show Figures

Figure 1

22 pages, 2000 KB  
Article
Low Computational Coding-Efficient Distributed Video Coding: Adding a Decision Mode to Limit Channel Coding Load
by Shahzad Khursheed, Nasreen Badruddin, Varun Jeoti, Dejan Vukobratovic and Manzoor Ahmed Hashmani
Entropy 2023, 25(2), 241; https://doi.org/10.3390/e25020241 - 28 Jan 2023
Cited by 2 | Viewed by 2349
Abstract
Distributed video coding (DVC) is based on distributed source coding (DSC) concepts in which video statistics are used partially or completely at the decoder rather than the encoder. The rate-distortion (RD) performance of distributed video codecs substantially lags the conventional predictive video coding. [...] Read more.
Distributed video coding (DVC) is based on distributed source coding (DSC) concepts in which video statistics are used partially or completely at the decoder rather than the encoder. The rate-distortion (RD) performance of distributed video codecs substantially lags the conventional predictive video coding. Several techniques and methods are employed in DVC to overcome this performance gap and achieve high coding efficiency while maintaining low encoder computational complexity. However, it is still challenging to achieve coding efficiency and limit the computational complexity of the encoding and decoding process. The deployment of distributed residual video coding (DRVC) improves coding efficiency, but significant enhancements are still required to reduce these gaps. This paper proposes the QUAntized Transform ResIdual Decision (QUATRID) scheme that improves the coding efficiency by deploying the Quantized Transform Decision Mode (QUAM) at the encoder. The proposed QUATRID scheme’s main contribution is a design and integration of a novel QUAM method into DRVC that effectively skips the zero quantized transform (QT) blocks, thus limiting the number of input bit planes to be channel encoded and consequently reducing both the channel encoding and decoding computational complexity. Moreover, an online correlation noise model (CNM) is specifically designed for the QUATRID scheme and implemented at its decoder. This online CNM improves the channel decoding process and contributes to the bit rate reduction. Finally, a methodology for the reconstruction of the residual frame (R^) is developed that utilizes the decision mode information passed by the encoder, decoded quantized bin, and transformed estimated residual frame. The Bjøntegaard delta analysis of experimental results shows that the QUATRID achieves better performance over the DISCOVER by attaining the PSNR between 0.06 dB and 0.32 dB and coding efficiency, which varies from 5.4 to 10.48 percent. In addition to this, results determine that for all types of motion videos, the proposed QUATRID scheme outperforms the DISCOVER in terms of reducing the number of input bit-planes to be channel encoded and the entire encoder’s computational complexity. The number of bit plane reduction exceeds 97%, while the entire Wyner-Ziv encoder and channel coding computational complexity reduce more than nine-fold and 34-fold, respectively. Full article
Show Figures

Figure 1

14 pages, 551 KB  
Communication
A Novel Expectation-Maximization-Based Blind Receiver for Low-Complexity Uplink STLC-NOMA Systems
by Ki-Hun Lee and Bang Chul Jung
Sensors 2022, 22(20), 8054; https://doi.org/10.3390/s22208054 - 21 Oct 2022
Cited by 1 | Viewed by 2399
Abstract
In this paper, we revisit a two-user space-time line coded uplink non-orthogonal multiple access (STLC-NOMA) system for Internet-of-things (IoT) networks and propose a novel low-complexity STLC-NOMA system. The basic idea is that both IoT devices (stations: STAs) employ amplitude-shift keying (ASK) modulators and [...] Read more.
In this paper, we revisit a two-user space-time line coded uplink non-orthogonal multiple access (STLC-NOMA) system for Internet-of-things (IoT) networks and propose a novel low-complexity STLC-NOMA system. The basic idea is that both IoT devices (stations: STAs) employ amplitude-shift keying (ASK) modulators and align their modulated symbols to in-phase and quadrature axes, respectively, before the STLC encoding. The phase distortion caused by wireless channels becomes compensated at the receiver side with the STLC, and thus each STA’s signals are still aligned on their axes at the access point (AP) in the proposed uplink STLC-NOMA system. Then, the AP can decode the signals transmitted from STAs via a single-user maximum-likelihood (ML) detector with low-complexity, while the conventional uplink STLC-NOMA system exploits a multi-user joint ML detector with relatively high-complexity. We mathematically analyze the exact BER performance of the proposed uplink STLC-NOMA system. Furthermore, we propose a novel expectation-maximization (EM)-based blind energy estimation (BEE) algorithm to jointly estimate both transmit power and effective channel gain of each STA without the help of pilot signals at the AP. Somewhat interestingly, the proposed BEE algorithm works well even in short-packet transmission scenarios. It is worth noting that the proposed uplink STLC-NOMA architecture outperforms the conventional STLC-NOMA technique in terms of bit-error-rate (BER), especially with high-order modulation schemes, even though it requires lower computation complexity than the conventional technique at the receiver. Full article
(This article belongs to the Special Issue Advanced Antenna Techniques for IoT and 5G Applications)
Show Figures

Figure 1

18 pages, 1541 KB  
Article
No-Reference Quality Assessment of Transmitted Stereoscopic Videos Based on Human Visual System
by Md Mehedi Hasan, Md. Ariful Islam, Sejuti Rahman, Michael R. Frater and John F. Arnold
Appl. Sci. 2022, 12(19), 10090; https://doi.org/10.3390/app121910090 - 7 Oct 2022
Cited by 7 | Viewed by 3090
Abstract
Provisioning the stereoscopic 3D (S3D) video transmission services of admissible quality in a wireless environment is an immense challenge for video service providers. Unlike for 2D videos, a widely accepted No-reference objective model for assessing transmitted 3D videos that explores the Human Visual [...] Read more.
Provisioning the stereoscopic 3D (S3D) video transmission services of admissible quality in a wireless environment is an immense challenge for video service providers. Unlike for 2D videos, a widely accepted No-reference objective model for assessing transmitted 3D videos that explores the Human Visual System (HVS) appropriately has not been developed yet. Distortions perceived in 2D and 3D videos are significantly different due to the sophisticated manner in which the HVS handles the dissimilarities between the two different views. In real-time video transmission, viewers only have the distorted or receiver end content of the original video acquired through the communication medium. In this paper, we propose a No-reference quality assessment method that can estimate the quality of a stereoscopic 3D video based on HVS. By evaluating perceptual aspects and correlations of visual binocular impacts in a stereoscopic movie, the approach creates a way for the objective quality measure to assess impairments similarly to a human observer who would experience the similar material. Firstly, the disparity is measured and quantified by the region-based similarity matching algorithm, and then, the magnitude of the edge difference is calculated to delimit the visually perceptible areas of an image. Finally, an objective metric is approximated by extracting these significant perceptual image features. Experimental analysis with standard S3D video datasets demonstrates the lower computational complexity for the video decoder and comparison with the state-of-the-art algorithms shows the efficiency of the proposed approach for 3D video transmission at different quantization (QP 26 and QP 32) and loss rate (1% and 3% packet loss) parameters along with the perceptual distortion features. Full article
(This article belongs to the Special Issue Computational Intelligence in Image and Video Analysis)
Show Figures

Figure 1

28 pages, 10117 KB  
Article
Novel Projection Schemes for Graph-Based Light Field Coding
by Nguyen Gia Bach, Chanh Minh Tran, Tho Nguyen Duc, Phan Xuan Tan and Eiji Kamioka
Sensors 2022, 22(13), 4948; https://doi.org/10.3390/s22134948 - 30 Jun 2022
Cited by 2 | Viewed by 2648
Abstract
In light field compression, graph-based coding is powerful to exploit signal redundancy along irregular shapes and obtains good energy compaction. However, apart from high time complexity to process high dimensional graphs, their graph construction method is highly sensitive to the accuracy of disparity [...] Read more.
In light field compression, graph-based coding is powerful to exploit signal redundancy along irregular shapes and obtains good energy compaction. However, apart from high time complexity to process high dimensional graphs, their graph construction method is highly sensitive to the accuracy of disparity information between viewpoints. In real-world light field or synthetic light field generated by computer software, the use of disparity information for super-rays projection might suffer from inaccuracy due to vignetting effect and large disparity between views in the two types of light fields, respectively. This paper introduces two novel projection schemes resulting in less error in disparity information, in which one projection scheme can also significantly reduce computation time for both encoder and decoder. Experimental results show projection quality of super-pixels across views can be considerably enhanced using the proposals, along with rate-distortion performance when compared against original projection scheme and HEVC-based or JPEG Pleno-based coding approaches. Full article
(This article belongs to the Section Sensing and Imaging)
Show Figures

Figure 1

21 pages, 1406 KB  
Article
A Novel Online Correlation Noise Model Based on Band Coefficients Mean to Achieve Low Computational and Coding-Efficient Distributed Video Codec
by Shahzad Khursheed, Nasreen Badruddin, Varun Jeoti and Manzoor Ahmed Hashmani
Appl. Sci. 2022, 12(13), 6505; https://doi.org/10.3390/app12136505 - 27 Jun 2022
Cited by 1 | Viewed by 2080
Abstract
Distributed video coding (DVC) is a novel coding paradigm that offers low computational encoding relative to conventional video-coding framework at the expense of high-decoding computational complexity. The challenging part of this video-coding framework is achieving better rate-distortion (RD) compared with conventional codec performance. [...] Read more.
Distributed video coding (DVC) is a novel coding paradigm that offers low computational encoding relative to conventional video-coding framework at the expense of high-decoding computational complexity. The challenging part of this video-coding framework is achieving better rate-distortion (RD) compared with conventional codec performance. A suitable and accurate correlation noise model (CNM) is crucial in improving the RD performance by achieving high coding efficiency and making decoding less computationally demanding. Since the correlation is nonstationary and time-variant and can vary from frame to frame, offline CNM estimation is not feasible for practical applications and real-time decoding. An online CNM may be the solution to this problem. In DVC, neither Wyner–Ziv frame (WZF) nor estimated side information (SI) of the corresponding WZF is available at the encoder. Therefore, online estimation of the CNM and its parameters can be quite challenging. The contribution of this research work is a novel online CNM which is computed by taking the mean of each transformed coefficient band and deployed for two different codecs. Our proposed codec, DIVCOM, which stands for “Distributed Video Coding with Online Band Mean Correlation Noise Model”, outperforms the existing baseline codec, DISCOVER (DIS), in both coding efficiency and peak signal-to-noise ratio (PSNR). The DIVCOM codec achieves coding efficiency of up to 8.05 kbps, and PSNR ranges from 0.0245 dB to 0.18 dB. An extended version of DIVCOM incorporating phase-based side information called PDIVCOM achieves coding efficiency up to 10.9 kbps, and PSNR ranges from 0.019 to 0.17 dB compared to DIS. Full article
(This article belongs to the Special Issue Advances on Image, Video and Signal Processing)
Show Figures

Figure 1

24 pages, 651 KB  
Article
Multimode Tree-Coding of Speech with Pre-/Post-Weighting
by Ying-Yi Li, Pravin Ramadas and Jerry Gibson
Appl. Sci. 2022, 12(4), 2026; https://doi.org/10.3390/app12042026 - 15 Feb 2022
Cited by 1 | Viewed by 2538
Abstract
As speech-coding standards have improved over the years, so complexity has increased, and less emphasis been placed on low encoding/decoding delay. We present a low-complexity, low-delay speech codec based on tree-coding with sample-by-sample adaptive long- and short-code generators that incorporates pre- and post-filtering [...] Read more.
As speech-coding standards have improved over the years, so complexity has increased, and less emphasis been placed on low encoding/decoding delay. We present a low-complexity, low-delay speech codec based on tree-coding with sample-by-sample adaptive long- and short-code generators that incorporates pre- and post-filtering for perceptual weighting and multimode speech classification with comfort noise generation (CNG). The pre-/post-weighting filters adapt based on the code generator parameters available at both the encoder and decoder rather than the usual method that uses the input speech. The coding of the multiple speech modes and comfort noise generation is accomplished using the code generator adaptation algorithms, again, rather than using the input speech. Codec complexity comparisons are presented and operational rate distortion curves for several standardized speech codecs and the new codec are given. Finally, codec performance is shown in relation to theoretical rate distortion bounds. Full article
(This article belongs to the Section Electrical, Electronics and Communications Engineering)
Show Figures

Figure 1

Back to TopTop