Abstract
Matrix embedding (ME) code is a commonly used steganography technique, which uses linear block codes to improve embedding efficiency. However, its main disadvantage is the inability to perform maximum likelihood decoding due to the high complexity of decoding large ME codes. As such, it is difficult to improve the embedding efficiency. The proposed q-ary embedding code can provide excellent embedding efficiency and is suitable for various embedding rates (large and small payloads). This article discusses that by using perforation technology, a convolutional code with a high embedding rate can be easily converted into a convolutional code with a low embedding rate. By keeping the embedding rate of the (2, 1) convolutional code unchanged, convolutional codes with different embedding rates can be designed through puncturing.
1. Introduction
Among the numerous steganography techniques that have been developed, matrix embedding (ME) [1,2] provides high undetectability and embedding efficiency, which result in efficient steganographic security. Steganography refers to embedding data to conceal objects such as images, videos, or audio. In steganography, the covered object is modified to obtain a stego.
Numerous ME codes based on covering codes [3,4,5,6] have been developed because they exhibit high embedding efficiency due to their favorable structural characteristics, such as excellent weight distribution of the coset leaders of linear codes. In [5], several coverage code series are constructed using factorized block-by-block direct sum (BDS). BDS(6) and BDS(8) provided the highest embedding efficiency. The use of nonlinear covering codes considerably improved efficiency. Fridrich et al. [7] proposed an ME-based embedding technique that comprises two types of linear block codes, namely simplex codes and a random code. The technique exhibited high efficiency for large payloads [7], which resulted in superior steganographic security. Furthermore, they use structured simple codes (including decoding by using fast Hadamard decoding) to obtain effective ME codes and approach the efficiency limit of large payloads. Generally, good ME codes are based on suitable linear block codes that are long enough. Due to the complexity of maximum likelihood (ML) decoding, it is difficult to determine the coset preamble of a large linear block code. Numerous approaches using structured codes [8,9,10,11,12,13] have been developed.
In coding theory, the best ME code (ME code that can approach the upper limit of the embedding efficiency of the rate-distortion function) requires a well-structured code and a sufficiently long effective decoding algorithm, such as a low-density generator matrix code [14]. Researchers have developed a great number of embedding techniques in adaptive steganography. The study in [15] proposes an adaptive algorithm called the linear independent approximation embedding (LIAE) algorithm. The LIAE algorithm has the ability to perform data embedding at an arbitrarily specified cover location. The method presented in this study used a family of convolutional codes known as convolutional embedding (CE) codes for q-ary payloads. The CE code can be used as an alternative approach to the theoretical upper limit of embedding efficiency. The CE code is based on a grid structure and Viterbi decoding (this is an ML algorithm). The CE code is suitable for encoding a payload with a sufficiently large block length to increase the embedding efficiency and change the embedding rate. Additionally, the optimal design of current CEs can be used to obtain the embedding scheme. Moreover, a puncturing technique is suitable for altering the embedding rate of CE codes. For the q-element payload, the CE code can be easily obtained at an embedding rate of 1/2 CE by using a piercing strategy. Experimental results show that the embedding efficiency of CE code is better than that of ME code.
The rest of this article is organized as follows. The second section briefly introduces the basic theory and the scope of the embedding scheme. The third section introduces the embedding algorithm of q-element payload using CE. The fourth part provides experimental results and constructive analysis of the performance of various embedded algorithms. Finally, Section 5 presents conclusions.
2. Preliminaries
- Cover for multitone images
The q-ary convolutional codes were applied to multitone images. The procedure for constructing the proposed embedding scheme involves the following aspects: (i) how to generate multiple-level tone images from a grayscale image; and (ii) how to construct an embedding system using an optimal decoding algorithm.
Error diffusion is a popular halftoning technique. Two-level representations are used in this technique to replace the original grayscale image or color image. This technique can be considered a generalization of multiple tones. Let and be coordinates of a grayscale image and a multiple-tone image, respectively, after quantizing point . The quantization error is expressed as . The multiple-tone point can be obtained using the following expression:
To transform the grayscale image into a multiple-tone image, an error filter is used in all filtering areas to obtain the final multiple-tone image. By contrast, in the recovery procedure, the multiple-tone image is recovered as the grayscale image. A low-pass filter is used to filter the points in the grayscale image and to obtain a continuous tone image.
- 2
- Embedding scheme and efficiency bound
The goal of the binary embedding scheme is to quantify the source limited by the distortion theory. The embedded model and extraction model are shown in Figure 1.
Figure 1.
Block diagram of a binary embedding system.
Under the assumption that a logo embedded into a cover is transmitted to the receiver, the optimal stego is provided by the embedder. Thus, a message , which is modified from , corresponds to syndrome . Given a cover , which is a Bernoulli-1/2 process in the binary symmetric source, it subtracts some toggle ; thus . Even though the embedder knows the cover , it cannot simply cancel this known interference due to the constraint that the average number of 1s cannot exceed , where is the block length and . We define the optimal or minimum quantized error , where denotes the Hamming distance between stego and cover . The optimal quantization error is the optimal modified vector such that the host and stego are of optimal quantization error. The rate distortion can be calculated as , where denotes the bound and denotes a binary entropy function, by an linear code with a code rate . Thus, the embedding rate is . Therefore, for a given good linear code with embedding rate , optimal distortion can be approached. Theoretically, the codeword of a linear code can be regarded as a quantized message set , with as the average distance between an arbitrary cover set . The upper bound of the embedding capacity can thus be expressed as follows:
where is the embedding rate corresponding to the optimal distortion . If a well-designed linear code exists, then the theoretical upper bound can be approached by an associated embedder. However, the major concern is to determine a parity-check matrix with a well-behaved linear code and a code rate . Furthermore, with the embedding rate requested in such a linear code , the aforementioned equation can then be expressed as follows:
For a binary symmetric source and an bit source sequence , the average distortion per bit is defined as follows:
where represents a quantized codeword existing in code , and is the average Hamming distortion between and per block. For an linear block code, the minimum average distortion can be expressed as follows:
where is the inverse function of the binary entropy function . The aforementioned equation is the rate-distortion function. The lower bound of average distortion for each bit in a code block is . The lower bound of each bit average distortion in blocks is displayed in Figure 2.
Figure 2.
Rate-distortion function.
When performing the binary data embedding of a sequence of length bits, the embedding efficiency is defined as follows:
3. Embedding Algorithm for Small and Large Payloads
Binary data embedding was achieved by using a standard array as follows: with linear code , we developed a standard array with a size of , as displayed in Figure 3.
Figure 3.
Standard array for the embedding algorithm.
Alternatively, the required coset leader can be determined precisely to perform binary data embedding or optimal embedding. An linear code can be characterized with a parity-check matrix of size as follows:
where the sequence is . Based on (7), the syndrome of the sequence is defined as . Furthermore, the set composed of all the sequences corresponding to identical is referred to as the coset of code and is defined as follows:
where denotes the coset leader in the standard array. The term can be derived through from an arbitrary sequence , and can be expressed using an ML decoding function as follows:
where represents the decoding function of the linear codes. Using ML decoding, the coset leader is added to to recover the codeword , which is closest to the sequence .
As displayed in Figure 3, for convolutional codes, it is necessary to determine the minimal toggle vector, , namely the coset leader, for a convolutional code in the vector domain to solve the equation , where . We considered the following simple embedding method. Using a systematic form CE in the vector domain, the equation can be used to solve the following expression:
Assuming that is a solution for , toggle vector can be determined immediately with and ; toggle can also be determined immediately. This section focuses on the efficient identification of the toggle vector using a systematic encoding technique. A symbol must be defined to describe the embedding of algorithms based on convolutional codes. Assuming that the convolutional code is a non-system generator matrix, it can be converted into a system generator matrix using basic row operations. Alternatively, the code can be expressed in a system recursive form. We use CE to embed binary messages as follows:
An embedding scheme with small payloads is used for numerous applications. However, for a case of CE codes with a low embedding rate, the trellis structure has high branches per state because of a large , which indicates that a complex mechanism is required when performing the Viterbi algorithm. To avoid this disadvantage, we constructed a CE code at a low embedding rate. The CE code was obtained through puncturing. In the time domain, we constructed an embedding rate , where is the puncturing period and is the number of deleted bits. Systematic recursive CE codes can be obtained by puncturing the output of a convolutional code with the puncturing matrix as follows:
Based on (10), we selected and to obtain the required . In puncturing matrix , if , the corresponding output bit from the CE code is embedded. Otherwise, the corresponding output bit from the CE code is deleted. To construct a systematic CE code by puncturing, the embedding algorithm must first locate the matrix corresponding to message with length bits in period . However, because of a systematic encoder, is set in assigned locations with respect to the set of indices . Here, of the second-row sequence in period . We located assigned indexes in corresponding to the location as follows:
where and is expressed as follows:
Thus, provided a cover matrix corresponding to , we can obtain the toggle matrix as in a puncturing period .
A notation must be defined to describe the embedding of a convolutional code-based algorithm. Here, CE is a nonsystematic generator matrix that can be translated into a systematic generator matrix using elementary row operations. Alternatively, it can be expressed in systematic recursive form. We used a CE to embed the binary message as follows:
A CE with a generator matrix is defined as follows:
where the information sequence is and the codeword sequence is . Codeword is closest to a random binary sequence with respect to the Hamming distance over the binary symmetric source. Convolutional code was used to generate the minimum error sequence from a quantization perspective as follows:
where the is the quantizer, which can be expressed as follows:
where and
The nearest neighboring quantizer , which we interpreted as the minimum error vector in quantizing using and , can be realized using the Viterbi algorithm for CE with a trellis structure. Finally, we defined the Voronoi cell of as the set
Consider the use of algebraic equations for the coset code of CEs. Furthermore, assume the shifted coset code of a convolutional code , where is defined as the sum of and a minimum error sequence . Subsequently, by using , an arbitrary binary sequence is quantized using coset code as follows:
where the shift sequence , that is, and , denotes the error sequence or coset leader sequence in quantizing toggle sequence by . It is assumed that cover sequence is uniformly distributed in ; moreover, the toggle sequence , which is obtained by subtracting message sequence from cover sequence , is also uniformly distributed. The minimum distance sequence between cover sequence and message sequence is equal to (16). By quantizing a random binary sequence by , an average quantized distortion level is represented as follows:
Similar to the linear block codes, the optimal toggle vector must be determined using convolutional systematic codes. A simple method similar to the systematic coding approach is data embedding using linear block codes with a coset vector associated with . The method in which the toggle vector was obtained in a systematic block code binary embedding was applied to the systematic CE binary embedding. The embedding procedure for systematic CE is as follows:
For a message syndrome sequence of length , it is necessary to determine sequence of length with syndrome as the linear code. For a special systematic convolutional code case, a generator matrix can be defined as follows:
where . The transposition of yields the following expression:
where is an matrix and embedded sequence and is derived using the following expression:
which is used to solve the following expression:
This equation is complex. Due to the systematic encoder, of size can be solved. Furthermore, the toggle sequence is obtained by subtracting from . Embedder quantizes the arbitrary toggle sequence to generate the optimal stego sequence as follows:
Finally, sequence closest to sequence , corresponding to syndrome , is derived as follows:
At the receiver, message sequence is extracted as follows: . To illustrate the nested CE algorithm, the following example is based on a systematic convolutional code to describe the embedding procedure, displayed as follows. Consider an embedded message sequence and a cover sequence . As the systematic convolutional codes are used, we easily obtain solution corresponding to . Subsequently, a systematic CE binary embedding is performed. Assuming that is the symbol intended for embedding, vector represents a sequence, that is, a (2, 1) systematic CE with the syndrome , , and the toggle sequence corresponding to and falling within the coset can be determined as follows:
The optimal toggle sequence corresponding to syndrome can be discovered by performing Viterbi decoding of as follows:
where is a Viterbi decoding function. The procedure for finding an optimal toggle sequence is displayed in as above. Finally, the stego sequence can be obtained as follows:
In the receiver, we reconstructed the message sequence as follows:
A crucial factor of CE codes is the method for determining the optimal generator matrix for large payloads.
4. Optimal Design for Q-Ary CEs
We used q-ary CE codes to achieve a high-performance ME. A structured code is commonly required in an efficient embedding algorithm. The ML decoding algorithm can be used to determine the optimal embedding algorithm. Optimal ML decoding (Viterbi decoding) can be performed with an existing CE. The CE features a sufficiently larger length of codeword compared with the block code, in which a short block of fixed length is used as a codeword set. A good CE code with a sufficiently large codeword length has cosets.
Next, the nonbinary embedding algorithm using CE codes, which involved modification of the cover samples by , was demonstrated. The algorithm was applied to an arbitrary selection of cover location. Although the proposed scheme used the embedding algorithm over , it can be used in the nonbinary domain for various applications for increasing embedding efficiency. We used q-ary random codes and searched the generator matrix to implement the nonbinary embedding algorithm. By applying embedding, we could generate embedding codes with optimal embedding efficiency and obtain numerous designs of generator polynomials of CE codes over , as displayed in Figure 4, Figure 5, Figure 6, Figure 7, Figure 8, Figure 9 and Figure 10.
Figure 4.
3-Ary, constraint length = 4, = 4.37.
Figure 5.
3-Ary, constraint length = 5, = 4.53.
Figure 6.
5-Ary, constraint length = 4, = 5.1.
Figure 7.
7-Ary, constraint length = 4, = 5.42.
Figure 8.
7-Ary, constraint length = 5, = 5.56.
Figure 9.
5-Ary, constraint length = 3, = 4.8.
Figure 10.
5-Ary, constraint length = 5, = 5.1.
Constructing a structured q-ary CE code with embedding efficiency close to the theoretical limit is a key open problem, involving the following aspects: (1) the embedding scheme requires a structured code of sufficient length, and must have an excellent parity check matrix or generator matrix; (2) the structured code is computationally efficient, and an effective encoding/decoding process has been developed based on the structured code.
5. Simulation Results
In this section, we describe how the packet form of q-ary CE codes is used in the application field. Consider the initial first packet data in Figure 11. This packet information is embedded into the least significant bit channel.
Figure 11.
Packet form in the application field.
Next, Figure 12 presents an example of the practical packet form.
Figure 12.
Example of the packet form.
Figure 13 displays the graphical user interface (GUI) of image steganography performed using MATLAB for embedding various images of different sizes over the proposed packet form. Total embedding involves cryptography with a random key. The embedding system includes the cyclic redundancy check (CRC) detection model. In the model, we examined the sensitivity of computer simulation results to the various images of different sizes to represent the rand errors and retransmission in the simulation. In the CRC model, it is assumed that the size of image capable for embedding can be protected for total length of the packet. The party check of CRC mode uses the following standard of ITU-IEEE:
Figure 13.
GUI of the image steganography system.
The GUI embedding system has the following characteristics: (1) image visualization: the image logo message stream is embedded with a q-ary cover and the packet of the embedding message is used in the security system. (2) Optimal design: some reordering or permutation of the image cover can be used to optimize q-ary CE codes. (3) Error detection mode: the component is used to protect the message packet from the attack channel.
Figure 13 indicates that the recovery logo messages in the GUI were the same as the original logo message. Thus, the transmitted and received message were the same. The square of CRC indicated simulation results that use the above party check polynomial to protect the packet, and the generator polynomial with various sizes was selected. The generator polynomials of q-ary convolutional with q = 3–7 and length 2–5 were used in the simulation. For convolutional codes with low embedding rate, the grid structure has a large number of branches per state, which means that a smaller number of metric operations are required to execute the Viterbi algorithm, and vice versa.
For constructing a high complexity code, a convolutional code with a high embedding rate is structured using a convolutional code with a low embedding rate through puncturing. Ultimately, the embedding rate of a (2, 1) convolutional code is maintained constant to design a convolutional code with various embedding rates through puncturing, and the complexity of the designed convolutional code is compared with that of the (2, 1) convolutional code. Moreover, the work of [16] proposed that the LDGM embedding codes and the embedding efficiency was the best performance for the study in steganography. Figure 14 shows the comparison of embedding efficiency between [16] and this study.
Figure 14.
The embedding efficiency between [16] and convolutional embedding codes.
6. Conclusions
A novel decoding method based on q-ary convolutional codes for ME in steganography was proposed in this study. Generally, the q-ary embedding scheme is applied to multiple-tone images. The q-ary level and the simulation of optimal embedding is run using the full search method. In q-ary CE, we used a q value of 3–7 and a code length of 3–5. The proposed method not only performed optimal decoding but also achieved optimal embedding efficiency for ME convolutional codes. The Viterbi decoding procedure was also used for this study. Moreover, the operation can be performed using a GUI for embedding applications.
Author Contributions
Conceptualization, J.-J.W., C.-Y.L., and H.-Y.C.; methodology, J.-J.W., C.-Y.L., and H.-Y.C.; software, J.-J.W. and H.-Y.C.; validation, S.-C.Y.; formal analysis, C.-Y.L.; investigation, S.-C.Y.; resources, C.-Y.L., S.-C.Y., H.-Y.C., and Y.-C.L.; data curation, J.-J.W. and H.-Y.C.; writing—original draft preparation, J.-J.W.; writing—review and editing, J.-J.W.; visualization, Y.-C.L.; supervision, S.-C.Y.; project administration, C.-Y.L.; funding acquisition, C.-Y.L. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Institutional Review Board Statement
Not applicable.
Informed Consent Statement
Not applicable.
Conflicts of Interest
The authors declare no conflict of interest.
References
- Crandall, R. Some Notes on Steganography. Post on Steganography Mailing List. 1998. Available online: http://os.inf.tu-dresden.de/west-feld/crandall.pdf (accessed on 18 June 2021).
- Bierbrauer, J. On Crandall’s Problem. 1998, unpublished. Available online: http://www.ws.binghamton.edu/fridrich/covcodes.pdf (accessed on 18 June 2021).
- Galand, F.; Kabatiansky, G. Information hiding by coverings. In Proceedings of the 2003 IEEE Information Theory Workshop (Cat. No.03EX674), Paris, France, 31 March–4 April 2003 ; pp. 151–154. [Google Scholar] [CrossRef] [Scilit]
- Zhang, W.; Wang, S.; Zhang, X. Improving embedding efficiency of covering codes for applications in steganography. IEEE Commun. Lett. 2007, 11, 680–682. [Google Scholar] [CrossRef] [Scilit]
- Bierbrauer, J.; Fridrich, J. Constructing good covering codes for applications in Steganography. In LNCS Transactions on Data Hiding and Multimedia Security; Springer: Berlin/Heidelberg, Germany, 2008; Volume 4920, pp. 1–22. [Google Scholar]
- Fridrich, J.; Filler, T. Practical methods for minimizing embedding impact in steganography. In Proceedings of the Security, Steganography, and Watermarking of Multimedia Contents IX, San Jose, CA, USA, 26 February 2007; Volume 6050, p. 2V3. [Google Scholar]
- Fridrich, J.; Soukal, D. Matrix embedding for large payloads. IEEE Trans. Inf. Forensics Secur. 2006, 1, 390–395. [Google Scholar] [CrossRef] [Scilit]
- Tseng, Y.-C.; Chen, Y.-Y.; Pan, H.-K. A secure data hiding scheme for binary images. IEEE Trans. Commun. 2002, 50, 1227–1231. [Google Scholar] [CrossRef]
- Li, R.Y.; Au, O.C.; Lai, K.K.; Yuk, C.K.; Lam, S.-Y. Data hiding with tree based parity check. In Proceedings of the IEEE International Conference, Beijing, China, 2–5 July 2007; pp. 635–638. [Google Scholar]
- Li, R.Y.; Au, O.C.; Yuk, C.K.M.; Yip, S.-K.; Lam, S.-Y. Halftone Image Data Hiding with Block-Overlapping Parity Check. In Proceedings of the 2007 IEEE International Conference on Acoustics, Speech and Signal Processing—ICASSP’ 07, Honolulu, HI, USA, 15–20 April 2007; Volume 2, pp. 193–196. [Google Scholar]
- Chen, J.; Zhu, Y.; Shen, Y.; Zhang, W. Efficient Matrix Embedding Based on Random Linear Codes. In Proceedings of the MINES 2010, Jiangsu, China, 4–6 November 2010; pp. 879–883. [Google Scholar]
- Gao, Y.; Li, X.; Yang, B. Employing optimal matrix for efficient matrix embedding. In Proceedings of the IIH-MSP2009, Kyoto, Japan, 12–14 September 2009; pp. 161–165. [Google Scholar]
- Sch¨onfeld, D.; Winkler, A. Embedding with syndrome coding based on BCH codes. In Proceedings of the ACM 8th Workshop on Multimedia and Security, Geneva, Switzerland, 26–27 September 2006; pp. 214–223. [Google Scholar]
- Wainwright, M.J. Sparse graph codes for side information and binning. IEEE Signal Process. Mag. 2007, 24, 47–57. [Google Scholar] [CrossRef]
- Wang, J.J.; Chen, H. An Adaptive Matrix Embedding Technique for Binary Hiding with an Efficient LIAE Algorithm. WSEAS Trans. Signal Process. 2012, 8, 64–75. [Google Scholar]
- Filler, T.; Fridrich, J. Binary quantization using Belief Propagation with decimation over factor graphs of LDGM codes. arXiv 2007, arXiv:0710.0192. [Google Scholar]
Publisher’s Note: MDPI stays neutral with regard to jurisdictional claims in published maps and institutional affiliations. |
© 2021 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).













