A Survey on Data Compression Methods for Biological Sequences
AbstractThe ever increasing growth of the production of high-throughput sequencing data poses a serious challenge to the storage, processing and transmission of these data. As frequently stated, it is a data deluge. Compression is essential to address this challenge—it reduces storage space and processing costs, along with speeding up data transmission. In this paper, we provide a comprehensive survey of existing compression approaches, that are specialized for biological data, including protein and DNA sequences. Also, we devote an important part of the paper to the approaches proposed for the compression of different file formats, such as FASTA, as well as FASTQ and SAM/BAM, which contain quality scores and metadata, in addition to the biological sequences. Then, we present a comparison of the performance of several methods, in terms of compression ratio, memory usage and compression/decompression time. Finally, we present some suggestions for future research on biological data compression. View Full-Text
Share & Cite This Article
Hosseini, M.; Pratas, D.; Pinho, A.J. A Survey on Data Compression Methods for Biological Sequences. Information 2016, 7, 56.
Hosseini M, Pratas D, Pinho AJ. A Survey on Data Compression Methods for Biological Sequences. Information. 2016; 7(4):56.Chicago/Turabian Style
Hosseini, Morteza; Pratas, Diogo; Pinho, Armando J. 2016. "A Survey on Data Compression Methods for Biological Sequences." Information 7, no. 4: 56.
Note that from the first issue of 2016, MDPI journals use article numbers instead of page numbers. See further details here.