Ensemble Deep Learning Models for Multi-Class DNA Sequence Classification: A Comparative Study of CNN, BiLSTM, and GRU Architectures
Abstract
1. Introduction
- (1)
- In this paper, we present a new ensemble deep learning model that combines CNN, BiLSTM, and GRU for classifying DNA sequences. This approach has not been extensively investigated in previous genomic research.
- (2)
- In this paper, we show that the above ensemble not only enhances the classification accuracy but can also balance precision and recall as well.
- (3)
- We offer a repeatable, at an architectural level implementation, specifying hyperparameters, to establish a baseline that future hybrid deep learning methods ought to be measured against in bioinformatics.
2. Background and Related Work
3. Exploratory Data Analysis
3.1. Dataset Composition
- G-protein coupled receptors (GPCR);
- Ion channels;
- Synthases;
- Tyrosine kinases;
- Tyrosine phosphatases.
3.2. Data Head
3.3. Class Distribution Visualization
3.4. Data Distribution
- G-protein coupled receptors (GPCR);
- Ion channels;
- Synthases;
- Tyrosine kinases;
- Tyrosine phosphatases.

4. Proposed Model
4.1. Pre-Processing
4.1.1. Experimental Configuration
4.1.2. Model Architecture
4.2. Experimental Setup and Reproducibility
4.3. GRU Model
4.4. Ensemble Model Strategy
- Independent Training: The training data independently trains CNN, BiLSTM, and GRU.
- Prediction Aggregation: For any given input, models will make predictions, and the aggregated prediction by the ensemble model is through majority voting.
- Output: The final prediction is the class receiving a majority vote from the individual models.
- Input: DNA sequences with corresponding labels.
- Preprocessing: Preprocess DNA sequences by normalizing and encoding.
- Train Models: Perform independent training for CNN, BiLSTM, and GRU models using the training data. Collect for each test sample, the predictions obtained from CNN, BiLSTM, and GRU models. Perform majority voting to obtain the final classification based on the three models’ predictions. Return the final classification result. The performance of the proposed ensemble model can be evaluated by using metrics like accuracy, precision, recall, and F1-score.
4.5. Evaluation Metrics
- Accuracy: Accuracy represents the overall percentage of correct predictions for any model across classes. It is calculated as the ratio between the number of correctly classified instances and the total number of instances within the dataset. This, in a very immediate and intuitive way, helps to understand the general performance of the model when the class distributions are fairly balanced. However, in this context of an imbalanced dataset, accuracy may be found misleading since a model can be accurate by repeatedly predicting a majority class, even though it does less well on minority classes. For this reason, accuracy is interpreted alongside complementary metrics such as precision, recall, and F1 score.
- Precision: Precision quantifies the reliability of the model’s positive predictions by measuring the proportion of true positive instances among all instances predicted as positive. A high precision value means the model commits very few false positive errors, which is of particular importance in applications where the cost of an incorrect positive prediction is high, such as defect detection or medical diagnosis. Because precision focuses on the quality of predictions and not on the coverage, precision helps assess how well the classifier avoids labelling negative samples as positives.
- Recall: The metric for evaluating the value of Recall, often termed sensitivity or positive predictive value, assesses how effectively a model detects all the positive data in the set while providing accurate predictions with its positive class values. The value for Recall is calculated by considering the ratio of the actual positive values in the set relative to all positive values in the set. An improved value for Recall points towards an effective coverage by the modeling process in such a way that there are no false cases detected for positive values in the set.
- F1 Score: The harmonic mean of precision and recall, providing a balanced measure of model performance.
4.5.1. CNN Confusion Matrix, ROC, and AUC
4.5.2. BiLSTM Confusion Matrix, ROC, and AUC
4.5.3. GRU Confusion Matrix, ROC, and AUC
5. Performance Evaluations
6. Discussion
6.1. Transferability to Other Application Domains
6.2. Limitations and Future Work
7. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Abbreviations
| ADAM | Adaptive Moment Estimation |
| AUC | Area Under the Curve |
| AUROC | Area Under the Receiver Operating Characteristic Curve |
| BiLSTM | Bidirectional Long Short-Term Memory |
| BZ2 | Bzip2 Compression Algorithm |
| CNN | Convolutional Neural Network |
| DNA | Deoxyribonucleic Acid |
| d-BM | Derivative Boyer–Moore |
| FLPM | Fast Local Pattern Matching |
| FNR | False Negative Rate |
| FPR | False Positive Rate |
| GRU | Gated Recurrent Unit |
| GWAS | Genome-Wide Association Study |
| KNN | k-Nearest Neighbors |
| LSTM | Long Short-Term Memory |
| LSTM + CNN | Long Short-Term Memory and Convolutional Neural Network Hybrid |
| LZ4 | Lempel–Ziv 4 Compression Algorithm |
| LZMA | Lempel–Ziv–Markov Chain Algorithm |
| ML | Machine Learning |
| MLP | Multi-Layer Perceptron |
| PAPM | Pattern-Aware Pattern Matching |
| ReLU | Rectified Linear Unit |
| RNA | Ribonucleic Acid |
| RNN | Recurrent Neural Network |
| ROC | Receiver Operating Characteristic |
| SVM | Support Vector Machine |
| XGBoost | Extreme Gradient Boosting |
References
- Bojanowski, P.; Grave, E.; Joulin, A.; Mikolov, T. Enriching word vectors with subword information. Trans. Assoc. Comput. Linguist. 2018, 6, 135–146. [Google Scholar] [CrossRef]
- Cho, K.; Van Merriënboer, B.; Gulcehre, C.; Bahdanau, D.; Bougares, F.; Schwenk, H.; Bengio, Y. Learning phrase representations using an RNN encoder-decoder for statistical machine translation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing, Doha, Qatar, 25–29 October 2014; Association for Computational Linguistics: Doha, Qatar, 2014; pp. 1724–1734. [Google Scholar]
- Chung, J.; Gulcehre, C.; Cho, K.; Bengio, Y. Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv 2014, arXiv:1412.3555. [Google Scholar] [CrossRef]
- Dietterich, T.G. Ensemble methods in machine learning. In Multiple Classifier Systems, Proceedings of the First International Workshop, MCS 2000, Cagliari, Italy, 21–23 June 2000; Springer: Berlin/Heidelberg, Germany, 2000; pp. 1–15. [Google Scholar] [CrossRef]
- Dharaniya, N.G.; Raaj, R.K.; Vikramathithan, M.; Vishal, P.; Yugavanan, S. DNA sequencing using a machine learning algorithm. Int. J. Res. Publ. Rev. 2024, 5, 12272–12274. [Google Scholar] [CrossRef]
- Dixit, P.; Prajapati, I.G. Machine Learning in Bioinformatics: A Novel Approach for DNA Sequencing. In Proceedings of the 2015 Fifth International Conference on Advanced Computing & Communication Technologies, Haryana, India, 21–22 February 2015. [Google Scholar] [CrossRef]
- Fatumo, S.; Chikowore, T.; Choudhury, A.; Ayub, M. Diversity in Genomic Studies: A Roadmap to Address the Imbalance. Nat. Med. 2022, 28, 243–250. [Google Scholar] [CrossRef]
- Garcia, M.; Patel, S. Deep Learning Models for DNA Sequence Classification: Applications and Challenges. IEEE Trans. Comput. Biol. Bioinform. 2023, 20, 77–89. [Google Scholar]
- Hastie, T.; Tibshirani, R.; Friedman, J.H. The Elements of Statistical Learning: Data Mining, Inference, and Prediction, 2nd ed.; Springer: Berlin/Heidelberg, Germany, 2022. [Google Scholar]
- Hochreiter, S.; Schmidhuber, J. Long short-term memory. Neural Comput. 1997, 9, 1735–1780. [Google Scholar] [CrossRef] [PubMed]
- Hossain, P.S.; Kim, K.; Uddin, J.; Samad, A.; Choi, K. Enhancing Taxonomic Categorization of DNA Sequences with Deep Learning: A Multi-Label Approach. Bioengineering 2023, 10, 1293. [Google Scholar] [CrossRef] [PubMed]
- Hamed, B.A.; Ibrahim, O.A.S.; El-Hafeez, T.A. Optimizing classification efficiency with machine learning techniques for pattern matching. J. Big Data 2023, 10, 124. [Google Scholar] [CrossRef]
- Hu, W.; Li, Y.; Wu, Y.; Guan, L.; Li, M. A Deep Learning Model for DNA Enhancer Prediction based on Nucleotide Position Aware Feature Encoding. iScience 2024, 27, 110030. [Google Scholar] [CrossRef] [PubMed]
- Li, W.; Zhang, H.; Wang, Q. Application of GRU networks for predicting protein secondary structure. Comput. Biol. Chem. 2020, 85, 107–115. [Google Scholar]
- Li, X.; Liu, S.; Sun, Z. A Survey of Machine Learning Models for DNA Sequence Classification. J. Comput. Biol. 2020, 27, 503–518. [Google Scholar]
- Li, X.; Zhang, Z.; Lu, Y. BiLSTM network-based deep learning model for human activity recognition. IEEE Access 2018, 6, 29156–29164. [Google Scholar]
- Liu, J.; Zhang, W.; Zhuang, Y. A novel deep learning model for the classification of gene sequences using convolutional neural networks. Bioinformatics 2020, 36, 3415–3421. [Google Scholar]
- Miah, J.; Ayon, E.H.A.; Ghosh, B.P.G.; Mia, T.; Badruddowza; Sarker, S.U.S.; Islam, T. Enhancing Viral DNA Sequence Classification Using Hybrid Deep Learning Models and Genetic Algorithm Optimization. Available online: https://ssrn.com/abstract=4692259 (accessed on 11 January 2024).
- Mittal, S.; Jena, M.K. Machine learning empowered next-generation DNA sequencing: Perspective and prospectus. Chem. Sci. 2024, 15, 12169–12188. [Google Scholar] [CrossRef]
- Mathur, G.; Pandey, A.; Goyal, S. A comprehensive tool for rapid and accurate prediction of disease using a DNA sequence classifier. J. Ambient. Intell. Humaniz. Comput. 2022, 14, 13869–13885. [Google Scholar] [CrossRef] [PubMed]
- Nguyen, T.H.; Zhao, Y. Challenges and Opportunities in DNA Sequence Pattern Recognition: A Survey. IEEE Trans. Comput. Biol. Bioinform. 2022, 19, 350–363. [Google Scholar]
- O’Reilly, K.; Jones, D. Innovations in DNA Sequence Analysis: Addressing Gaps in Geometric and Correlation-Based Approaches. In Proceedings of the 2023 European Conference on Bioinformatics (ECBio), Lyon, France, 23–27 July 2023; pp. 77–85. [Google Scholar]
- Ashraf, S.; Ahmad, M.; Aslam, N. Analysis of DNA sequence classification using CNN and hybrid models. BMC Bioinform. 2021, 22, 1835056. [Google Scholar]
- Kaur, H.; Singh, A.; Malhotra, P. Comparison of deep learning approaches for DNA-binding protein classification using CNN and hybrid models. In Proceedings of the International Conference on Machine Intelligence and Data Science Applications; Springer: Singapore, 2024; pp. 123–135. Available online: https://link.springer.com/chapter/10.1007/978-981-99-5881-8_7 (accessed on 28 May 2025).
- Khan, M.A.; Tariq, U.; Sharif, M. SaPt-CNN-LSTM-AR-EA: A hybrid ensemble learning framework for time series-based multivariate DNA sequence prediction. PeerJ Comput. Sci. 2023, 9, e16192. Available online: https://peerj.com/articles/16192 (accessed on 28 May 2025).
- Li, Z.; Subasri, V.; Shen, Y.; Li, D.; Zhao, Y.; Stan, G.-B.; Shan, C. Omni-DNA: A Unified Genomic Foundation Model for Cross-Modal and Multi-Task Learning. arXiv 2025, arXiv:2502.03499. [Google Scholar] [CrossRef]
- Duan, Q.; Huang, B.; Song, Z.; Lehmann, I.; Gu, L.; Eils, R.; Wild, B. JanusDNA: A Powerful Bi-directional Hybrid DNA Foundation Model. arXiv 2025, arXiv:2505.17257. [Google Scholar] [CrossRef]
- Awasthi, R.; Mend Mend Arachchige, G.S.; Zhu, X. Unsupervised Evaluation of Pre-Trained DNA Language Model Embeddings. BMC Genom. 2025, 26, 710. [Google Scholar] [CrossRef] [PubMed]
- Benegas, G.; Albors, C.; Aw, A.J.; Ye, C.; Song, Y.S. A DNA Language Model Based on Multispecies Alignment Predicts the Effects of Genome-Wide Variants. Nat. Biotechnol. 2025, 43, 1960–1965. [Google Scholar] [CrossRef]
- Wang, S. Graph neural network–driven text classification for fire-door defect inspection in pre-completion construction. Sci. Rep. 2025, 15, 44382. [Google Scholar] [CrossRef] [PubMed]
- Zhang, J.; Fleyeh, H.; Wang, X.; Lu, M. Dynamic building defect categorization through enhanced unsupervised text classification with domain-specific corpus embedding methods. Autom. Constr. 2024, 157, 105182. [Google Scholar] [CrossRef]









| Sequence | Class | |
|---|---|---|
| 0 | ATGCCCCAACTAAATACTACCGTATGGCCCACCATAATTACCCCCA… | 4 |
| 1 | ATGAACGAAAATCTGTTCGCTTCATTCATTGCCCCCACAATCCTAG… | 4 |
| 2 | ATGTGTGGCATTTGGGCGCTGTTTGGCAGTGATGATTGCCTTTCTG… | 3 |
| 3 | ATGTGTGGCATTTGGGCGCTGTTTGGCAGTGATGATTGCCTTTCTG… | 3 |
| 4 | ATGCAACAGCATTTTGAATTTGAATACCAGACCAAAGTGGATGGTG… | 3 |
| Count | 4380.000000 |
| mean | 3.504566 |
| std | 2.132134 |
| min | 0.000000 |
| 25% | 2.000000 |
| 50% | 4.000000 |
| 75% | 6.000000 |
| max | 6.000000 |
| Model | Accuracy (%) | Precision (%) | Recall (%) | F1 (%) | AUC |
|---|---|---|---|---|---|
| CNN | 80.60 | 81.60 | 80.60 | 83.10 | 0.83 |
| BiLSTM | 90.98 | 73.09 | 82.83 | 77.99 | 0.90 |
| GRU | 81.20 | 74.20 | 80.00 | 76.00 | 0.82 |
| Ensemble | 90.60 | 91.00 | 91.00 | 91.00 | 0.95 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Tabane, E.; Mnkandla, E.; Wang, Z. Ensemble Deep Learning Models for Multi-Class DNA Sequence Classification: A Comparative Study of CNN, BiLSTM, and GRU Architectures. Appl. Sci. 2026, 16, 1545. https://doi.org/10.3390/app16031545
Tabane E, Mnkandla E, Wang Z. Ensemble Deep Learning Models for Multi-Class DNA Sequence Classification: A Comparative Study of CNN, BiLSTM, and GRU Architectures. Applied Sciences. 2026; 16(3):1545. https://doi.org/10.3390/app16031545
Chicago/Turabian StyleTabane, Elias, Ernest Mnkandla, and Zenghui Wang. 2026. "Ensemble Deep Learning Models for Multi-Class DNA Sequence Classification: A Comparative Study of CNN, BiLSTM, and GRU Architectures" Applied Sciences 16, no. 3: 1545. https://doi.org/10.3390/app16031545
APA StyleTabane, E., Mnkandla, E., & Wang, Z. (2026). Ensemble Deep Learning Models for Multi-Class DNA Sequence Classification: A Comparative Study of CNN, BiLSTM, and GRU Architectures. Applied Sciences, 16(3), 1545. https://doi.org/10.3390/app16031545

