Systematic Benchmarking of DNA Sequence Encoding Strategies for Predicting Regulatory Effects of Non-Coding SNPs
Abstract
1. Introduction
2. Results
2.1. Overview of Benchmark Framework
- How interpretable are these encoding strategies?
- Are encoding strategies with higher encoding abundance (representation dimensions per base) more effective?
- What are the key factors influencing performance in downstream tasks?
2.2. Quantifying Multi-Dimensional Encoding Strategies for Non-Coding SNPs
2.3. Evaluating Performance in Predicting Regulatory Directions of Non-Coding SNPs on Gene Expression
2.4. Evaluating Performance in Predicting Impact Magnitude of Non-Coding SNPs on Gene Expression
2.5. Evaluating Performance in Predicting Impact Magnitude of Non-Coding SNPs on DNA Methylation Levels
2.6. Univariate Analysis: Investigating the Relationship Between Model Complexity, Tissue Type, and Performance
2.7. Cross-QTL Task Analysis: Exploring the Effectiveness of Encoding Strategies and Models
2.8. Guided Analysis for Practical Use: Case Study on meQTL Prediction
3. Discussion
4. Materials and Methods
4.1. Datasets
4.2. Encoding Strategies
- OneHot. The DNA sequence is represented by a one-hot encoded vector with four possible values corresponding to the four bases (A, T, C, G). If a base is ‘N’ (unknown), it is encoded as four zeros.
- DNABert2 [14]. DNABert2 is a pre-trained DNA language model, which was implemented using the official API (https://github.com/MAGICS-LAB/DNABERT_2, accessed on 24 December 2024) in this study. For DNABert2, we designed a position-wise average cutting and pooling strategy, aiming to maximize mutation representation. The processing steps are as follows: given a DNA sequence centered on either the TSS (for eQTL) or CpG site (for meQTL), the two outer segments are cut into lengths of 250 bp, while the center segment (501 bp) is centered around the TSS or CpG site. The remaining parts of the DNA sequence between the two outer segments and the central segment are evenly divided into 500 bp fragments. For each fragment, an embedding is generated, and embeddings are concatenated along the embedding dimension. Finally, the concatenated embeddings are averaged over the DNA sequence dimension to produce a one-dimensional vector of the same dimension as the language model embedding.
- GPN [15]. GPN is a pre-trained DNA language model, implemented using the official API (https://github.com/songlab-cal/gpn, accessed on 24 December 2024) in this study. The same position-wise average cutting and pooling strategy applied to DNABert2 is also used here.
- HyenaDNA [16]. HyenaDNA is another pre-trained DNA language model, implemented via the official API (https://github.com/HazyResearch/hyena-dna, accessed on 24 December 2024) in this study. Like DNABert2, it applies the position-wise average cutting and pooling strategy.
- NT [17]. NT is a pre-trained DNA language model, implemented using the official API (https://github.com/instadeepai/nucleotide-transformer, accessed on 24 December 2024) in this study. It also utilizes the position-wise average cutting and pooling strategy, similar to DNABert2.
- Enformer [13]. Enformer is a model designed for gene expression prediction from DNA sequences. In this study, we extracted DNA sequences of length 196,608 bp centered around the TSS or CpG sites to match Enformer’s input requirements. For both pre-mutated and post-mutated sequences, we obtained output features with a dimension of 5313. When calculating quantitative metrics, we directly utilized these output features. For Enformer representations, we applied PCA to reduce the high-dimensional output features before downstream modeling. This step was introduced to reduce computational cost and to make Enformer representations more tractable for comparison with other encoding strategies in the benchmark. We retained the top 10 principal components as a compact representation for downstream models. Because this compression may discard information from the original Enformer outputs, the Enformer-PCA results should be interpreted as performance under a computationally reduced setting rather than as the full predictive capacity of the original Enformer representation.
4.3. Feature Profile of Encoding Strategies
- Abundance. Abundance measures the amount of information required to describe a mutation. Specifically, it refers to the dimensionality of the embedding per base, indicating how detailed the representation is for each SNP. The final value is normalized by the sequence length , measured as the length of the DNA sequence centered on the mutation:where is the dimensionality of the -th embedding axis. For example, a 501 bp sequence encoded by one-hot contains four base channels at each nucleotide position, so the normalized abundance is reported as 4.0000. GPN generates a position-wise representation with 512 features for each base, so its abundance is reported as 512.0000. In contrast, for DNABert2, NT, and Enformer, the reported abundance values may not be identical to their raw hidden dimensions or output-track numbers, because their representations are affected by tokenization, sequence segmentation, pooling, or binned functional outputs. Thus, the abundance values in Table 1 should be interpreted as normalized representation density rather than raw embedding dimensionality.
- Differentiation. Differentiation quantifies the change in embedding information per unit length of DNA before and after mutation, effectively capturing the ability of an encoding strategy to amplify mutation effects. This difference is computed for each dimension of the embedding vector, , and averaged over the entire sequence length centered around the mutation. The formula can be expressed as:where represents the normalized difference in the -th dimension of the embedding vector before and after mutation.
- Time cost. Time cost represents the computational expense required to generate embeddings for a unit length of DNA, both before and after mutation. This metric reflects the efficiency of the encoding strategy, which is particularly critical for large-scale genomic datasets:where and represent the computational time required to generate the embedding vectors for the unmutated and mutated DNA sequences, respectively.
- Interpretability. Interpretability assesses whether the encoding preserves key biological information, such as the relative positions of bases or functional annotations of the DNA sequence. This dimension is crucial for understanding the biological implications of predictions made by models using these encodings.
- (1)
- Positional Interpretability: the clarity with which the embedding can be directly mapped to individual nucleotides in the DNA sequence. This means the position of each nucleotide in the original sequence can be unambiguously represented in the embedding.
- (2)
- Functional Interpretability: whether the embedding captures biological features that can be derived from experimental techniques such as sequencing. This includes elements like transcription factor binding sites, enhancers, or other functional genomic elements that are detectable through high-throughput sequencing methods.
- (3)
- Semantic Interpretability: refers to the extent to which embeddings learned from large-scale DNA sequence pretraining capture contextual and high-order sequence patterns that may support downstream prediction. Unlike functional interpretability, semantic interpretability does not indicate a direct correspondence to experimentally defined regulatory annotations but reflects whether the learned representations provide biologically relevant sequence context for modeling non-coding SNP effects.
4.4. Machine Learning and Deep Learning Models
4.5. Computational Environment
4.6. Performance Evaluation
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Farrow, S.L.; Gokuladhas, S.; Akan, I.S.; Nyaga, D.; Cooper, A.A.; Grand, R.S.; O’Sullivan, J.M. Dissecting genotype-specific effects of disease-associated genetic variants. iScience 2026, 29, 116143. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zeitlinger, J.; Roy, S.; Ay, F.; Mathelier, A.; Medina-Rivera, A.; Mahony, S.; Sinha, S.; Ernst, J. Perspective on recent developments and challenges in regulatory and systems genomics. Bioinform. Adv. 2025, 5, vbaf106. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Liu, Z.; Bao, Y.; Gu, A.; Song, W.; Lin, G.N. Predicting the regulatory impacts of noncoding variants on gene expression through epigenomic integration across tissues and single-cell landscapes. Nat. Comput. Sci. 2025, 5, 927–939. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Liu, Z.; Gu, A.; Bao, Y.; Lin, G.N. Epigenetic Impacts of Non-Coding Mutations Deciphered Through Pre-Trained DNA Language Model at Single-Cell Resolution. Adv. Sci. 2025, 12, e2413571. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Xu, C.; Liu, Q.; Zhou, J.; Xie, M.; Feng, J.; Jiang, T. Quantifying functional impact of non-coding variants with multi-task Bayesian neural network. Bioinformatics 2019, 36, 1397–1404. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhu, Y.; Tazearslan, C.; Suh, Y. Challenges and progress in interpretation of non-coding genetic variants associated with human disease. Exp. Biol. Med. 2017, 242, 1325–1334. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Michaelson, J.J.; Loguercio, S.; Beyer, A. Detection and interpretation of expression quantitative trait loci (eQTL). Methods 2009, 48, 265–276. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Smith, A.K.; Kilaru, V.; Kocak, M.; Almli, L.M.; Mercer, K.B.; Ressler, K.J.; Tylavsky, F.A.; Conneely, K.N. Methylation quantitative trait loci (meQTLs) are consistently detected across ancestry, developmental stage, and tissue type. BMC Genom. 2014, 15, 145. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ji, Y.; Zhou, Z.; Liu, H.; Davuluri, R.V. DNABERT: Pre-trained Bidirectional Encoder Representations from Transformers model for DNA-language in genome. Bioinformatics 2021, 37, 2112–2120. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhou, J.; Troyanskaya, O.G. Predicting effects of noncoding variants with deep learning–based sequence model. Nat. Methods 2015, 12, 931–934. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Nguyen, E.; Poli, M.; Durrant, M.G.; Kang, B.; Katrekar, D.; Li, D.B.; Bartie, L.J.; Thomas, A.W.; King, S.H.; Brixi, G.; et al. Sequence modeling and design from molecular to genome scale with Evo. Science 2024, 386, eado9336. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zeng, H.; Gifford, D.K. Predicting the impact of non-coding variants on DNA methylation. Nucleic Acids Res. 2017, 45, e99. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Avsec, Ž.; Agarwal, V.; Visentin, D.; Ledsam, J.R.; Grabska-Barwinska, A.; Taylor, K.R.; Assael, Y.; Jumper, J.; Kohli, P.; Kelley, D.R. Effective gene expression prediction from sequence by integrating long-range interactions. Nat. Methods 2021, 18, 1196–1203. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhou, Z.; Ji, Y.; Li, W.; Dutta, P.; Davuluri, R.V.; Liu, H. DNABERT-2: Efficient Foundation Model and Benchmark For Multi-Species Genome. arXiv 2023, arXiv:2306.15006. [Google Scholar]
- Benegas, G.; Batra, S.S.; Song, Y.S. DNA language models are powerful predictors of genome-wide variant effects. Proc. Natl. Acad. Sci. USA 2023, 120, e2311219120. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Nguyen, E.; Poli, M.; Faizi, M.; Thomas, A.W.; Sykes, C.B.; Wornow, M.; Patel, A.; Rabideau, C.; Massaroli, S.; Bengio, Y.; et al. HyenaDNA: Long-range genomic sequence modeling at single nucleotide resolution. In Proceedings of the 37th International Conference on Neural Information Processing Systems; Curran Associates Inc.: New Orleans, LA, USA, 2024; p. 1872. [Google Scholar]
- Dalla-Torre, H.; Gonzalez, L.; Mendoza-Revilla, J.; Lopez Carranza, N.; Grzywaczewski, A.H.; Oteri, F.; Dallago, C.; Trop, E.; de Almeida, B.P.; Sirelkhatim, H.; et al. Nucleotide Transformer: Building and evaluating robust foundation models for human genomics. Nat. Methods 2024, 22, 287–297. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Shu, L.; Tang, J.; Guan, X.; Zhang, D. A comprehensive survey of genome language models in bioinformatics. Brief. Bioinform. 2026, 27, bbaf724. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Abiola, O.; Angel, J.M.; Avner, P.; Bachmanov, A.A.; Belknap, J.K.; Bennett, B.; Blankenhorn, E.P.; Blizard, D.A.; Bolivar, V.; Brockmann, G.A.; et al. The nature and identification of quantitative trait loci: A community’s view. Nat. Rev. Genet. 2003, 4, 911–916. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Doerge, R.W. Mapping and analysis of quantitative trait loci in experimental populations. Nat. Rev. Genet. 2002, 3, 43–52. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chen, T.; Guestrin, C. XGBoost: A Scalable Tree Boosting System. In Proceedings of the Kdd’16: The 22nd Acm Sigkdd International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, 13–17 August 2016; pp. 785–794. [Google Scholar]
- Ke, G.; Meng, Q.; Finley, T.; Wang, T.; Chen, W.; Ma, W.; Ye, Q.; Liu, T.-Y. Lightgbm: A highly efficient gradient boosting decision tree. Adv. Neural Inf. Process. Syst. 2017, 30. [Google Scholar]
- Breiman, L. Random forests. Mach. Learn. 2001, 45, 5–32. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Z. Introduction to machine learning: K-nearest neighbors. Ann. Transl. Med. 2016, 4, 218. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Noble, W.S. What is a support vector machine? Nat. Biotechnol. 2006, 24, 1565–1567. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Botalb, A.; Moinuddin, M.; Al-Saggaf, U.M.; Ali, S.S.A. Contrasting Convolutional Neural Network (CNN) with Multi-Layer Perceptron (MLP) for Big Data Analysis. In Proceedings of the 2018 International Conference on Intelligent and Advanced System (ICIAS), Kuala Lumpur, Malaysia, 13–14 August 2018. [Google Scholar]
- Sherstinsky, A. Fundamentals of Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM) network. Phys. D Nonlinear Phenom. 2020, 404, 132306. [Google Scholar] [CrossRef] [Scilit]
- Islam, S.; Elmekki, H.; Elsebai, A.; Bentahar, J.; Drawel, N.; Rjoub, G.; Pedrycz, W. A comprehensive survey on applications of transformers for deep learning tasks. Expert Syst. Appl. 2024, 241, 122666. [Google Scholar] [CrossRef] [Scilit]
- Ruggiero, R.P.; Boissinot, S. Variation in base composition underlies functional and evolutionary divergence in non-LTR retrotransposons. Mob. DNA 2020, 11, 14. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Consortium, G. The GTEx Consortium atlas of genetic regulatory effects across human tissues. Science 2020, 369, 1318–1330. [Google Scholar] [CrossRef] [Scilit]
- Greenacre, M.; Groenen, P.J.F.; Hastie, T.; D’Enza, A.I.; Markos, A.; Tuzhilina, E. Principal component analysis. Nat. Rev. Methods Prim. 2022, 2, 100, Correction in Nat. Rev. Methods Prim. 2023, 3, 22. [Google Scholar] [CrossRef] [Scilit]
- Zhou, J.; Theesfeld, C.L.; Yao, K.; Chen, K.M.; Wong, A.K.; Troyanskaya, O.G. Deep learning sequence-based ab initio prediction of variant effects on expression and disease risk. Nat. Genet. 2018, 50, 1171–1179. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Villicaña, S.; Castillo-Fernandez, J.; Hannon, E.; Christiansen, C.; Tsai, P.-C.; Maddock, J.; Kuh, D.; Suderman, M.; Power, C.; Relton, C.; et al. Genetic impacts on DNA methylation help elucidate regulatory genomic processes. Genome Biol. 2023, 24, 176. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhou, Y.; Zhou, B.; Pache, L.; Chang, M.; Khodabakhshi, A.H.; Tanaseichuk, O.; Benner, C.; Chanda, S.K. Metascape provides a biologist-oriented resource for the analysis of systems-level datasets. Nat. Commun. 2019, 10, 1523. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Avsec, Ž.; Latysheva, N.; Cheng, J.; Novati, G.; Taylor, K.R.; Ward, T.; Bycroft, C.; Nicolaisen, L.; Arvaniti, E.; Pan, J.; et al. Advancing regulatory variant effect prediction with AlphaGenome. Nature 2026, 649, 1206–1218. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yang, C.; Liu, Y.; Ling, L.; Liu, Y.; Li, F.; Chen, C.; Wang, L.; Yu, F.; Qiao, L.; Zeng, X.; et al. Explicit dynamic cross-strand interactions for DNA sequence language modelling. Nat. Mach. Intell. 2026, 8, 880–900. [Google Scholar] [CrossRef] [Scilit]
- Cherif, I.L.; Kortebi, A. On using eXtreme Gradient Boosting (XGBoost) Machine Learning algorithm for Home Network Traffic Classification. In Proceedings of the 2019 Wireless Days (WD), Manchester, UK, 24–26 April 2019. [Google Scholar]




| Encoding Group | Encoding Strategy | Normalized Abundance (Representation Elements per Base) | Differentiation | Normalized Time Cost (CPU Seconds per Base) |
|---|---|---|---|---|
| Categorical | One-Hot | 4.0000 | 0.0040 | 1.9978 × 10−7 |
| Semantic | DNABert2 [14] | 469.5377 | 9.3144 | 1.0755 × 10−3 |
| Semantic | GPN [15] | 512.0000 | 1.3086 | 5.8716 × 10−3 |
| Semantic | HyenaDNA [16] | 128.5110 | 0.5234 | 6.8183 × 10−5 |
| Semantic | NT [17] | 1277.4451 | 4.6028 | 1.3727 × 10−2 |
| Functional | Enformer [13] | 9501.8922 | 0.0241 | 8.8423 × 10−2 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Jin, H.; Bao, Y.; Li, W.; Yang, C.; Wang, W.; Cai, W.; Liu, Z.; Lin, G.N. Systematic Benchmarking of DNA Sequence Encoding Strategies for Predicting Regulatory Effects of Non-Coding SNPs. Int. J. Mol. Sci. 2026, 27, 6657. https://doi.org/10.3390/ijms27156657
Jin H, Bao Y, Li W, Yang C, Wang W, Cai W, Liu Z, Lin GN. Systematic Benchmarking of DNA Sequence Encoding Strategies for Predicting Regulatory Effects of Non-Coding SNPs. International Journal of Molecular Sciences. 2026; 27(15):6657. https://doi.org/10.3390/ijms27156657
Chicago/Turabian StyleJin, Hui, Yihang Bao, Wenhao Li, Chengyi Yang, Weidi Wang, Wenxiang Cai, Zhe Liu, and Guan Ning Lin. 2026. "Systematic Benchmarking of DNA Sequence Encoding Strategies for Predicting Regulatory Effects of Non-Coding SNPs" International Journal of Molecular Sciences 27, no. 15: 6657. https://doi.org/10.3390/ijms27156657
APA StyleJin, H., Bao, Y., Li, W., Yang, C., Wang, W., Cai, W., Liu, Z., & Lin, G. N. (2026). Systematic Benchmarking of DNA Sequence Encoding Strategies for Predicting Regulatory Effects of Non-Coding SNPs. International Journal of Molecular Sciences, 27(15), 6657. https://doi.org/10.3390/ijms27156657

