iBitter-HF: A Method for Bitter Peptide Sequence Identification Based on Hybrid Feature Embedding
Abstract
1. Introduction
2. Materials and Methods
2.1. Dataset
2.2. Feature Extraction
2.2.1. Hand-Crafted Feature Extraction
- (1)
- Composition and sequence-pattern descriptors.
- (2)
- Physicochemical-property and substitution-matrix descriptors.
2.2.2. UniRep Deep Representation Learning Feature Extraction
2.2.3. Construction of the Hybrid Feature Embedding
2.3. Machine-Learning Models
2.4. LGBM-Based Feature-Importance Ranking and Model Optimization
2.5. Evaluation Metrics
2.6. t-SNE Visualization Procedure
2.7. Webserver Implementation and Usage
3. Results and Discussion
3.1. Effect of Feature Hybridization on Bitter Peptide Identification
3.2. Comparison of Machine-Learning Models Under Hybrid Feature Embedding
3.3. Model Optimization Based on Feature-Importance Ranking

3.4. Dimensionality Reduction Visualization of the Feature Space

3.5. Comparison with Existing Bitter Peptide Predictors
4. Conclusions
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Saha, B.C.; Hayashi, K. Debittering of protein hydrolyzates. Biotechnol. Adv. 2001, 19, 355–370. [Google Scholar] [CrossRef] [Scilit]
- Maehashi, K.; Huang, L. Bitter peptides and bitter taste receptors. Cell. Mol. Life Sci. 2009, 66, 1661–1671. [Google Scholar] [CrossRef] [Scilit]
- Luo, Y.; Shi, L.; Li, Y.; Zhuang, A.; Gong, Y.; Liu, L.; Lin, C. From intention to implementation: Automating biomedical research via LLMs. Sci. China-Inf. Sci. 2025, 68, 170105. [Google Scholar] [CrossRef] [Scilit]
- Aluri, S.; Imambi, S.S. Brain tumour classification using MRI images based on lenet with golden teacher learning optimization. Netw.-Comput. Neural Syst. 2024, 35, 27–54. [Google Scholar] [CrossRef] [Scilit]
- Ren, X.; Wei, J.; Luo, X.; Liu, Y.; Li, K.; Zhang, Q.; Gao, X.; Yan, S.; Wu, X.; Jiang, X.; et al. HydrogelFinder: A Foundation Model for Efficient Self-Assembling Peptide Discovery Guided by Non-Peptidal Small Molecules. Adv. Sci. 2024, 11, 2400829. [Google Scholar] [CrossRef] [Scilit]
- Lai, L.; Liu, Y.; Song, B.; Li, K.; Zeng, X. Deep Generative Models for Therapeutic Peptide Discovery: A Comprehensive Review. ACM Comput. Surv. 2025, 57, 155. [Google Scholar] [CrossRef] [Scilit]
- Liu, S.; Shi, T.; Yu, J.; Li, R.; Lin, H.; Deng, K.J. Research on Bitter Peptides in the Field of Bioinformatics: A Comprehensive Review. Int. J. Mol. Sci. 2024, 25, 9844. [Google Scholar] [CrossRef] [Scilit]
- Charoenkwan, P.; Yana, J.; Schaduangrat, N.; Nantasenamat, C.; Hasan, M.M.; Shoombuatong, W. IBitter-SCM: Identification and characterization of bitter peptides using a scoring card method with propensity scores of dipeptides. Genomics 2020, 112, 2813–2822. [Google Scholar] [CrossRef] [Scilit]
- Charoenkwan, P.; Nantasenamat, C.; Hasan, M.M.; Moni, M.A.; Lio, P.; Shoombuatong, W. IBitter-Fuse: A novel sequence-based bitter peptide predictor by fusing multi-view features. Int. J. Mol. Sci. 2021, 22, 8958. [Google Scholar] [CrossRef] [Scilit]
- Charoenkwan, P.; Nantasenamat, C.; Hasan, M.M.; Manavalan, B.; Shoombuatong, W. BERT4Bitter: A bidirectional encoder representations from transformers (BERT)-based model for improving the prediction of bitter peptides. Bioinformatics 2021, 37, 2556–2562. [Google Scholar] [CrossRef] [Scilit]
- Jiang, J.; Lin, X.; Jiang, Y.; Jiang, L.; Lv, Z. Identify bitter peptides by using deep representation learning features. Int. J. Mol. Sci. 2022, 23, 7877. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Y.F.; Wang, Y.H.; Gu, Z.F.; Pan, X.R.; Li, J.; Ding, H.; Zhang, Y.; Deng, K.J. Bitter-RF: A random forest machine model for recognizing bitter peptides. Front. Med. 2023, 10, 1052923. [Google Scholar] [CrossRef] [Scilit]
- Lv, J.; Geng, A.; Pan, Z.; Wei, L.; Zou, Q.; Zhang, Z.; Cui, F. IBitter-GRE: A novel stacked bitter peptide predictor with ESM-2 and multi-view features. J. Mol. Biol. 2025, 437, 169005. [Google Scholar] [CrossRef] [Scilit]
- Ahmad, S.; Ahsan, M.; Asim, M.N.; Dengel, A.; Malik, M.I. IBitter-Stack: A multi-representation ensemble learning model for accurate bitter peptide identification. J. Mol. Biol. 2025, 437, 169448. [Google Scholar] [CrossRef] [Scilit]
- Ren, F.; Yang, L.; Zhang, M.; Tan, Y.; Fang, Y.; Lei, H.; Tan, H.; Wei, X. Developing machine learning-driven QSAR models for predicting bitter activity and bitterness thresholds of oligopeptides. Food Chem. 2026, 512, 148934. [Google Scholar] [CrossRef] [Scilit]
- Qiao, J.; Gao, W.; Jin, J.; Wang, D.; Guo, X.; Manavalan, B.; Wei, L. Molecular pretraining models towards molecular property prediction. Sci. China Inf. Sci. 2025, 68, 170104. [Google Scholar] [CrossRef] [Scilit]
- Chen, Z.; Zhao, P.; Li, C.; Li, F.; Xiang, D.; Chen, Y.Z.; Akutsu, T.; Daly, R.J.; Webb, G.I.; Zhao, Q.; et al. ILearnPlus: A comprehensive and automated machine-learning platform for nucleic acid and protein sequence analysis, prediction and visualization. Nucleic Acids Res. 2021, 49, e60. [Google Scholar] [CrossRef] [Scilit]
- Chen, Z.; Zhao, P.; Li, F.; Leier, A.; Marquez-Lago, T.T.; Wang, Y.; Webb, G.I.; Smith, A.I.; Daly, R.J.; Chou, K.C.; et al. IFeature: A Python package and web server for features extraction and selection from protein and peptide sequences. Bioinformatics 2018, 34, 2499–2502. [Google Scholar] [CrossRef] [Scilit]
- Qiao, J.; Jin, J.; Wang, D.; Teng, S.; Zhang, J.; Yang, X.; Liu, Y.; Wang, Y.; Cui, L.; Zou, Q.; et al. A self-conformation-aware pre-training framework for molecular property prediction with substructure interpretability. Nat. Commun. 2025, 16, 4382. [Google Scholar] [CrossRef] [Scilit]
- Chen, Y.; Wang, Z.; Wang, J.; Chu, Y.; Zhang, Q.; Li, Z.A.; Zeng, X. Self-supervised learning in drug discovery. Sci. China Inf. Sci. 2025, 68, 170103. [Google Scholar] [CrossRef] [Scilit]
- Alley, E.C.; Khimulya, G.; Biswas, S.; AlQuraishi, M.; Church, G.M. Unified rational protein engineering with sequence-based deep representation learning. Nat. Methods 2019, 16, 1315–1322. [Google Scholar] [CrossRef] [Scilit]
- Xie, X.; Wu, C.; Qi, Y.; Liu, S.; Huang, J.; Lyu, H.; Dao, F.; Lin, H. BertADP: A fine-tuned protein language model for anti-diabetic peptide prediction. BMC Biol. 2025, 23, 210. [Google Scholar] [CrossRef] [Scilit]
- Riasat, H.; Alkhalifah, T.; Alturise, F.; Khan, Y.D. XCPP: A Multi-model Explainable Deep Learning Framework for Accurate Identification of Cell-Penetrating Peptides from Structured Sequence Features. Curr. Drug Targets 2026. [Google Scholar] [CrossRef] [Scilit]
- Kawashima, S.; Pokarowski, P.; Pokarowska, M.; Kolinski, A.; Katayama, T.; Kanehisa, M. AAindex: Amino acid index database, progress report 2008. Nucleic Acids Res. 2008, 36, D202–D205. [Google Scholar] [CrossRef] [Scilit]
- Henikoff, S.; Henikoff, J.G. Amino acid substitution matrices from protein blocks. Proc. Natl. Acad. Sci. USA 1992, 89, 10915–10919. [Google Scholar] [CrossRef] [Scilit]
- Duan, Z.X.; Liang, Y.F.; Xiu, X.; Ma, W.J.; Mei, H. ResUbiNet: A Novel Deep Learning Architecture for Ubiquitination Site Prediction. Curr. Genom. 2025, 26, 302–311. [Google Scholar] [CrossRef] [Scilit]
- Cox, D.R. The regression analysis of binary sequences. J. R. Stat. Soc. Ser. B Stat. Methodol. 1958, 20, 215–232. [Google Scholar] [CrossRef] [Scilit]
- Wu, Y.; Cao, F.; Li, J. Multinomial Logistic Regression with Adaptive Regularization for Cancer Subtype Classification via Multi-omics Data. Curr. Bioinform. 2025, 20, 506–521. [Google Scholar] [CrossRef] [Scilit]
- Cover, T.; Hart, P. Nearest neighbor pattern classification. IEEE Trans. Inf. Theory 1967, 13, 21–27. [Google Scholar] [CrossRef] [Scilit]
- Hand, D.J.; Yu, K. Idiot’s Bayes—Not so stupid after all? Int. Stat. Rev. 2001, 69, 385–398. [Google Scholar] [CrossRef] [Scilit]
- Cortes, C.; Vapnik, V. Support-vector networks. Mach. Learn. 1995, 20, 273–297. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Zhai, Y.; Ding, Y.; Zou, Q. SBSM-Pro: Support bio-sequence machine for proteins. Sci. China-Inf. Sci. 2024, 67, 212106. [Google Scholar] [CrossRef] [Scilit]
- Breiman, L. Random forests. Mach. Learn. 2001, 45, 5–32. [Google Scholar] [CrossRef] [Scilit]
- Su, R.; Liu, X.; Wei, L.; Zou, Q. Deep-Resp-Forest: A deep forest model to predict anti-cancer drug response. Methods 2019, 166, 91–102. [Google Scholar] [CrossRef] [Scilit]
- Friedman, J.H. Greedy function approximation: A gradient boosting machine. Ann. Stat. 2001, 29, 1189–1232. [Google Scholar] [CrossRef] [Scilit]
- Bentéjac, C.; Csörgő, A.; Martínez-Muñoz, G. A comparative analysis of gradient boosting algorithms. Artif. Intell. Rev. 2021, 54, 1937–1967. [Google Scholar] [CrossRef] [Scilit]
- Adler, A.I.; Painsky, A. Feature importance in gradient boosting trees with cross-validation feature selection. Entropy 2022, 24, 687. [Google Scholar] [CrossRef] [Scilit]
- Yan, J.; Xu, Y.; Cheng, Q.; Jiang, S.; Wang, Q.; Xiao, Y.; Ma, C.; Yan, J.; Wang, X. LightGBM: Accelerated genomically designed crop breeding through ensemble learning. Genome Biol. 2021, 22, 271. [Google Scholar] [CrossRef] [Scilit]
- Liang, C.; Wang, D.; Zhang, H.; Zhang, S.; Guo, F. Robust Tensor Subspace Learning for Incomplete Multi-View Clustering. IEEE Trans. Knowl. Data Eng. 2024, 36, 6934–6948. [Google Scholar] [CrossRef] [Scilit]
- Dou, M.; Tang, J.; Tiwari, P.; Ding, Y.; Guo, F. Drug-Drug Interaction Relation Extraction Based on Deep Learning: A Review. ACM Comput. Surv. 2024, 56, 158. [Google Scholar] [CrossRef] [Scilit]
- Liang, C.; Wang, L.; Liu, L.; Zhang, H.; Guo, F. Multi-view unsupervised feature selection with tensor robust principal component analysis and consensus graph learning. Pattern Recognit. 2023, 141, 141. [Google Scholar] [CrossRef] [Scilit]
- Dao, F.; Lebeau, B.; Ling, C.C.Y.; Yang, M.; Xie, X.; Fullwood, M.J.; Lin, H.; Lyu, H. RepliChrom: Interpretable machine learning predicts cancer-associated enhancer-promoter interactions using DNA replication timing. iMeta 2025, 4, e70052. [Google Scholar] [CrossRef] [Scilit]
- Huang, Z.; Xiao, Z.; Ao, C.; Guan, L.; Yu, L. Computational approaches for predicting drug-disease associations: A comprehensive review. Front. Comput. Sci. 2025, 19, 195909. [Google Scholar] [CrossRef] [Scilit]
- Huang, Z.; Guo, X.; Qin, J.; Gao, L.; Ju, F.; Zhao, C.; Yu, L. Accurate RNA velocity estimation based on multibatch network reveals complex lineage in batch scRNA-seq data. BMC Biol. 2024, 22, 290. [Google Scholar] [CrossRef] [Scilit]
- Guo, X.; Huang, Z.; Ju, F.; Zhao, C.; Yu, L. Highly Accurate Estimation of Cell Type Abundance in Bulk Tissues Based on Single-Cell Reference and Domain Adaptive Matching. Adv. Sci. 2024, 11, 2306329. [Google Scholar] [CrossRef] [Scilit]
- Kobak, D.; Berens, P. The art of using t-SNE for single-cell transcriptomics. Nat. Commun. 2019, 10, 5416. [Google Scholar] [CrossRef] [Scilit]



Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Yan, F.; Xiang, S.; Tang, Y.; Kuang, Z.; Liu, H.; Luo, X.; Lv, Z. iBitter-HF: A Method for Bitter Peptide Sequence Identification Based on Hybrid Feature Embedding. Foods 2026, 15, 3016. https://doi.org/10.3390/foods15173016
Yan F, Xiang S, Tang Y, Kuang Z, Liu H, Luo X, Lv Z. iBitter-HF: A Method for Bitter Peptide Sequence Identification Based on Hybrid Feature Embedding. Foods. 2026; 15(17):3016. https://doi.org/10.3390/foods15173016
Chicago/Turabian StyleYan, Feng, Shicheng Xiang, Yi Tang, Zhengran Kuang, Hengxi Liu, Ximei Luo, and Zhibin Lv. 2026. "iBitter-HF: A Method for Bitter Peptide Sequence Identification Based on Hybrid Feature Embedding" Foods 15, no. 17: 3016. https://doi.org/10.3390/foods15173016
APA StyleYan, F., Xiang, S., Tang, Y., Kuang, Z., Liu, H., Luo, X., & Lv, Z. (2026). iBitter-HF: A Method for Bitter Peptide Sequence Identification Based on Hybrid Feature Embedding. Foods, 15(17), 3016. https://doi.org/10.3390/foods15173016

