Self-Attention over Parallel Dense Embeddings for High-Dimensional Omic Data
Abstract
1. Introduction
2. Materials and Methods
2.1. Self-Attention over Parallel Dense Embeddings
- 1.
- Input Layer: Each input sample consists of p features (columns). The input layer directly receives these features.
- 2.
- Parallel Dense Layers: The input layer is mapped to q parallel dense embedding branches of dimension d. These hyperparameters are chosen to provide a compact latent representation, thereby reducing the computational burden and parameter count relative to a standard fully connected network. These layers transform the input into multiple representations, producing q vectors of dimension d. Formally:where ReLU is the Rectified Linear Unit activation function used to introduce nonlinearity, denotes the weight matrix of the i-th parallel branch, x the input vector, and the bias vector of the i-th parallel branch.
- 3.
- Reshaping into Sequence (Embedding Matrix): The outputs of the parallel layers are stacked to form a matrix H of shape , which can be interpreted as a sequence of q “tokens” with embedding dimension d. This matrix acts as the input embedding for the attention mechanism.
- 4.
- Multi-head Attention Layer: We apply Transformer-style self-attention:Let denote the matrix whose rows are the embeddings produced by the q parallel branches. These embeddings are treated as the input tokens of the multi-head self-attention layer and are projected into query, key, and value representations using the learnable projection matrices , , and :where , , and are the learnable query, key, and value projection matrices for the i-th attention head and h denotes the number of attention heads.The outputs of the h attention heads are concatenated and projected through the learnable output projection matrix :where Concat(·) concatenates the outputs of all attention heads and denotes the learnable output projection matrix. This allows the model to capture interactions between the parallel dense representations.
- 5.
- Residual Connections and Normalization: We employ residual connections and layer normalization after the attention layer, as in standard Transformers [3].
- 6.
- Feed-Forward Network (FFN): Following the attention layer, each token passes through a small feed-forward network (two dense layers with a ReLU activation in between), again with residual connections and layer normalization [3].
- 7.
- Output Layer: The resulting sequence is flattened into a single vector, which is then fed into a final dense layer with sigmoid activation for binary classification (or softmax for multi-class tasks).
2.2. Architectures
2.3. Simulation Scenarios
2.3.1. Latent Variable and Class Assignment
2.3.2. Gene Expression Model
- a continuous effect ,
- a class-specific effectwhere .
2.3.3. Negative Binomial Observation Model
2.3.4. Calibration of Classification Difficulty
2.3.5. Hyperparameter Selection and Ablation Study
| Algorithm 1 Nested hyperparameter selection and evaluation protocol |
|
2.4. Real Data
3. Results
3.1. Simulation Results
3.1.1. Relevant Features Rate
3.1.2. Lightweight Architectures
3.2. Real Data Results
3.2.1. Comparative Performance
3.2.2. Attention-Based Model Explainability
4. Discussion
5. Conclusions
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Gorishniy, Y.; Kotelnikov, A.; Babenko, A. TabM: Advancing Tabular Deep Learning with Parameter-Efficient Ensembling. In Proceedings of the International Conference on Learning Representations; ICLR: Appleton, WI, USA, 2025. [Google Scholar]
- O’Brien Quinn, H.; Sedky, M.; Francis, J.; Streeton, M. Literature Review of Explainable Tabular Data Analysis. Electronics 2024, 13, 3806. [Google Scholar] [CrossRef] [Scilit]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention Is All You Need. In Proceedings of the Advances in Neural Information Processing Systems; American Institute of Physics: College Park, MD, USA, 2017; Volume 30. [Google Scholar]
- Niu, Z.; Zhong, G.; Yu, H. A Review on the Attention Mechanism of Deep Learning. Neurocomputing 2021, 452, 48–62. [Google Scholar] [CrossRef] [Scilit]
- Duru, I.; Sunar, A.S. Transformer and Pre-Transformer Model-Based Sentiment Prediction with Various Embeddings: A Case Study on Amazon Reviews. Entropy 2025, 27, 1202. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An Image Is Worth 16×16 Words: Transformers for Image Recognition at Scale. In Proceedings of the International Conference on Learning Representations; ICLR: Appleton, WI, USA, 2025. [Google Scholar]
- Zhao, S.; Wu, Y.; Tong, M.; Yao, Y.; Qian, W.; Qi, S. CoT-XNet: Contextual Transformer with Xception Network for Diabetic Retinopathy Grading. Phys. Med. Biol. 2022, 67, 245003. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Arik, S.Ö.; Pfister, T. TabNet: Attentive Interpretable Tabular Learning. In Proceedings of the AAAI Conference on Artificial Intelligence; PKP: Warsaw, Poland, 2021; Volume 35, pp. 6679–6687. [Google Scholar]
- Somepalli, G.; Goldblum, M.; Schwarzschild, A.; Bruss, C.B.; Goldstein, T. SAINT: Improved Neural Networks for Tabular Data via Row Attention and Contrastive Pre-Training. In Proceedings of the NeurIPS Workshop on Table Representation Learning; American Institute of Physics: College Park, MD, USA, 2022. [Google Scholar]
- Thielmann, A.; Reuter, A.; Säfken, B. Beyond Black-Box Predictions: Identifying Marginal Feature Effects in Tabular Transformer Networks. In Proceedings of the 29th International Conference on Artificial Intelligence and Statistics (AISTATS), Proceedings of Machine Learning Research; PMLR: New York, NY, USA, 2026; Volume 300. [Google Scholar]
- Gorishniy, Y.; Rubachev, I.; Khrulkov, V.; Babenko, A. Revisiting Deep Learning Models for Tabular Data. Adv. Neural Inf. Process. Syst. 2021, 34, 18932–18943. [Google Scholar]
- Huang, X.; Khetan, A.; Cvitkovic, M.; Karnin, Z. TabTransformer: Tabular Data Modeling Using Contextual Embeddings. arXiv 2020, arXiv:2012.06678. [Google Scholar]
- Aggarwal, C.C. Neural Networks and Deep Learning; Springer: Cham, Switzerland, 2018. [Google Scholar]
- Oberg, A.L.; Bot, B.M.; Grill, D.E.; Poland, G.A.; Therneau, T.M. Technical and Biological Variance Structure in mRNA-Seq Data: Life in the Real World. BMC Genom. 2012, 13, 304. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Cancer Genome Atlas Network. Comprehensive Molecular Portraits of Human Breast Tumours. Nature 2012, 490, 61–70. [CrossRef] [Scilit] [PubMed]
- Weinstein, J.N.; Collisson, E.A.; Mills, G.B.; Shaw, K.R.M.; Ozenberger, B.A.; Ellrott, K.; Shmulevich, I.; Sander, C.; Stuart, J.M. The Cancer Genome Atlas Pan-Cancer Analysis Project. Nat. Genet. 2013, 45, 1113–1120. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Geyer, F.C.; Rodrigues, D.N.; Weigelt, B.; Reis-Filho, J.S. Molecular Classification of Estrogen Receptor-Positive/Luminal Breast Cancers. Adv. Anat. Pathol. 2012, 19, 39–53. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Johnson, K.S.; Conant, E.F.; Soo, M.S. Molecular Subtypes of Breast Cancer: A Review for Breast Radiologists. J. Breast Imaging 2021, 3, 12–24. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Sheu, R.K.; Pardeshi, M.S. A Survey on Medical Explainable AI (XAI): Recent Progress, Explainability Approach, Human Interaction and Scoring System. Sensors 2022, 22, 8068. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhang, K.; Wang, D.; Lin, F.; Xie, J.; Zhou, W. A Comprehensive Review of Explainable Artificial Intelligence in Healthcare: Methods, Evaluation, and Clinical Integration. iScience 2026, 29, 115026. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Nedeljković, M.; Tanić, N.; Prvanović, M.; Milovanović, Z.; Tanić, N. Friend or Foe: ABCG2, ABCC1 and ABCB1 Expression in Triple-Negative Breast Cancer. Breast Cancer 2021, 28, 727–736. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Segovia-Mendoza, M.; Morales-Montor, J. Immune Tumor Microenvironment in Breast Cancer and the Participation of Estrogen and Its Receptors in Cancer Physiopathology. Front. Immunol. 2019, 10, 348. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Voglstaetter, M.; Thomsen, A.R.; Nouvel, J.; Koch, A.; Jank, P.; Navarro, E.G.; Gainey-Schleicher, T.; Khanduri, R.; Groß, A.; Rossner, F.; et al. Tspan8 Is Expressed in Breast Cancer and Regulates E-Cadherin/Catenin Signalling and Metastasis Accompanied by Increased Circulating Extracellular Vesicles. J. Pathol. 2019, 248, 421–437. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Gandhi, N.; Oturkar, C.C.; Das, G.M. Estrogen Receptor-Alpha and p53 Status as Regulators of AMPK and mTOR in Luminal Breast Cancer. Cancers 2021, 13, 3612. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hunt, E.N.; Kopacz, J.P.; Vestal, D.J. Unraveling the Role of Guanylate-Binding Proteins (GBPs) in Breast Cancer: A Comprehensive Literature Review and New Data on Prognosis in Breast Cancer Subtypes. Cancers 2022, 14, 2794. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Quintero, M.; Adamoski, D.; Reis, L.M.; Ascenção, C.F.R.; Oliveira, K.R.S.; Goncalves, K.A.; Dias, M.M.; Carazzolle, M.F.; Dias, S.M.G. Guanylate-Binding Protein-1 Is a Potential New Therapeutic Target for Triple-Negative Breast Cancer. BMC Cancer 2017, 17, 727. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Provance, O.K.; Lewis-Wambi, J. Deciphering the Role of Interferon Alpha Signaling and Microenvironment Crosstalk in Inflammatory Breast Cancer. Breast Cancer Res. 2019, 21, 59. [Google Scholar] [CrossRef] [Scilit] [PubMed]









| Scenario (p) | Best q | Best d | Best h | Parameters |
|---|---|---|---|---|
| 1000 | 4 | 64 | 4 | 289985 |
| 5000 | 8 | 16 | 1 | 642481 |
| Scenario | Model | Params | ROC-AUC | PR-AUC | Balanced Acc | Sensitivity | Specificity |
|---|---|---|---|---|---|---|---|
| Baseline MLP | 128257 | ||||||
| PLAT | 289985 | ||||||
| FT-Transformer | 266081 | ||||||
| Baseline MLP | 640257 | ||||||
| PLAT | 642481 | ||||||
| FT-Transformer | 649953 |
| Scenario | Model A | Model B | ROC-AUC (A−B) | p-Value |
|---|---|---|---|---|
| PLAT | Baseline MLP | 0.0239 | 0.2422 | |
| PLAT | FT-Transformer | 0.0041 | 0.8262 | |
| FT-Transformer | Baseline MLP | 0.0198 | 0.2812 | |
| PLAT | Baseline MLP | 0.6953 | ||
| PLAT | FT-Transformer | 0.3613 | ||
| FT-Transformer | Baseline MLP | 0.0145 | 0.5508 |
| Configuration | Parameters (Thousands) | % of Baseline |
|---|---|---|
| Baseline MLP | 128 | 100% |
| Very Lightweight Attention | 15 | 12% |
| Lightweight Attention | 30 | 23% |
| Semi-Lightweight Attention | 64 | 50% |
| Standard Attention | 133 | 104% |
| Metrics | PLAT | FT-Transformer | MLP |
|---|---|---|---|
| Accuracy | 0.933 ± 0.015 | 0.918 ± 0.010 | 0.947 ± 0.017 |
| Balanced Accuracy | 0.900 ± 0.028 | 0.874 ± 0.024 | 0.920 ± 0.033 |
| Sensitivity | 0.959 ± 0.010 | 0.955 ± 0.012 | 0.967 ± 0.009 |
| Specificity | 0.842 ± 0.056 | 0.793 ± 0.053 | 0.873 ± 0.063 |
| AUC | 0.949 ± 0.027 | 0.936 ± 0.021 | 0.948 ± 0.030 |
| PR-AUC | 0.980 ± 0.013 | 0.973 ± 0.014 | 0.977 ± 0.016 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Elatifi, K.; Gallego, N.J.; Sánchez-Pla, A.; Reverter, F. Self-Attention over Parallel Dense Embeddings for High-Dimensional Omic Data. Appl. Sci. 2026, 16, 7890. https://doi.org/10.3390/app16167890
Elatifi K, Gallego NJ, Sánchez-Pla A, Reverter F. Self-Attention over Parallel Dense Embeddings for High-Dimensional Omic Data. Applied Sciences. 2026; 16(16):7890. https://doi.org/10.3390/app16167890
Chicago/Turabian StyleElatifi, Kamal, Nicolas Jäger Gallego, Alex Sánchez-Pla, and Ferran Reverter. 2026. "Self-Attention over Parallel Dense Embeddings for High-Dimensional Omic Data" Applied Sciences 16, no. 16: 7890. https://doi.org/10.3390/app16167890
APA StyleElatifi, K., Gallego, N. J., Sánchez-Pla, A., & Reverter, F. (2026). Self-Attention over Parallel Dense Embeddings for High-Dimensional Omic Data. Applied Sciences, 16(16), 7890. https://doi.org/10.3390/app16167890

