Evaluation of Plate Homogeneity in Cell-Based Potency Assays Using Large Language Models
Abstract
1. Introduction
2. Materials and Methods
2.1. Experimental Setup for Plate Homogeneity Evaluation
2.2. Large Language Models for Plate Homogeneity Assessment
2.3. Prompt Design for LLM-Based Assessment
2.4. Conventional “Manual” Approach to Plate Homogeneity Assessment
3. Results
3.1. Overview of Analytical Outputs Across Models
3.2. Google Gemini Results
3.3. ChatGPT Results Overview
3.4. Microsoft Copilot Analyst Result Overview
3.5. Manual Plate Effect Analysis Overview
4. Discussion
5. Conclusions
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AI | Artificial Intelligence |
| ANOVA | Analysis of Variance |
| CI | Confidence Interval |
| CV | Coefficient of Variation |
| EC50 | Half-Maximal Effective Concentration |
| LISA | Local Indicators of Spatial Association |
| LLM | Large Language Model |
| LOESS | Locally Estimated Scatterplot Smoothing |
| QC | Quality Control |
| RGA | Reporter Gene Assay |
| Z-score | Standardized Score (Z-value) |
| AI | Artificial Intelligence |
| ANOVA | Analysis of Variance |
| CI | Confidence Interval |
| CV | Coefficient of Variation |
References
- Wouters, O.J.; McKee, M.; Luyten, J. Estimated Research and Development Investment Needed to Bring a New Medicine to Market, 2009–2018. JAMA-J. Am. Med. Assoc. 2020, 323, 844–853. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zheng, Y.; Koh, H.Y.; Ju, J.; Yang, M.; May, L.T.; Webb, G.I.; Li, L.; Pan, S.; Church, G. Large language models for drug discovery and development. Patterns 2025, 6, 101346. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Othman, Z.K.; Ahmed, M.M.; Okesanya, O.J.; Ibrahim, A.M.; Musa, S.S.; Hassan, B.A.; Saeed, L.I.; Lucero-Prisno, D.E. Advancing drug discovery and development through GPT models: A review on challenges, innovations and future prospects. Intell. Based Med. 2025, 11, 100233. [Google Scholar] [CrossRef] [Scilit]
- Kant, S.; Deepika; Roy, S. Artificial intelligence in drug discovery and development: Transforming challenges into opportunities. Discov. Pharm. Sci. 2025, 1, 7. [Google Scholar] [CrossRef] [Scilit]
- Lu, J.; Choi, K.; Eremeev, M.; Gobburu, J.; Goswami, S.; Liu, Q.; Mo, G.; Musante, C.J.; Shahin, M.H. Large Language Models and Their Applications in Drug Discovery and Development: A Primer. Clin. Transl. Sci. 2025, 18, e70205. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zheng, Y.; Koh, H.Y.; Yang, M.; Li, L.; May, L.T.; Webb, G.I.; Pan, S.; Church, G. Large Language Models in Drug Discovery and Development: From Disease Mechanisms to Clinical Trials. arXiv 2024, arXiv:2409.04481. [Google Scholar]
- Liao, Q.; Zhang, Y.; Chu, Y.; Ding, Y.; Liu, Z.; Zhao, X.; Wang, Y.; Wan, J.; Ding, Y.; Tiwari, P.; et al. Application of Artificial Intelligence in Drug-target Interactions Prediction: A Review. npj Biomed. Innov. 2025, 2, 1. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- White, J.R.; Abodeely, M.; Ahmed, S.; Debauve, G.; Johnson, E.; Meyer, D.M.; Mozier, N.M.; Naumer, M.; Pepe, A.; Qahwash, I.; et al. Best practices in bioassay development to support registration of biopharmaceuticals. Biotechniques 2019, 67, 126–137. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Singh, R.; Maheshwari, P. Evaluation of plate edge effects in in-vitro cell based assay. Int. J. Latest Trans. Eng. Sci. 2019, 7, 52–57. [Google Scholar]
- Lundholt, B.K.; Scudder, K.M.; Pagliaro, L. A simple technique for reducing edge effect in cell-based assays. J. Biomol. Screen 2003, 8, 566–570. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Nori, H.; King, N.; McKinney, S.M.; Carignan, D.; Horvitz, E. Capabilities of GPT-4 on Medical Challenge Problems. arXiv 2023, arXiv:2303.13375. [Google Scholar]
- Ji, Z.; Lee, N.; Frieske, R.; Yu, T.; Su, D.; Xu, Y.; Ishii, E.; Bang, Y.J.; Madotto, A.; Fung, P. Survey of Hallucination in Natural Language Generation. ACM Comput Surv. 2023, 55, 1–38. [Google Scholar] [CrossRef] [Scilit]
- Wei, J.; Bosma, M.; Zhao, V.Y.; Guu, K.; Yu, A.W.; Lester, B.; Du, N.; Dai, A.M.; Le, Q.V. Finetuned Language Models Are Zero-Shot Learners. arXiv 2022, arXiv:2109.01652. [Google Scholar]
- Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. Training language models to follow instructions with human feedback. Adv. Neural Inf. Process. Syst. 2022, 35, 27730–27744. [Google Scholar] [CrossRef] [Scilit]
- Tabassi, E. Artificial Intelligence Risk Management Framework (AI RMF 1.0); National Institute of Standards and Technology: Gaithersburg, MD, USA, 2023. [CrossRef] [Scilit]
- Maynez, J.; Narayan, S.; Bohnet, B.; McDonald, R. On Faithfulness and Factuality in Abstractive Summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics; Association for Computational Linguistics: Stroudsburg, PA, USA, 2020; pp. 1906–1919. [Google Scholar] [CrossRef] [Scilit]
- Hebenstreit, K.; Praas, R.; Kiesewetter, L.P.; Samwald, M. A comparison of chain-of-thought reasoning strategies across datasets and models. PeerJ Comput. Sci. 2024, 10, e1999. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chen, Y.; Benton, J.; Radhakrishnan, A.; Uesato, J.; Denison, C.; Schulman, J.; Somani, A.; Hase, P.; Wagner, M.; Roger, F.; et al. Reasoning Models Don’t Always Say What They Think. arXiv 2025, arXiv:2505.05410. [Google Scholar]
- Bivand, R.S.; Pebesma, E.; Gómez-Rubio, V. Applied Spatial Data Analysis with R; Springer: New York, NY, USA, 2013. [Google Scholar] [CrossRef] [Scilit]
- Lachmann, A.; Giorgi, F.M.; Alvarez, M.J.; Califano, A. Detection and removal of spatial bias in multiwell assays. Bioinformatics 2016, 32, 1959–1965. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Schmal, C.; Myung, J.; Herzel, H.; Bordyugov, G. Moran’s I quantifies spatio-temporal pattern formation in neural imaging data. Bioinformatics 2017, 33, 3072–3079. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- León, F.; Pizarro, E.; Noll, D.; Pertierra, L.R.; Parker, P.; Espinaze, M.P.A.; Luna-Jorquera, G.; Simeone, A.; Frere, E.; Dantas, G.P.M.; et al. Comparative genomics supports ecologically induced selection as a putative driver of banded penguin diversification. Mol. Biol. Evol. 2024, 41, msae166. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Dávid, C.; Giber, K.; Kerti-Szigeti, K.; Kollo, M.; Nusser, Z.; Acsády, L. An image segmentation method based on the spatial correlation coefficient of Local Moran’s I—Identification of A-type potassium channel clusters in the thalamus. eLife 2023, 12, RP89361. [Google Scholar] [CrossRef] [Scilit]
- Emons, M.; Gunz, S.; Crowell, H.L.; Mallona, I.; Kuehl, M.; Furrer, R.; Robinson, M.D. Harnessing the potential of spatial statistics for spatial omics data with pasta. Nucleic Acids Res. 2025, 53, gkaf870. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Malo, N.; Hanley, J.A.; Cerquozzi, S.; Pelletier, J.; Nadon, R. Statistical practice in high-throughput screening data analysis. Nat. Biotechnol. 2006, 24, 167–175. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Cleveland, W.S. Robust Locally Weighted Regression and Smoothing Scatterplots. J. Am. Stat. Assoc. 1979, 74, 829–836. [Google Scholar] [CrossRef]
- Hastie, T.; Tibshirani, R.; Friedman, J. The Elements of Statistical Learning; Springer: New York, NY, USA, 2009. [Google Scholar] [CrossRef]
- Smyth, G.K.; Speed, T. Normalization of cDNA microarray data. Methods 2003, 31, 265–273. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Fang, R.; Wey, A.; Bobbili, N.K.; Leke, R.F.G.; Taylor, D.W.; Chen, J.J. An analytical approach to reduce between-plate variation in multiplex assays that measure antibodies to Plasmodium falciparum antigens. Malar. J. 2017, 16, 287. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hong, M.-G.; Lee, W.; Nilsson, P.; Pawitan, Y.; Schwenk, J.M. Multidimensional Normalization to Minimize Plate Effects of Suspension Bead Array Data. J. Proteome Res. 2016, 15, 3473–3480. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wu, L.; Hall, T.; Ssewanyana, I.; Oulton, T.; Patterson, C.; Vasileva, H.; Singh, S.; Affara, M.; Mwesigwa, J.; Correa, S.; et al. Optimisation and standardisation of a multiplex immunoassay of diverse Plasmodium falciparum antigens to assess changes in malaria transmission using sero-epidemiology. Wellcome Open Res. 2020, 4, 26. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yang, H.; Zhao, Y.; Wu, Y.; Wang, S.; Zheng, T.; Zhang, H.; Ma, Z.; Che, W.; Wang, S.; Wei, S.; et al. Large language models meet text-centric multimodal sentiment analysis: A survey. Sci. China Inf. Sci. 2025, 68, 200101. [Google Scholar] [CrossRef] [Scilit]
- Herrera-Poyatos, D.; Peláez-González, C.; Zuheros, C.; Herrera-Poyatos, A.; Tejedor, V.; Herrera, F.; Montes, R. An overview of model uncertainty and variability in LLM-based sentiment analysis: Challenges, mitigation strategies, and the role of explainability. Front. Artif. Intell. 2025, 8, 1609097. [Google Scholar] [CrossRef] [Scilit] [PubMed]




| Type of LLM | Description | Examples | Key Advantages | Limitations |
|---|---|---|---|---|
| General-Purpose LLMs | Trained on diverse datasets; highly versatile for tasks like literature reviews, summarization, and technical document understanding. Suitable for broad applications across domains. | ChatGPT, Microsoft Copilot, Google Gemini, GPT-4, |
|
|
| Specialized LLMs | Designed for specific scientific purposes; trained on biomedical and pharmaceutical corpora. Optimized for domain-specific tasks such as functional genomics, protein interaction modeling, and drug-target discovery. | DNABERT, Nucleotide Transformer, HyenaDNA, Evo, GenSLMsTran ESM3, ESMFold, AlphaFold2, AlphaFold3, MolGPT, Molformer 1 |
|
|
| Requested Output | LLM | ||
|---|---|---|---|
| Google Gemini 1 | ChatGPT 2 | Microsoft Copilot Analyst | |
| ANOVA to detect row-wise effects. | Yes (tabular data with p-values) | Yes (tabular data with p-values) | Yes (description of significance only) |
| ANOVA to detect column-wise effects. | Yes (tabular data with p-values) | Yes (tabular data with p-values) | Yes (description of significance only) |
| Raw luminescence signal heatmap visualization | Yes | Yes | Yes |
| Average signal with 95% CI 3 | Yes | Yes | Yes |
| Z-score analysis | Yes | Yes | Yes |
| Spatial analysis | In the form of heat-maps not as Moran’s spatial analysis | Yes, Moran’s I, LOESS surface fit and two-way ANOVA (after additional input) | Yes, Moran’s I (after additional input) |
| Generated the averaged dataset | Yes | Yes | Yes |
| Generation of R-code | Yes | Yes | Yes |
| Generation of figures | Yes:
| Yes:
| Yes:
|
| Download links to figures | No direct links provided; figures displayed and downloadable from chat window. | Yes (after additional prompting) | No direct links provided; figures available after additional prompting |
| Proposal for additional analysis | No | Yes, detailed analysis was done with additional prompt input based on provided ChatGPT responses. | Yes, detailed analysis was done with additional prompt input based on provided Analyst responses. |
| Metric (Averaged Plate) | Manual (Excel M365) | Google Gemini | ChatGPT | Microsoft Copilot Analyst |
|---|---|---|---|---|
| Row effect (one-way ANOVA p) | Significant p = 4.31 × 10−37 | Significant p < 0.000001 | Significant p ≈ 4.3 × 10−37 | Significant p < 10−22 |
| Column effect (one-way ANOVA p) | Not significant p = 0.765 | Not significant p = 0.765 | Not significant p = 0.765 | Not significant p ≫ 0.05 |
| Row A mean signal | 37,099.4 (95% CI: 36,338.63–37,860.16) | 37,099.4 (95% CI: 36,338.63–37,860.16) | 37,099.40 | 37,099.40 (95% CI: 36,338.63–37,860.16) |
| Row H mean signal | 29,864.25 (95% CI: 29,346.80–30,381.70) | 29,864.25 (95% CI: 29,346.80–30,381.70) | 29,864.25 | 29,864.25 (95% CI: 29,346.80–30,381.70) |
| Column 1 mean signal | 33,928.47 (95% CI: 32,166.13–35,690.81) | 33,928.47 (95% CI: 32,166.13–35,690.81) | (Not explicitly stated) | 33,928.47 (95% CI: 32,166.13–35,690.81) |
| Column 12 mean signal | 32,189.78 (95% CI: 29,988.53–34,391.03) | 32,189.78 (95% CI: 29,988.53–34,391.03) | (Not explicitly stated) | 32,189.78 (95% CI: 29,988.53–34,391.03) |
| Z-score outliers | No wells with |Z| > 3.0 | No wells with |Z| > 3.0 | No isolated anomalies by Z-score analysis. Z-score heatmaps confirm absence of extreme outliers. | No isolated anomalies by Z-score analysis. Z-score heatmaps confirm absence of extreme outliers |
| Spatial autocorrelation (Moran’s I/p) | NA | NA | Moran’s I significant; permutation test p = 0.001 | Moran’s I = 0.82 |
| Primary conclusion | Strong row-wise gradient/edge effect, minimal column effect | Same | Same | Same |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Kosir, R.; Uzar, A.; Brvar, L.; Oven, I. Evaluation of Plate Homogeneity in Cell-Based Potency Assays Using Large Language Models. Biophysica 2026, 6, 66. https://doi.org/10.3390/biophysica6040066
Kosir R, Uzar A, Brvar L, Oven I. Evaluation of Plate Homogeneity in Cell-Based Potency Assays Using Large Language Models. Biophysica. 2026; 6(4):66. https://doi.org/10.3390/biophysica6040066
Chicago/Turabian StyleKosir, Rok, Aleksandra Uzar, Luka Brvar, and Irena Oven. 2026. "Evaluation of Plate Homogeneity in Cell-Based Potency Assays Using Large Language Models" Biophysica 6, no. 4: 66. https://doi.org/10.3390/biophysica6040066
APA StyleKosir, R., Uzar, A., Brvar, L., & Oven, I. (2026). Evaluation of Plate Homogeneity in Cell-Based Potency Assays Using Large Language Models. Biophysica, 6(4), 66. https://doi.org/10.3390/biophysica6040066

