The Use of ChatGPT to Write Assessment Questions
Abstract
1. Introduction
2. Methods
2.1. The Educational Setting
2.2. Question Generation and Administration
2.3. Survey Design and Administration
2.4. Bloom’s Taxonomy Levels
2.5. Statistical Analysis
3. Results
3.1. Student Demographics
3.2. Student Academic Performance
3.3. Student Perceptions
3.4. Bloom’s Taxonomy Levels
4. Discussion
4.1. The Context of the Study
4.2. Student Academic Performance
4.3. Student Perceptions
4.4. Bloom’s Taxonomy Levels
4.5. Study Limitations
5. Conclusions
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AI | Artificial intelligence |
| SD | Standard deviation |
References
- Coşkun, Ö.; Kıyak, Y.S.; Budakoğlu, I.İ. ChatGPT to generate clinical vignettes for teaching and multiple-choice questions for assessment: A randomized controlled experiment. Med. Teach. 2025, 47, 268–274. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Fuller, K.A.; Morbitzer, K.A.; Zeeman, J.M.; Persky, A.M.; Savage, A.C.; McLaughlin, J.E. Exploring the use of ChatGPT to analyze student course evaluation comments. BMC Med. Educ. 2024, 24, 423. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ozturk, N.; Yakak, I.; Ağ, M.B.; Aksoy, N. Is ChatGPT reliable and accurate in answering pharmacotherapy-related inquiries in both Turkish and English? Curr. Pharm. Teach. Learn. 2024, 16, 102101. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Knobloch, J.; Cozart, K.; Halford, Z.; Hilaire, M.; Richter, L.M.; Arnoldi, J. Students’ perception of the use of artificial intelligence (AI) in pharmacy school. Curr. Pharm. Teach. Learn. 2024, 16, 102181. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Alrazeeni, D.M.; Alharrasi, M.; Rony, M.K.K.; Biswas, R.K.; Tama, I.J.; Halder, C.R.; Deb, B.; Bashar, F.; Akter, F. Transforming nursing education with artificial intelligence: A systematic review (2010–2025). SAGE Open Nurs. 2026, 12, 23779608261424597. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Khakpaki, A. Advancements in artificial intelligence transforming medical education: A comprehensive overview. Med. Educ. Online 2025, 30, 2542807. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Schönwetter, D.J.; MacDonald, L.L.; Reynolds, P.A.; Eaton, K.A. Bridging borders, bridging barriers: Artificial intelligence for dental education. Eur. J. Dent. Educ. 2026, 30, 769–775. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Abdel Aziz, M.H.; Rowe, C.; Southwood, R.; Nogid, A.; Berman, S.; Gustafson, K. A scoping review of artificial intelligence within pharmacy education. Am. J. Pharm. Educ. 2024, 88, 100615. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ali, M. Will AI reshape or deform pharmacy education? Curr. Pharm. Teach. Learn. 2025, 17, 102274. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Cain, J.; Malcom, D.R.; Aungst, T.D. The role of artificial intelligence in the future of pharmacy education. Am. J. Pharm. Educ. 2023, 87, 100135. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Cheung, B.H.H.; Lau, G.K.K.; Wong, G.T.C.; Lee, E.Y.P.; Kulkarni, D.; Seow, C.S.; Wong, R.; Co, M.T. ChatGPT versus human in generating medical graduate exam multiple choice questions—A multinational prospective study (Hong Kong S.A.R., Singapore, Ireland, and the United Kingdom). PLoS ONE 2023, 18, e0290691. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Edwards, C.J.; Erstad, B.L. Evaluation of a generative language model tool for writing examination questions. Am. J. Pharm. Educ. 2024, 88, 100684. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Khan, H.F.; Qayyum, S.; Beenish, H.; Khan, R.A.; Iltaf, S.; Faysal, L.R. Determining the alignment of assessment items with curriculum goals through document analysis by addressing identified item flaws. BMC Med. Educ. 2025, 25, 200. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Nasution, N.E.A. Using artificial intelligence to create biology multiple choice questions for higher education. Agric. Environ. Educ. 2023, 2, em002. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Han, Z.; Battaglia, F.; Udaiyar, A.; Fooks, A.; Terlecky, S.R. An explorative assessment of ChatGPT as an aid in medical education: Use it with caution. Med. Teach. 2024, 46, 657–664. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Laupichler, M.C.; Rother, J.F.; Grunwald Kadow, I.C.; Ahmadi, S.; Raupach, T. Large language models in medical education: Comparing ChatGPT- to human-generated exam questions. Acad. Med. 2024, 99, 508–512. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zuckerman, M.; Flood, R.; Tan, R.J.B.; Kelp, N.; Ecker, D.J.; Menke, J.; Lockspeiser, T. ChatGPT for assessment writing. Med. Teach. 2023, 45, 1224–1227. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Joncas, S.X.; St-Onge, C.; Bourque, S.; Farand, P. Re-using questions in classroom-based assessment: An exploratory study at the undergraduate medical education level. Perspect. Med. Educ. 2018, 7, 373–378. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Law, A.K.K.; So, J.; Lui, C.T.; Choi, Y.F.; Cheung, K.H.; Kei-Ching Hung, K.; Graham, C.A. AI versus human-generated multiple-choice questions for medical education: A cohort study in a high-stakes examination. BMC Med. Educ. 2025, 25, 208. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Elzayyat, M.; Mohammad, J.N.; Zaqout, S. Assessing LLM-generated vs. expert-created clinical anatomy MCQs: A student perception-based comparative study in medical education. Med. Educ. Online 2025, 30, 2554678. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kiyani, A.; Hanif, F.; Muhammad, M.; Iqbal, S.; Zaib, N.; Bashir, U.; Ali, K. Benchmarking ChatGPT-generated multiple-choice questions against faculty-authored items in dental education. Sci. Rep. 2025, 15, 44805. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Cain, J.; Rajan, A.S. Proof of concept of ChatGPT as a virtual tutor. Am. J. Pharm. Educ. 2024, 88, 101333. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- McLaughlin, J.E.; Ponte, C.D.; Lyons, K. Student perceptions of GenAI as a virtual tutor to support collaborative research training for health professionals. BMC Med. Educ. 2025, 25, 895. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chauhan, A.; Khaliq, F.; Nayak, K.R. Title: Assessing quality of scenario-based multiple-choice questions in physiology: Faculty-generated vs. ChatGPT-generated questions among Phase I medical students. Int. J. Artif. Intell. Educ. 2025, 35, 2315–2344. [Google Scholar] [CrossRef] [Scilit]
- Edwards, C.; Erstad, B.; Cornelison, B. Evaluating NAPLEX preparation exam questions generated by a large language model. Am. J. Pharm. Educ. 2026, 90, 101983. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kıyak, Y.S.; Emekli, E.; Coşkun, Ö.; Budakoğlu, I. Keeping humans in the loop efficiently by generating question templates instead of questions using AI: Validity evidence on hybrid AIG. Med. Teach. 2025, 47, 744–747. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ngo, A.; Gupta, S.; Perrine, O.; Reddy, R.; Ershadi, S.; Remick, D. ChatGPT 3.5 fails to write appropriate multiple choice practice exam questions. Acad. Pathol. 2024, 11, 100099. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Klang, E.; Portugez, S.; Gross, R.; Kassif Lerner, R.; Brenner, A.; Gilboa, M.; Ortal, T.; Ron, S.; Robinzon, V.; Meiri, H.; et al. Advantages and pitfalls in utilizing artificial intelligence for crafting medical examinations: A medical education pilot study with GPT-4. BMC Med. Educ. 2023, 23, 772. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Balu, A.; Prvulovic, S.T.; Fernandez Perez, C.; Kim, A.; Donoho, D.A.; Keating, G. Evaluating the value of AI-generated questions for USMLE step 1 preparation: A study using ChatGPT-3.5. Med. Teach. 2025, 47, 1645–1653. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Sridharan, K.; Sequeira, R.P. Artificial intelligence and medical education: Application in classroom instruction and student assessment using a pharmacology & therapeutics case study. BMC Med. Educ. 2024, 24, 431. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Emekli, E.; Karahan, B.N. Artificial intelligence in radiology examinations: A psychometric comparison of question generation methods. Diagn. Interv. Radiol. 2025, 32, 548–554. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Shultz, B.; DiDomenico, R.J.; Goliak, K.; Mucksavage, J. Exploratory assessment of GPT-4’s effectiveness in generating valid exam items in pharmacy education. Am. J. Pharm. Educ. 2025, 89, 101405. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kıyak, Y.S.; Coşkun, Ö.; Budakoğlu, I.İ.; Uluoğlu, C. ChatGPT for generating multiple-choice questions: Evidence on the use of artificial intelligence in automatic item generation for a rational pharmacotherapy exam. Eur. J. Clin. Pharmacol. 2024, 80, 729–735. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Abozaid, H.; Park, Y.S.; Tekian, A. Peer review improves psychometric characteristics of multiple choice questions. Med. Teach. 2017, 39, S50–S54. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ahmed, A.; Kerr, E.; O’Malley, A. Quality assurance and validity of AI-generated single best answer questions. BMC Med. Educ. 2025, 25, 300. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Klausner, E.A. In-class exercises regarding the roles of excipients in a pharmaceutics course. Innov. Pharm. 2023, 14, 3. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Klausner, E.A.; Nagel, K. Aulton’s Pharmaceutics: The Design and Manufacture of Medicines, Sixth Edition, Kevin M.G. Taylor, Michael E. Aulton (Eds.), Elsevier (2021), 968 pp, US $67.99 (paperback), ISBN: 978-0-7020-8154-5. Curr. Pharm. Teach. Learn. 2022, 14, 809–810. [Google Scholar] [CrossRef] [Scilit]
- Medina, M.S.; Farland, M.Z.; Conry, J.M.; Culhane, N.; Kennedy, D.R.; Lockman, K.; Malcom, D.R.; Mirzaian, E.; Vyas, D.; Steinkopf, M.; et al. The AACP Academic Affairs Committee’s guidance for use of the Curricular Outcomes and Entrustable Professional Activities (COEPA) for pharmacy graduates. Am. J. Pharm. Educ. 2023, 87, 100562. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Khan, M.O.F.; Rashrash, M.; Drouin, A.; Huynh, T. Evaluating curriculum differences in US PharmD programs: A peer evaluation. Am. J. Pharm. Educ. 2024, 88, 100712. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Jasti, B.R.; Fincher, T.K.; Mobley, W.C.; Holladay, J.W.; Klausner, E.A.; Nagel, K.; Bhalla, S.; Dutta, A.K.; Brazeau, G.A.; Vadlapatla, R. A Didactic Curriculum Toolkit for Pharmaceutics and Related Disciplines: Recommendations to ACPE Accredited Programs. Am. J. Pharm. Educ. 2026, 90, 101934. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hall, E.A.; Spivey, C.; Kendrex, H.; Havrda, D.E. Effects of remote proctoring on composite examination performance among doctor of pharmacy students. Am. J. Pharm. Educ. 2021, 85, 8410. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Klausner, E.A.; Persky, A.M. An integrative review of approaches used to assess course interventions. Am. J. Pharm. Educ. 2023, 87, ajpe8896. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Woodruff, A.E.; Jensen, M.; Loeffler, W.; Avery, L. Advanced screencasting with embedded assessments in pathophysiology and therapeutics course modules. Am. J. Pharm. Educ. 2014, 78, 128. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Competencygenie: The AI-Powered Tool for Tagging Exam Items, Classifying Competencies, and Enhancing Curriculum Evaluation in Health Professions Programs. Available online: https://enflux.com/competencygenie/ (accessed on 9 September 2026).
- Larsen, T.M.; Endo, B.H.; Yee, A.T.; Do, T.; Lo, S.M. Probing internal assumptions of the revised Bloom’s taxonomy. CBE—Life Sci. Educ. 2022, 21, ar66. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Krathwohl, D.R. A revision of Bloom’s taxonomy: An overview. Theory Pract. 2002, 41, 212–218. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lakens, D. Equivalence tests: A practical primer for t tests, correlations, and meta-analyses. Soc. Psychol. Personal. Sci. 2017, 8, 355–362. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Boscardin, C.K.; Sewell, J.L.; Tolsgaard, M.G.; Pusic, M.V. How to use and report on p-values. Perspect. Med. Educ. 2024, 13, 250–254. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Alexander, K.M.; Johnson, M.; Farland, M.Z.; Blue, A.; Bald, E.K. Exploring generative artificial intelligence to enhance reflective writing in pharmacy education. Am. J. Pharm. Educ. 2025, 89, 101416. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hernandez, T.; Magid, M.S.; Polydorides, A.D. Assessment question characteristics predict medical student performance in general pathology. Arch. Pathol. Lab. Med. 2021, 145, 1280–1288. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kim, M.K.; Patel, R.A.; Uchizono, J.A.; Beck, L. Incorporation of Bloom’s taxonomy into multiple-choice examination questions for a pharmacotherapeutics course. Am. J. Pharm. Educ. 2012, 76, 114. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Klender, S.; Ferriby, A.; Notebaert, A. Differences in item statistics between positively and negatively worded stems on histology examinations. HAPS Educ. 2019, 23, 476–486. [Google Scholar] [CrossRef] [Scilit]
- Ray, M.E.; Rudolph, M.J.; Daugherty, K.K. Bloom’s taxonomy in health professions education: Associations with exam scores, clinical reasoning, and instructional effectiveness. Curr. Pharm. Teach. Learn. 2025, 17, 102444. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Indran, I.R.; Paranthaman, P.; Gupta, N.; Mustafa, N. Twelve tips to leverage AI for efficient and effective medical question generation: A guide for educators using Chat GPT. Med. Teach. 2024, 46, 1021–1026. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kıyak, Y.S.; Emekli, E. ChatGPT prompts for generating multiple-choice questions in medical education and evidence on their validity: A literature review. Postgrad. Med. J. 2024, 100, 858–865. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Edwards, C.J.; Cornelison, B.; Erstad, B.L. Comparison of a generative large language model to pharmacy student performance on therapeutics examinations. Curr. Pharm. Teach. Learn. 2025, 17, 102394. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ehlert, A.; Ehlert, B.; Cao, B.; Morbitzer, K. Large language models and the North American Pharmacist Licensure Examination (NAPLEX) practice questions. Am. J. Pharm. Educ. 2024, 88, 101294. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Khatri, S.; Sengul, A.; Moon, J.; Jackevicius, C.A. Accuracy and reproducibility of ChatGPT responses to real-world drug information questions. J. Am. Coll. Clin. Pharm. 2025, 8, 432–438. [Google Scholar] [CrossRef] [Scilit]
- Caetano, M.L.; Pawasauskas, J. A retrospective analysis of the impact of disabling item review on item performance on computerized fixed-item tests in a doctor of pharmacy program. Curr. Pharm. Teach. Learn. 2020, 12, 539–543. [Google Scholar] [CrossRef] [Scilit] [PubMed]

| 2024 Cohort, % | 2025 Cohort, % | Combined Classes, % | ||
|---|---|---|---|---|
| Response rate | 90 | 96 | 93 | |
| Age, years | 18–24 | 49 | 38 | 43 |
| 25–30 | 37 | 40 | 38 | |
| ≥31 | 14 | 23 | 19 | |
| Prefer not to answer | 0 | 0 | 0 | |
| Gender | Female | 77 | 69 | 73 |
| Male | 23 | 31 | 27 | |
| Non-binary | 0 | 0 | 0 | |
| Transgender | 0 | 0 | 0 | |
| Prefer not to answer | 0 | 0 | 0 | |
| Highest earned academic degree | No degree | 14 | 6 | 9 |
| Associate | 28 | 23 | 25 | |
| Bachelor’s | 49 | 65 | 57 | |
| Master’s | 9 | 6 | 7 | |
| PhD | 0 | 0 | 0 |
| 2024 Cohort, n = 41 a | 2025 Cohort, n = 48 | ||||
|---|---|---|---|---|---|
| Course | Question Source | Number of Questions | Percent Correct, Mean ± SD | Number of Questions | Percent Correct, Mean ± SD |
| Pharmaceutics I | ChatGPT | 42 | 78 ± 8 b,c | 47 | 76 ± 11 b |
| Instructor | 74 | 75 ± 11 d | 77 | 73 ± 11 d | |
| ChatGPT and instructor | 116 | 76 ± 10 | 124 | 74 ± 11 | |
| Pharmaceutics II | ChatGPT | 55 | 73 ± 10 b | 56 | 74 ± 13 b |
| Instructor | 82 | 79 ± 11 | 84 | 79 ± 11 | |
| ChatGPT and instructor | 137 | 76 ± 10 | 140 | 77 ± 11 | |
| Combined courses | ChatGPT | 97 | 75 ± 9 b | 103 | 75 ± 11 |
| Instructor | 156 | 77 ± 11 | 161 | 76 ± 11 | |
| ChatGPT and the instructor | 253 | 76 ± 10 | 264 | 76 ± 10 | |
| Course | Question Source | Bloom’s Level b | 2024, Number of Questions, % | 2025, Number of Questions, % |
|---|---|---|---|---|
| Pharmaceutics I | ChatGPT c | Remember | 25, 60% | 27, 57% |
| Understand | 17, 40% | 18, 38% | ||
| Apply | 0, 0% | 1, 2% | ||
| Analyze | 0, 0% | 1, 2% | ||
| Instructor | Remember | 24, 32% | 26, 34% | |
| Understand | 40, 54% | 46, 60% | ||
| Apply | 0, 0% | 1, 1% | ||
| Analyze | 10, 14% | 4, 5% | ||
| Pharmaceutics II | ChatGPT | Remember | 34, 62% | 32, 57% |
| Understand | 19, 35% | 22, 39% | ||
| Apply | 0, 0% | 0, 0% | ||
| Analyze | 2, 4% | 2, 4% | ||
| Instructor | Remember | 40, 49% | 43, 51% | |
| Understand | 33, 40% | 31, 37% | ||
| Apply | 4, 5% | 4, 5% | ||
| Analyze | 5, 6% | 6, 7% |
| Course | Question Source | Bloom’s Level b | 2024, Percent Correct (Mean ± SD) | 2025, Percent Correct (Mean ± SD) |
|---|---|---|---|---|
| Pharmaceutics I | ChatGPT c | Remember | 76 ± 11 | 75 ± 13 |
| Understand | 81 ± 10 | 78 ± 12 | ||
| Apply | N/A c | N/A c | ||
| Analyze | N/A c | N/A c | ||
| Instructor | Remember | 72 ± 13 | 74 ± 13 | |
| Understand | 77 ± 13 | 72 ± 13 | ||
| Apply | N/A c | N/A c | ||
| Analyze | 77 ± 13 | 68 ± 21 | ||
| Pharmaceutics II | ChatGPT | Remember | 73 ± 10 | 73 ± 12 |
| Understand | 73 ± 13 | 75 ± 15 | ||
| Apply | N/A c | N/A c | ||
| Analyze | N/A c | N/A c | ||
| Instructor | Remember | 81 ± 11 d | 78 ± 12 | |
| Understand | 76 ± 14 | 79 ±12 | ||
| Apply | 87 ± 17 | 79 ± 21 | ||
| Analyze | 75 ± 21 | 75 ± 19 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Klausner, E.A.; Mark, K.S. The Use of ChatGPT to Write Assessment Questions. Pharmacy 2026, 14, 142. https://doi.org/10.3390/pharmacy14070142
Klausner EA, Mark KS. The Use of ChatGPT to Write Assessment Questions. Pharmacy. 2026; 14(7):142. https://doi.org/10.3390/pharmacy14070142
Chicago/Turabian StyleKlausner, Eytan A., and Karen S. Mark. 2026. "The Use of ChatGPT to Write Assessment Questions" Pharmacy 14, no. 7: 142. https://doi.org/10.3390/pharmacy14070142
APA StyleKlausner, E. A., & Mark, K. S. (2026). The Use of ChatGPT to Write Assessment Questions. Pharmacy, 14(7), 142. https://doi.org/10.3390/pharmacy14070142

