1. Introduction
The public release of ChatGPT (OpenAI, CA, USA) in November 2022 was met with both enthusiasm and apprehension from the medical education community [
1]. Within weeks, it was recognized that ChatGPT, a large language model (LLM) representing a disruptive breakthrough in artificial intelligence (AI) technology, could pass standardized licensing examinations for medical students [
2,
3]. Since that time, further development of ChatGPT has improved its capabilities, other competing LLMs have been launched, and medical educators have been rapidly seeking to understand how to appropriately incorporate AI-augmented education methods into their practice [
4,
5,
6,
7].
With regard to medical student assessment and examination, the literature on the usage of LLMs focuses almost exclusively on multiple-choice format questions [
8,
9,
10,
11,
12,
13,
14,
15,
16]. For assessment of preclinical medical student learning at McMaster University (Hamilton, ON, Canada), the free-text short-answer question format is relied upon, which is not as well-represented in the LLM literature. Our group has previously studied ChatGPT-generated responses to free-text assessment problems compared to student-generated responses, focusing on the risks posed by ChatGPT to assessment validity, and we have established that ChatGPT-assisted grading of student-generated responses may be a viable assistant to instructors grading student work [
15,
16].
In this study, we turn our attention to the generation of new free-text assessment problems. Producing realistic clinical vignettes and questions of appropriate difficulty and pedagogical value, with corresponding answer keys and rubrics for instructors, is a time-consuming and challenging task. This study evaluates the quality and efficiency of AI-generated short-answer questions and their corresponding answer keys compared to human-developed resources.
2. Materials and Methods
2.1. Context
One of the assessment tools used for preclinical medical learners at the Michael G. DeGroote School of Medicine is called the Concept Application Exercise (CAE). This format has been described elsewhere [
17,
18,
19]. In brief, a CAE includes between 4 and 6 clinical vignettes, each with corresponding questions. The medical student responds in a free-text format to each vignette, answering the questions and explaining their answers and thought processes, completing the CAE in 60 min. In total, medical students write approximately 12–14 CAEs throughout their preclinical studies.
Each CAE is graded by a tutor, who is a physician–educator working with the learners in small-group problem-based learning (PBL) sessions between each CAE. CAE problems are developed centrally by the education leaders for each curriculum sub-unit, and the same problems are administered to all medical learners. Answer keys are also generated by the curriculum leads for the tutors to assign a score of “Novice”, “Proficient”, or “Accomplished” to each of the CAE problems. Tutors additionally provide written feedback to each individual free-text response.
2.2. Generation of CAE Problems
In this study, we compare CAE problems and answer keys generated by ChatGPT model version 4o to those produced by human instructors. We focus on renal and hematology problems, which are taught as part of the same preclinical curriculum block in our program.
A total of 9 representative human-generated CAE problems and answer keys were selected from a bank of previous assessments (5 renal, 4 hematology). A further 12 problems and answer keys (6 renal, 6 hematology) were generated by using ChatGPT, for a total of 21 problem and answer key pairs (11 renal, 10 hematology) for use in this study.
To generate problems and answer keys, a prompt template was developed using best assessment practices and numerous process iterations, as shown in
Figure 1. Detailed discussions of prompt engineering in this context are available elsewhere [
10,
15,
20]. For each problem, the same prompt was used, changing only a placeholder reserved for the primary testable concept intended for the CAE problem. In this study, a list of 12 testable concepts was created by a group of subject matter experts, and each testable concept was used one time to produce one of the ChatGPT-generated problem and answer key pairs. For each concept, ChatGPT was asked to generate a single problem and answer key; multiple versions were not requested and no selection among alternatives was performed. Problems were generated one at a time, and the prompt was not reiterated or modified after submission. For each problem, a new ChatGPT session was initiated, and prior conversation history was cleared to minimize potential bias across prompts. All ChatGPT-generated problems were produced in February 2025. The 12 ChatGPT-generated problems and answer keys were reviewed by the renal and hematology subunit education leaders to ensure feasibility for inclusion in a CAE, and minor adjustments (e.g., grammar, text format, phrasing) were made to ensure a high-fidelity comparison between human-generated and ChatGPT-generated groups. These education leaders were not involved in other aspects of this study.
2.3. Assessment of CAE Problem Quality
Four experienced faculty reviewers and one student assessment representative, blinded to question origin, independently reviewed each of the 21 CAE problem and answer key pairs. Each reviewer was asked to assign a single integer score from 1 (very poor quality) to 5 (excellent quality) based on predefined criteria for high-quality CAE questions that are distributed to current faculty instructors who produce problems. These criteria include medical accuracy, clarity, clinical relevance, cognitive demand, difficulty level, feasibility to complete in 10 min, curricular alignment, and clarity of the produced answer key. These scores constituted the primary endpoint for analysis in this study. Reviewers had the option to provide comments in addition to their numeric score for each question. Reviewers were free to score the problems in any order, without a fixed time constraint, and were able to adjust scores to previous problems or edit their comments as they progressed through the collection of 21 problems.
2.4. Data Analysis
The average of the five numerical scores for each CAE question and answer key pair were calculated, in addition to basic descriptive statistics for the subsets of ChatGPT-generated and human-generated CAE problems. The mean scores for each subset were compared using a two-tailed Student’s t-test (α = 0.05). Further comparisons were carried out with the hematology and renal topic subgroups. One-way ANOVA testing was used to compare the scores assigned by the five reviewers. To account for inter-rater and inter-item random effects, an ordinal mixed-effects model was additionally used to compare groups, using the “ordinal” package in R, version 4.5.2.
For the optional comments, common elements were extracted for thematic analysis. Additionally, each comment was manually assigned a sentiment of negative, neutral or mixed, or positive, corresponding to a value of −1, 0, or +1, which was used to provide aggregate sentiment statistics for each reviewer and for each CAE problem. Basic descriptive statistics were applied to compare the content and sentiment of faculty comments, both as a whole and between the ChatGPT-generated and human-generated responses.
2.5. Ethics
The study was approved by the local institutional ethics review board.
3. Results
Twelve ChatGPT-generated problems and answer keys were successfully produced, with an example problem shown in
Figure 2. The average and standard deviation of the five scores applied by faculty reviewers for each of the 21 CAE problems is shown in
Figure 3.
Expert reviewers rated ChatGPT-generated pairs of questions and answer keys significantly higher in quality than human-generated content, with a mean rating of 4.00 compared to 2.71 (
p < 0.001). This statistically significant difference was preserved when the analysis was stratified by medical subunit (i.e., renal or hematology). In addition, the average scores applied to ChatGPT-generated problems were more consistent across the five faculty reviewers, with a range of 1.0 for this subset compared to a range of 2.0 for the human-generated problems. Further details are given in
Table 1. In ordinal mixed-effect modeling, accounting for rater and item random effects and treating score as an ordinal rather than continuous variable, human-generated item pairs had significantly lower odds of receiving higher ratings than ChatGPT-generated item pairs (β = −2.43, 95% confidence interval: −3.34 to −1.51,
p < 0.001), corresponding to an odds ratio of 0.089 (95% confidence interval: 0.035 to 0.221) for human-generated problems receiving higher scores than ChatGPT-generated problems.
One-way ANOVA of the scores assigned by each reviewer did not suggest a statistically significant difference in the mean scores assigned by any reviewers (
p = 0.728), although there were subtle differences appreciated in reviewer-specific patterns. In particular, only one reviewer assigned a score of 1 to any problem, while another reviewer also did not assign any scores of 5. Score distributions assigned by each reviewer are shown in
Figure 4.
In total, with 5 scores assigned to each of the 21 problem and answer key pairs, there were 105 scores collected in this study. A total of 76 of these scores were accompanied by optional comments (72.4%).
Table 2 characterizes the distribution and sentiment of these comments. Two reviewers provided comments on all the 21 problems, resulting in all individual problems having at least two comments for analysis. Overall, outright positive comments constituted the minority of the comments (22/76, 28.9%), with only one of these comments being attributed to a human-generated problem.
When converting comments to a sentiment score of +1, 0, or −1 and taking the sum for each problem and answer key pair, the lowest total score was −5, reflecting human-generated Q16 which was criticized for having an ambiguous answer key, complex question wording, and too many testable concepts for a 10 min completion window. The highest total score was +3, achieved by a ChatGPT-generated problem (Q11, shown in
Figure 2), representing three positive comments and two mixed or neutral comments. The average total sentiment score for ChatGPT-generated problems was +1.0 (range: −2.0 to +3.0), compared to −2.4 for human-generated problems (range: −5.0 to +1.0).
The faculty comments can be roughly grouped into six themes. The first three pertain to the medical content or difficulty of the CAE problem. (a) In terms of medical accuracy, only two problems received negative comments: ChatGPT-generated Q1, where the vignette for a patient with Crohn’s disease taking no medications was not felt to be in keeping with their symptom presentation timeline, which drew three of the nine outright negative comments towards ChatGPT-generated problems across the dataset; and human-generated Q3, where one reviewer felt the problem did not realistically conform with Canadian resource stewardship guidelines (i.e., availability of MR imaging). (b) Comments on problem difficulty and timing were often heterogeneous with slightly differing views between reviewers, with the exception of the two problems described in the previous paragraph (Q11, Q16). (c) A total of 11 of the problems received comments with regard to the level of Bloom’s taxonomy reasoning involved in the problem, 8 of which included some concern that the problem focused too heavily on recall of memorized facts rather than higher-order domains such as application or integration (3/12 ChatGPT, 5/9 human).
The remaining three themes focused on writing style and presentation. (d) The integration of the clinical vignette into the CAE prompt arose in comments for six problems; in five of these, it was felt that the vignette was generally unrelated to the problem or not necessary to respond to the prompt (four human-generated problems, one ChatGPT-generated responses). (e) Of the nine problems where grammar or writing style drew attention, eight were in a negative light; two ChatGPT-generated responses were noted to be somewhat vague or containing extraneous information, while six human-generated responses were noted for ambiguity, questions having multiple parts or disconnected concepts, confusing writing, or in one instance, a spelling mistake. (f) Finally, seven problems specifically had comments on their answer keys. Five of these were negative comments for human-generated problems, noting the difficulty of assigning a single score or applying the grading rubric consistently to student work.
4. Discussion
Our study sought to compare the quality of free-text assessment problems for preclinical medical students produced by either ChatGPT or McMaster University faculty members. The primary result of our experiment indicates that free-text problems and answer keys generated by ChatGPT were deemed by our blinded panel of faculty reviewers to be of overall favorable quality compared to similar problems generated by human curriculum leads, despite some observed deficiencies. Analysis of optional comments was useful to identify areas of strength or weakness for the two methods of producing free-text problems. More negative comments were directed towards human-generated problems, consistent with the lower average score of these problems.
There are several factors and limitations that may have exacerbated this scoring difference between ChatGPT-generated and human-generated problems. For example, this experiment asked reviewers to consider the quality of the answer key in their scoring of the overall problem and answer key pair. In reality, our CAEs are graded by tutors who work with the students in small group settings in the weeks leading up to the assessments, and who are in regular contact with the curriculum leaders to discuss nuances in grading that may not be explicitly written down in human-generated answer keys.
Seven of our problems drew negative comments for answer keys, with the most severe five of these negative comments being directed towards human-generated answer keys. From the student perspective, the quality of writing in the answer key does not factor as strongly into the educational experience as the question itself, and adjustments to answer keys may have improved the quality scores of these problems. Moreover, while some human-generated problems were deemed vague by our reviewer panel, CAE problems are often influenced by learning materials (e.g., lectures, tutorials, anatomy labs, clinical skills sessions) delivered to students preceding the assessment date, which may have resulted in these problems being less vague or ambiguous when they were delivered to students in the context of their full preclinical learning experience in the years these questions were selected.
In terms of other study limitations, we highlight that ChatGPT-generated problems were reviewed with occasional stylistic or grammatical changes to ensure a high-fidelity comparison between groups. Such adjustments from our curriculum experts, while minor, may have slightly elevated the scores of ChatGPT-generated problems. As our primary intention was to compare ChatGPT-generated problems to previously used assessment materials, the historical set of problems was altered minimally, though we consider that more involved editing of historical problems (e.g., writing style) may reduce bias between groups for future iterations of this experiment. The current study includes only two curriculum subunits of preclinical medical education, and extrapolation to subunits beyond nephrology or hematology should be done with caution. Future iterations of this work for other subunits may include a multi-domain rubric (e.g., to separate comments on writing style and medical appropriateness), or may ask reviewers to guess whether each problem was a historical control problem or ChatGPT-generated as a means to measure bias towards higher or lower scores awarded for one group of problems over the other.
Regardless of these factors, our study suggests that using LLMs to produce free-text medical assessment problems is a feasible strategy in our undergraduate medical program. As with other AI-enabled tools in medicine, we consider this to be an augmentation rather than a replacement to our existing processes. While this study is unique in the literature for its focus on the free-text format for medical learners, several studies focusing on multiple-choice problems have reached similar conclusions, with unanimous concern that LLM-generated problems can contain mistakes [
11,
12,
13,
14,
15,
16] and must all be carefully reviewed by faculty with appropriate medical expertise prior to inclusion in an examination. This was illustrated with Q1 in our study, regarding megaloblastic anemia in the context of impaired absorption in a Crohn’s disease patient, where reviewers felt the timeline of the presenting illness and medication list for the patient were not realistic. One research group attempted to use an ensemble of several independent LLMs to identify multiple-choice problems with incorrect medical information, which improved efficiency, but was still not able to reach 100% recall in identifying problems with medical inaccuracies [
15]. We make further note of our reviewer comments regarding some problems being aligned with lower levels of Bloom’s taxonomy, such as “Recall”. This is despite our first prompt criterion in
Figure 1 requiring “the highest levels” of Bloom’s taxonomy, the highest of which is “Create”, which none of our problems reflected. Even with hyperbolic prompt language, several ChatGPT-generated problems remained among lower levels of Bloom’s taxonomy, suggesting an avenue for future study in prompt engineering.
While newer versions of ChatGPT are able to handle image-based media, both as inputs alongside prompts and as outputs, our study did not evaluate this capability. CAEs at our institution may include images, such as medical imaging, blood films, electrocardiograms, or clinical photographs with pertinent physical examination findings. Additionally, visual aids are often useful when tutors review correct answers with students following each CAE. The use of images in producing problem and answer key pairs would be a useful area of future research, whether images were produced originally by the LLM or supplied by the curriculum leaders in the prompting process.
Applying LLM-based technology to develop assessment problems is scalable and accessible, and has several advantages for medical students. Developing problems of appropriate difficulty and clinical relevance is a challenging and time-consuming process. Learners may use LLMs such as ChatGPT to produce supplementary practice problems according to their individual learning goals and educational backgrounds, focusing on testable concepts that they have not yet mastered. It has been suggested that learners may be more comfortable using an LLM than approaching professors or peers in this scenario, to avoid embarrassment or appearing unprepared [
21]. From the perspective of the curriculum leader, reducing the amount of time required to produce high-quality CAEs may allow for valuable time to be redirected to activities that are more learner-facing, such as providing support to underperforming students. Additionally, this would reduce the likelihood of assessment problems being recycled in consecutive years, reducing the risk of knowledge sharing between cohorts of students, therefore offering a benefit to academic integrity and the validity of the CAE assessments.
Beyond medical inaccuracy, another significant shortcoming of ChatGPT-generated problems for CAEs is hypothesized reliance on overly common or basic medical scenarios. We consider that LLMs are trained on a wide range of textual data, although this body of literature may be biased towards preparation for standardized and/or multiple-choice format examinations such as the USMLE. As a result, ChatGPT may struggle to produce novel clinical scenarios for very rare medical conditions, possibly limiting its utility for learners in these circumstances. ChatGPT is also unlikely to incorporate the collection of in-person and online learning activities that take place in our program in the weeks preceding each CAE to adapt assessment problems to the learning experience in each cohort. We would encourage faculty using LLMs as a tool for assessment generation to continue relying on their unique clinical experiences and curriculum knowledge to design CAE problems, especially in the free-text assessment format, where there is a benefit to having students respond with their complete thought process to unusual or atypical presentations of certain diseases.
There are many factors that should be considered prior to implementing this technology in an undergraduate medical program. First and foremost, faculty must be transparent about the use of LLMs in the generation of assessments, including the intended benefit to learners, the potential shortcomings of this approach, how assessment scores are produced, and the ongoing role of human faculty in the design process for medical assessments. As discussed previously, all items and answer keys must be carefully reviewed by faculty experts to ensure suitability. At our institution, while CAEs are an important learning tool to ensure students are making sufficient progress in meeting curriculum learning goals, CAE results do not appear on student academic transcripts and are not solely responsible for student progression through our program. Because of their lower stakes and formative nature, ethical concerns with implementing LLMs to generate CAEs are perhaps less nuanced compared to their implementation for higher stakes assessments, such as clerkship or licensing examinations. In either scenario, prior to implementation, potential biases or unintended influences inherent in AI-generated education content must be thoughtfully considered, communicated proactively with learners, and re-examined on a continuous basis.