Performance of Large Language Models on Cognitive Aptitude Testing: A Multi-Run Evaluation on the German Medical School Admission Test (TMS)
Abstract
1. Introduction
1.1. LLM Performance on Knowledge-Based Medical Examinations
1.2. Crystallized Versus Fluid Intelligence in Medical Assessment
1.3. Research Gap and Educational Relevance
- (1)
- evaluate model accuracy across verbal-conceptual and visual-spatial domains using a multi-run design
- (2)
- assess inter-run reliability and reasoning stability, which are critical for educational and assessment-related applications
- (3)
- identify performance differences between proprietary and open-source models
- (4)
- compare aptitude-based performance patterns with previously reported results on knowledge-based medical examinations
2. Materials and Methods
2.1. Study Design
2.2. AI Models
- Proprietary models (n = 5): GPT-5 (OpenAI, San Francisco, CA, USA), Grok 4 (xAI, San Francisco, CA, USA), Claude Sonnet 4.5 (Anthropic, San Francisco, CA, USA), Gemini 2.5 (Google LLC, Mountain View, CA, USA), and ERNIE 4.5 Turbo (Baidu, Beijing, China).
- Open-source models (n = 3): DeepSeek V3 (DeepSeek, Hangzhou, China), Qwen3-Max (Alibaba Cloud, Hangzhou, China), and Mistral Medium 3.1 (Mistral AI, Paris, France).
2.3. Examination Material and TMS Subtests
2.4. Modality-Based Evaluation Strategy
- Standard item set (144 items): Administered to seven models (GPT-5, Grok 4, Claude Sonnet 4.5, Gemini 2.5, ERNIE 4.5 Turbo, Qwen3-Max, and Mistral Medium 3.1). These models processed all six subtests listed in Table 1, including tasks requiring visual or visual-analytic reasoning.
- Text-only item set (72 items): Administered to DeepSeek V3. At the time of evaluation, image input in the evaluated model version was limited to optical character recognition (OCR) and did not support visual or visual-analytic reasoning. Consequently, the three subtests requiring such capabilities (Muster zuordnen, Schlauchfiguren, Diagramme und Tabellen) were excluded, and DeepSeek V3 was evaluated exclusively on the remaining text-based subtests.
- Aggregate text-based accuracy values reflect performance across all eight models, whereas visual-analytic accuracy values are based on the seven multimodal models.
2.5. Prompting Procedure and Data Collection
2.6. Statistical Analysis
3. Results
3.1. Overall Performance
3.2. Model-Level Performance and Ranking
3.2.1. Text-Based Subtests (72 Items)
3.2.2. Visual-Analytic Subtests (72 Items)
3.2.3. Combined Dataset (144 Items)
3.2.4. Inter-Run Variability
3.3. Proprietary Versus Open-Source Models
3.4. Inter-Run Reliability and Reasoning Stability
3.5. Pairwise Model Comparisons
4. Discussion
4.1. Main Findings
4.2. Crystallized Versus Fluid Intelligence
4.3. Visual-Analytic Items vs. Text-Based Items
4.4. Open-Source Versus Proprietary Models
4.5. Reliability and Implications for Medical Education
4.6. Limitations
- Restricted evaluation scope for DeepSeek V3: DeepSeek V3 was evaluated on a reduced, text-only subset of TMS items due to modality constraints, limiting direct comparability with multimodal models on visual and visual-analytic subtests.
- Zero-shot prompting paradigm: A standardized zero-shot prompting strategy was employed to reflect realistic user interaction. While this enhances ecological validity, alternative prompting approaches (e.g., structured reasoning prompts) may yield different performance profiles and were not examined in this study.
- Use of historical examination material: The TMS items used in this evaluation were drawn from previously administered examinations with officially published solutions. As these materials are available through public and print-based preparation resources, it cannot be excluded that some items were included in the training corpora of the evaluated models. However, this limitation applies broadly across models and does not explain the observed differences in inter-run reliability.
- Limited control over generation parameters: All models were accessed via official web interfaces without direct control over temperature or other decoding parameters. As a result, internal sampling strategies may have varied across models and over time, potentially contributing to inter-run variability.
- Exclusion of specific cognitive subtests: Subtests primarily targeting short-term memory (Learning Figures, Learning Facts) and sustained attention/processing speed (Concentrated Work) were excluded. Consequently, the present analysis does not represent a full benchmark of TMS performance across all cognitive domains. However, these subtests are conceptually and technically ill-suited for evaluation with current LLM architectures.
- Absence of direct human performance comparison: No direct comparison with human TMS performance was conducted, precluding assessment of whether LLM performance approaches, matches, or exceeds typical applicant outcomes. While human test participants receive both absolute scores and percentile ranks, admission decisions are primarily based on the percentile ranks (Test für Medizinische Studiengänge [TMS], n.d.). Since the present study evaluated a selected subset of items without access to a representative human reference cohort, a valid mapping of LLM performance to human standard values or percentile ranks was not feasible. Accordingly, the results are intended to characterize relative performance patterns across models rather than direct equivalence to human test-takers.
- Dynamic nature of deployed LLM systems: The evaluated models are continuously updated by their providers, and their underlying parameters, training data, and inference behavior may change over time (L. Chen et al., 2023). As a result, performance estimates may vary across evaluation dates. While offline evaluation of fixed model checkpoints would improve reproducibility, such an approach would not reflect the real-world usage scenario of most end users, who typically interact with continuously updated web-based systems and lack the computational resources or technical expertise required for local deployment.
- Choice of reliability metric: Cohen’s κ was selected as the primary inter-run reliability measure because it quantifies categorical agreement between paired runs while correcting for chance agreement, which is appropriate for evaluating whether models produce consistent responses to identical items. Importantly, κ captures response consistency rather than reasoning quality; high agreement may reflect stable but incorrect reasoning patterns, while low agreement indicates response variability regardless of underlying accuracy. Other studies may prefer Krippendorff’s α to estimate overall reliability across multiple runs.
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AI | Artificial Intelligence |
| CI | Confidence Interval |
| CV | Coefficient of Variation |
| ITB | Institute for Test and Talent Research (Institut für Test- und Begabungsforschung) |
| κ | Cohen’s Kappa |
| LLM | Large Language Model |
| M2 | Second State Examination (German medical licensing examination) |
| MLLM | Multimodal Large Language Model |
| OCR | Optical Character Recognition |
| SD | Standard Deviation |
| TMS | Test for Medical Studies (Test für Medizinische Studiengänge; “Medizinertest”) |
| USMLE | United States Medical Licensing Examination |
References
- Alvarado Gonzalez, M. A., Bruno Hernandez, M., Peñaloza Perez, M. A., Lopez Orozco, B., Cruz Soto, J. T., & Malagon, S. (2025). Do repetitions matter? Strengthening reliability in LLM evaluations. arXiv, arXiv:2509.24086. [Google Scholar] [CrossRef] [Scilit]
- Bicknell, B. T., Butler, D., Whalen, S., Ricks, J., Dixon, C. J., Clark, A. B., Spaedy, O., Skelton, A., Edupuganti, N., Dzubinski, L., Tate, H., Dyess, G., Lindeman, B., & Lehmann, L. S. (2024). ChatGPT-4 Omni performance in USMLE disciplines and clinical skills: Comparative analysis. JMIR Medical Education, 10, e63430. [Google Scholar] [CrossRef] [Scilit]
- Buckley, T. A., Crowe, B., Abdulnour, R. E., Rodman, A., & Manrai, A. K. (2025). Comparison of frontier open-source and proprietary large language models for complex diagnoses. JAMA Health Forum, 6(3), e250040. [Google Scholar] [CrossRef] [Scilit]
- Chen, L., Zaharia, M., & Zou, J. (2023). How is ChatGPT’s behavior changing over time? arXiv, arXiv:2307.09009. [Google Scholar] [CrossRef] [Scilit]
- Chen, Y., Huang, X., Yang, F., Lin, H., Lin, H., Zheng, Z., Liang, Q., Zhang, J., & Li, X. (2024). Performance of ChatGPT and Bard on the medical licensing examinations varies across different cultures: A comparison study. BMC Medical Education, 24(1), 1372. [Google Scholar] [CrossRef] [Scilit]
- Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37–46. [Google Scholar] [CrossRef] [Scilit]
- Elkin, P. L., Mehta, G., LeHouillier, F., Resnick, M., Mullin, S., Tomlin, C., Resendez, S., Liu, J., Nebeker, J. R., & Brown, S. H. (2025). Semantic clinical artificial intelligence vs. native large language model performance on the USMLE. JAMA Network Open, 8(4), e256359. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Gilson, A., Safranek, C. W., Huang, T., Socrates, V., Chi, L., Taylor, R. A., & Chartash, D. (2023). How does ChatGPT perform on the United States Medical Licensing Examination (USMLE)? The implications of large language models for medical education and knowledge assessment. JMIR Medical Education, 9, e45312. [Google Scholar] [CrossRef] [Scilit]
- Hell, B., Trapmann, S., & Schuler, H. (2007). Eine metaanalyse der validität von fachspezifischen studierfähigkeitstests im deutschsprachigen Raum. Empirische Pädagogik, 21(3), 251–270. Available online: https://www.researchgate.net/publication/260530228_Eine_Metaanalyse_der_Validitat_von_fachspezifischen_Studierfahigkeitstests_im_deutschsprachigen_Raum (accessed on 22 December 2025).
- Ilić, D., & Gignac, G. E. (2024). Evidence of interrelated cognitive-like capabilities in large language models: Indications of artificial general intelligence or achievement? Intelligence, 106, 101858. [Google Scholar] [CrossRef] [Scilit]
- ITB-Academic Tests. (n.d.). TMS—Test für medizinische Studiengänge. Available online: https://itb-academic-tests.org/hochschulvertreter/tms (accessed on 22 December 2025).
- Kadmon, G., & Kadmon, M. (2016). Academic performance of students with the highest and mediocre school-leaving grades: Does the aptitude test for medical studies (TMS) balance their prognoses? GMS Journal for Medical Education, 33(1), Doc7. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kung, T. H., Cheatham, M., Medenilla, A., Sillos, C., De Leon, L., Elepaño, C., Madriaga, M., Aggabao, R., Diaz-Candido, G., Maningo, J., & Tseng, V. (2023). Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models. PLoS Digital Health, 2(2), e0000198. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159–174. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Li, C., Wu, W., Zhang, H., Li, Q., Gao, Z., Xia, Y., Hernández-Orallo, J., Vulić, I., & Wei, F. (2025). 11Plus-Bench: Demystifying multimodal LLM spatial reasoning with cognitive-inspired analysis. arXiv, arXiv:2508.20068. [Google Scholar] [CrossRef] [Scilit]
- Liu, M., Okuhara, T., Chang, X., Shirabe, R., Nishiie, Y., Okada, H., & Kiuchi, T. (2024). Performance of ChatGPT across different versions in medical licensing examinations worldwide: Systematic review and meta-analysis. Journal of Medical Internet Research, 26, e60807. [Google Scholar] [CrossRef] [Scilit]
- López Espejel, J., Ettifouri, E. H., Alassan, M. S. Y., Chouham, E. M., & Dahhane, W. (2023). GPT-3.5, GPT-4, or BARD? Evaluating LLMs’ reasoning ability in zero-shot setting and performance boosting through prompts. Natural Language Processing Journal, 5, 100032. [Google Scholar] [CrossRef] [Scilit]
- Mavrych, V., Yaqinuddin, A., & Bolgova, O. (2025). Claude, ChatGPT, Copilot, and Gemini performance versus students in different topics of neuroscience. Advances in Physiology Education, 49(2), 430–437. [Google Scholar] [CrossRef] [Scilit]
- Meyer, A., Riese, J., & Streichert, T. (2024). Comparison of the performance of GPT-3.5 and GPT-4 with that of medical students on the written German medical licensing examination: Observational study. JMIR Medical Education, 10, e50965. [Google Scholar] [CrossRef] [Scilit]
- Nori, H., King, N., Mayer McKinney, S., Carignan, D., & Horvitz, E. (2023). Capabilities of GPT-4 on medical challenge problems. arXiv, arXiv:2303.13375. [Google Scholar] [CrossRef] [Scilit]
- Schult, J., Hofmann, A., & Stegt, S. J. (2019). Leisten fachspezifische studierfähigkeitstests im deutschsprachigen raum eine valide studienerfolgsprognose? Ein metaanalytisches update. Zeitschrift für Entwicklungspsychologie und Pädagogische Psychologie, 51(1), 16–30. [Google Scholar] [CrossRef] [Scilit]
- Test für Medizinische Studiengänge (TMS). (n.d.). Results and scoring of the test for medical studies (TMS). TMS-info. Available online: https://www.tms-info.org/ergebnis-und-auswertung/ (accessed on 23 December 2025).
- TMS. (n.d.). Informationen der am TMS beteiligten organisationen. Available online: https://www.tms-info.org/informationen-der-am-tms-beteiligten-fakultaeten (accessed on 22 December 2025).
- Trost, G. (1992). Erfahrungen mit dem test für medizinische Studiengänge (TMS). Medizinische Ausbildung, 9(2), 67–76. Available online: https://gesellschaft-medizinische-ausbildung.org/files/ZMA-Archiv/1992/2/Trost_G-n.pdf (accessed on 22 December 2025).
- Yang, Y., Chen, M., Liu, Q., Hu, M., Chen, Q., Zhang, G., Hu, S., Zhai, G., Qiao, Y., Wang, Y., Shao, W., & Luo, P. (2025). Truly assessing fluid intelligence of large language models through dynamic reasoning evaluation. arXiv, arXiv:2506.02648. [Google Scholar] [CrossRef] [Scilit]
- Zimmerhofer, A., Hofmann, A., Wittenberg, T., Amelung, D., & Kadmon, M. (2019, September 17). Test for medical studies (TMS): Testing cognitive competence in applicants for medical school [Conference presentation]. 16th DPPD Conference, Heidelberg, Germany. Available online: https://www.researchgate.net/publication/355046109_Test_for_Medical_Studies_TMS_Testing_Cognitive_Competence_in_Applicants_for_Medical_School (accessed on 22 December 2025).
- Zong, H., Wu, R., Cha, J., Wang, J., Wu, E., Li, J., Zhou, Y., Zhang, C., Feng, W., & Shen, B. (2024). Large language models in worldwide medical exams: Platform development and comprehensive analysis. Journal of Medical Internet Research, 26, e66114. [Google Scholar] [CrossRef] [Scilit] [PubMed]


| Subtest (German Original) | Cognitive Domain & Task Description | N Items | Modality |
|---|---|---|---|
| Medizinisch- naturwissenschaftliches Grundverständnis | Deductive reasoning based on short medical–scientific texts without requiring prior knowledge | 24 | Text |
| Quantitative und formale Probleme | Mathematical and logical reasoning embedded in biomedical contexts | 24 | Text |
| Textverständnis | Synthesis and interpretation of complex scientific texts | 24 | Text |
| Diagramme und Tabellen | Visual-analytic reasoning involving graphs, curves, and tabulated data | 24 | Visual |
| Muster zuordnen | Perceptual precision and error detection in complex graphical representations | 24 | Visual |
| Schlauchfiguren | Spatial reasoning requiring mental rotation of three-dimensional objects | 24 | Visual |
| Total included items | 144 |
| Panel A: Text-Based Subtests (72 Items)—All 8 Models | ||||||||
| Rank | Model | Type | N | Mean Accuracy (%) | 95% CI | SD | CV (%) | Range |
| 1 | Claude Sonnet 4.5 | Prop. | 72 | 77.8 | 71.6–84.0 | 7.08 | 9.1 | 18 |
| 2 | Grok 4 | Prop. | 72 | 75.6 | 67.2–83.9 | 9.55 | 12.6 | 26 |
| 3 | Qwen3-Max | Open | 72 | 73.3 | 68.3–78.4 | 5.76 | 7.9 | 15 |
| 4 | Gemini 2.5 | Prop. | 72 | 72.5 | 65.0–80.0 | 8.53 | 11.8 | 16 |
| 5 | ERNIE 4.5 Turbo | Prop. | 72 | 71.9 | 69.2–74.7 | 3.17 | 4.4 | 8 |
| 6 | GPT-5 | Prop. | 72 | 71.1 | 66.9–75.4 | 4.85 | 6.8 | 12 |
| 7 | Mistral Medium 3.1 | Open | 72 | 69.4 | 64.1–74.8 | 6.13 | 8.8 | 15 |
| 8 | DeepSeek V3 | Open | 72 | 68.6 | 65.0–72.2 | 4.12 | 6.0 | 9 |
| Panel B: Visual-Analytic Subtests (72 Items)—7 Models | ||||||||
| Rank | Model | Type | N | Mean Accuracy (%) | 95% CI | SD | CV (%) | Range |
| 1 | Qwen3-Max | Open | 72 | 63.1 | 59.1–67.0 | 4.46 | 7.1 | 11 |
| 2 | ERNIE 4.5 Turbo | Prop. | 72 | 58.3 | 54.1–62.6 | 4.81 | 8.2 | 11 |
| 3 | GPT-5 | Prop. | 72 | 58.1 | 52.1–64.0 | 6.76 | 11.6 | 15 |
| 4 | Mistral Medium 3.1 | Open | 72 | 57.8 | 53.8–61.8 | 4.56 | 7.9 | 12 |
| 5 | Claude Sonnet 4.5 | Prop. | 72 | 56.9 | 50.7–63.2 | 7.15 | 12.6 | 18 |
| 6 | Gemini 2.5 | Prop. | 72 | 49.7 | 41.1–58.3 | 9.79 | 19.7 | 25 |
| 7 | Grok 4 | Prop. | 72 | 46.4 | 37.6–55.2 | 10 | 21.5 | 22 |
| Panel C: Combined Dataset (144 Items)—7 Models | ||||||||
| Rank | Model | Type | N | Mean Accuracy (%) | 95% CI | SD | CV (%) | Range |
| 1 | Qwen3-Max | Open | 144 | 68.2 | 66.9–69.4 | 1.42 | 2.1 | 2 |
| 2 | Claude Sonnet 4.5 | Prop. | 144 | 67.4 | 62.9–71.9 | 5.13 | 7.6 | 12 |
| 3 | ERNIE 4.5 Turbo | Prop. | 144 | 65.1 | 62.6–67.7 | 2.88 | 4.4 | 6 |
| 4 | GPT-5 | Prop. | 144 | 64.6 | 59.9–69.2 | 5.29 | 8.2 | 13 |
| 5 | Mistral Medium 3.1 | Open | 144 | 63.6 | 59.8–67.4 | 4.3 | 6.8 | 10 |
| 6 | Gemini 2.5 | Prop. | 144 | 61.1 | 53.7–68.5 | 8.43 | 13.8 | 20 |
| 7 | Grok 4 | Prop. | 144 | 61.0 | 53.3–68.6 | 8.75 | 14.3 | 19 |
| Model | Mean κ | Min κ | Max κ | Interpretation |
|---|---|---|---|---|
| DeepSeek V3 | 0.676 | 0.543 | 0.788 | Substantial |
| Mistral Medium 3.1 | 0.452 | 0.363 | 0.567 | Moderate |
| Qwen3-Max | 0.424 | 0.234 | 0.617 | Moderate |
| Gemini 2.5 | 0.355 | 0.143 | 0.724 | Fair |
| Claude Sonnet 4.5 | 0.347 | 0.223 | 0.567 | Fair |
| Grok 4 | 0.340 | 0.060 | 0.671 | Fair |
| GPT-5 | 0.307 | 0.040 | 0.487 | Fair |
| ERNIE 4.5 Turbo | 0.208 | 0.025 | 0.338 | Slight |
| Model A | Model B | N | χ2 | p-Value | OR | Sig |
|---|---|---|---|---|---|---|
| Claude Sonnet 4.5 | DeepSeek V3 † | 72 | 10.29 | 0.0013 | 13.0 | *** |
| Claude Sonnet 4.5 | Gemini 2.5 | 144 | 8.00 | 0.0047 | 3.0 | ** |
| ERNIE 4.5 Turbo | Gemini 2.5 | 144 | 6.43 | 0.0112 | 2.5 | * |
| Qwen3-Max | Gemini 2.5 | 144 | 5.49 | 0.0191 | 2.15 | * |
| DeepSeek V3 † | Grok 4 | 72 | 5.33 | 0.0209 | 0.2 | * |
| GPT-5 | Gemini 2.5 | 144 | 4.83 | 0.0280 | 2.18 | * |
| ERNIE 4.5 Turbo | DeepSeek V3 † | 72 | 4.00 | 0.0455 | 3.0 | * |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Published by MDPI on behalf of the University Association of Education and Psychology. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Stelling, H.; Kraus, A.; Grieb, G.; Güler, I. Performance of Large Language Models on Cognitive Aptitude Testing: A Multi-Run Evaluation on the German Medical School Admission Test (TMS). Eur. J. Investig. Health Psychol. Educ. 2026, 16, 23. https://doi.org/10.3390/ejihpe16020023
Stelling H, Kraus A, Grieb G, Güler I. Performance of Large Language Models on Cognitive Aptitude Testing: A Multi-Run Evaluation on the German Medical School Admission Test (TMS). European Journal of Investigation in Health, Psychology and Education. 2026; 16(2):23. https://doi.org/10.3390/ejihpe16020023
Chicago/Turabian StyleStelling, Henrik, Armin Kraus, Gerrit Grieb, and Ibrahim Güler. 2026. "Performance of Large Language Models on Cognitive Aptitude Testing: A Multi-Run Evaluation on the German Medical School Admission Test (TMS)" European Journal of Investigation in Health, Psychology and Education 16, no. 2: 23. https://doi.org/10.3390/ejihpe16020023
APA StyleStelling, H., Kraus, A., Grieb, G., & Güler, I. (2026). Performance of Large Language Models on Cognitive Aptitude Testing: A Multi-Run Evaluation on the German Medical School Admission Test (TMS). European Journal of Investigation in Health, Psychology and Education, 16(2), 23. https://doi.org/10.3390/ejihpe16020023

