Next Article in Journal
How Problem-Solving Attitudes Link Catastrophic Thinking to Environmental Awareness Among Egyptian University Students: A Structural Equation Modeling Approach
Previous Article in Journal
Experiential Avoidance and Psychoactive Substance Use: Systematic Review
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Performance of Large Language Models on Cognitive Aptitude Testing: A Multi-Run Evaluation on the German Medical School Admission Test (TMS)

1
Practices for Nuclear Medicine, Rubensstraße 125, 12157 Berlin, Germany
2
Department of Plastic, Aesthetic and Hand Surgery, Otto-von-Guericke University, 39120 Magdeburg, Germany
3
Department of Plastic Surgery and Hand Surgery, Gemeinschaftskrankenhaus Havelhoehe, Kladower Damm 221, 14089 Berlin, Germany
4
Department of Plastic Surgery and Hand Surgery, Burn Center, Medical Faculty, RWTH Aachen University, Pauwelsstrasse 30, 52074 Aachen, Germany
5
Department of Health Management, Friedrich-Alexander-Universität Erlangen-Nürnberg (FAU), Lange Gasse 20, 90403 Nürnberg, Germany
*
Author to whom correspondence should be addressed.
Eur. J. Investig. Health Psychol. Educ. 2026, 16(2), 23; https://doi.org/10.3390/ejihpe16020023
Submission received: 7 January 2026 / Revised: 3 February 2026 / Accepted: 9 February 2026 / Published: 12 February 2026

Abstract

Background and Objectives: Large language models (LLMs) have demonstrated high performance on knowledge-based medical examinations but their capabilities on cognitive aptitude tests emphasizing reasoning and abstraction remain underexplored. The Test for Medical Studies (TMS), a German medical school admission test, provides a standardized framework to examine these capabilities. This study aimed to evaluate the performance and consistency of multiple LLMs on text-based and visual-analytic TMS items. Materials and Methods: Eight contemporary LLMs, comprising proprietary and open-source systems, were evaluated using a multi-run design on standardized TMS items spanning text-based and visual-analytic cognitive domains. Results: Mean accuracy remained substantially below levels typically reported for knowledge-based medical examinations, with marked performance differences between text-based and visual-analytic subtests. Open-source models performed competitively compared with proprietary systems. Inter-run reliability was heterogeneous, indicating notable variability across repeated evaluations. Conclusions: Current LLMs show limited and domain-dependent performance on cognitive aptitude tasks relevant to medical school admission. High accuracy on knowledge-based examinations does not translate into stable performance on aptitude tests emphasizing fluid intelligence. The observed modality-dependent performance patterns and inter-run variability highlight the importance of differentiated, multi-run evaluation strategies when assessing LLMs for applications in medical education.
Keywords: large language models; artificial intelligence; Test for Medical Studies (TMS); medical school admission test; medical education; cognitive aptitude; clinical reasoning; medical student selection; psychometrics; benchmarking large language models; artificial intelligence; Test for Medical Studies (TMS); medical school admission test; medical education; cognitive aptitude; clinical reasoning; medical student selection; psychometrics; benchmarking

Share and Cite

MDPI and ACS Style

Stelling, H.; Kraus, A.; Grieb, G.; Güler, I. Performance of Large Language Models on Cognitive Aptitude Testing: A Multi-Run Evaluation on the German Medical School Admission Test (TMS). Eur. J. Investig. Health Psychol. Educ. 2026, 16, 23. https://doi.org/10.3390/ejihpe16020023

AMA Style

Stelling H, Kraus A, Grieb G, Güler I. Performance of Large Language Models on Cognitive Aptitude Testing: A Multi-Run Evaluation on the German Medical School Admission Test (TMS). European Journal of Investigation in Health, Psychology and Education. 2026; 16(2):23. https://doi.org/10.3390/ejihpe16020023

Chicago/Turabian Style

Stelling, Henrik, Armin Kraus, Gerrit Grieb, and Ibrahim Güler. 2026. "Performance of Large Language Models on Cognitive Aptitude Testing: A Multi-Run Evaluation on the German Medical School Admission Test (TMS)" European Journal of Investigation in Health, Psychology and Education 16, no. 2: 23. https://doi.org/10.3390/ejihpe16020023

APA Style

Stelling, H., Kraus, A., Grieb, G., & Güler, I. (2026). Performance of Large Language Models on Cognitive Aptitude Testing: A Multi-Run Evaluation on the German Medical School Admission Test (TMS). European Journal of Investigation in Health, Psychology and Education, 16(2), 23. https://doi.org/10.3390/ejihpe16020023

Article Metrics

Back to TopTop