Comparing Human and ChatGPT Performance on the Force Concept Inventory: An Item-Level Analysis Through Rasch Modeling and the Theory of Conceptual Fields
Abstract
1. Introduction
2. Theoretical Framework
2.1. The Force Concept Inventory
2.2. The Rasch Model as an Analytical Framework in Science Education
2.3. The Theory of Conceptual Fields
3. Materials and Methods
3.1. Study Design
3.2. Participants
3.3. Data Collection
3.4. Analytical Strategy
3.5. Methodological Limitations
4. Results2
4.1. Overall Conceptual Performance
4.2. Item-Level Comparison
4.3. Item Characteristic Curves
5. Discussion
5.1. Beyond Overall Scores: What Do Rasch Estimates Reveal?
5.2. Interpreting Human–AI Differences Through the Theory of Conceptual Fields
5.3. Educational Implications
5.4. Methodological Limitations and Future Research
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Abbreviations
| ChatGPT | Chat Generative Pre-trained Transformer |
| GenAI | Generative Artificial Intelligence |
| FCI | Force Concept Inventory |
| IRT | Item Response Theory |
| TCF | Theory of Conceptual Fields |
| ECCE | Electric Circuits Conceptual Evaluation (ECCE) |
Appendix A
Items Used in Our Analysis








Appendix B
Item’s Difficulties

| 1 | The Force Concept Inventory and answer key are available from https://www.physport.org/assessments/assessment.cfm?A=FCI (accessed on 10 January 2026). |
| 2 | The items included in this analysis are presented in the Appendix A. |
References
- Aldazharova, S., Issayeva, G., Maxutov, S., & Balta, N. (2024). Assessing AI’s problem solving in physics: Analyzing reasoning, false positives and negatives through the Force Concept Inventory. Contemporary Educational Technology, 16(4), ep538. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Artopoulos, A., & Lliteras, A. (2024). La emergencia de la alfabetización crítica en IA: La reconstrucción social de la ciudadanía en democracias bajo acecho digital. Revista Diálogo Educacional, 24(83), 1283–1304. [Google Scholar] [CrossRef] [Scilit]
- Azambuja, C. C. d., & Ferreira da Silva, G. (2024). Novos desafios para a educação na era da inteligência artificial. Filosofia Unisinos, 25(1), e25107. [Google Scholar] [CrossRef] [Scilit]
- Bang, Y., Cahyawijaya, S., Lee, N., Dai, W., Su, D., Wilie, B., Lovenia, H., Ji, Z., Yu, T., Chung, W., Do, Q. V., Xu, Y., & Fung, P. (2023). A multitask, multilingual, multimodal evaluation of ChatGPT on reasoning, hallucination, and interactivity. arXiv. [Google Scholar] [CrossRef] [Scilit]
- Bond, T. G., & Fox, C. M. (2015). Applying the Rasch model: Fundamental measurement in the human sciences (4th ed.). Routledge. [Google Scholar]
- Boone, W. J., Staver, J. R., & Yale, M. S. (2014). Rasch analysis in the human sciences. Springer. [Google Scholar]
- Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M. T., & Zhang, Y. (2023). Sparks of artificial general intelligence: Early experiments with GPT-4. arXiv. [Google Scholar] [CrossRef] [Scilit]
- Carvalho Junior, G. D. (2024). Les invariants opératoires dans le processus de conceptualisation en physique thermique. Carrefours de l’Éducation, 57, 115–130. [Google Scholar] [CrossRef] [Scilit]
- Castañeda, L., Haba-Ortuño, I., Villar-Onrubia, D., Marín, V. I., Tur, G., Ruipérez-Valiente, J. A., & Wasson, B. (2024). Desarrollando el marco DALI de alfabetización en datos para la ciudadanía. RIED—Revista Iberoamericana de Educación a Distancia, 27(1), 289–318. [Google Scholar] [CrossRef] [Scilit]
- Chalmers, R. P. (2012). mirt: A multidimensional item response theory package for the R environment. Journal of Statistical Software, 48(6), 1–29. [Google Scholar] [CrossRef] [Scilit]
- Clement, J. (1982). Students’ preconceptions in introductory mechanics. American Journal of Physics, 50(1), 66–71. [Google Scholar] [CrossRef] [Scilit]
- Consoli, T., Schmitz, M.-L., Antonietti, C., Gonon, P., Cattaneo, A., & Petko, D. (2025). Quality of technology integration matters: Positive associations between high-quality digital use and student outcomes. Education and Information Technologies, 30, 7719–7752. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Cuban, L. (2001). Oversold and underused: Computers in the classroom. Harvard University Press. [Google Scholar]
- Deane, T., Nomme, K., Jeffery, E., Pollock, C., & Birol, G. (2016). Development of the Statistical Reasoning in Biology Concept Inventory (SRBCI): A Rasch model approach. CBE—Life Sciences Education, 15(1), ar5. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Fernández, C. G., & Calderón-Garrido, D. (2024). De la educabilidad a la aceptación de la tecnología y alfabetización en inteligencia artificial: Validación de un instrumento. Digital Education Review, (45), 8–14. [Google Scholar] [CrossRef] [Scilit]
- Hake, R. R. (1998). Interactive-engagement versus traditional methods: A six-thousand-student survey of mechanics test data for introductory physics courses. American Journal of Physics, 66(1), 64–74. [Google Scholar] [CrossRef] [Scilit]
- Hestenes, D., Wells, M., & Swackhamer, G. (1992). Force concept inventory. The Physics Teacher, 30(3), 141–158. [Google Scholar] [CrossRef] [Scilit]
- Kaddour, J., Harris, J., Mozes, M., Bradley, H., Raileanu, R., & McHardy, R. (2023). Challenges and applications of large language models. arXiv. [Google Scholar] [CrossRef] [Scilit]
- Kortemeyer, G. (2023). Could an artificial-intelligence agent pass an introductory physics course? Physical Review Physics Education Research, 19, 010132. [Google Scholar] [CrossRef] [Scilit]
- Lobet, M., Honet, A., Romainville, M., & Wathelet, V. (2024). ChatGPT: Quel en a été l’usage spontané d’étudiants de première année universitaire à son arrivée ? Erudit, 18, 67–90. [Google Scholar] [CrossRef] [Scilit]
- López-Simó, V., & Rezende, M. F. (2024). Challenging ChatGPT with different types of physics education questions. The Physics Teacher, 62(4), 290–294. [Google Scholar] [CrossRef] [Scilit]
- Maharana, A., Lee, D.-H., Tulyakov, S., Bansal, M., Barbieri, F., & Fang, Y. (2024, August 11–16). Evaluating very long-term conversational memory of large language models. 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024), Bangkok, Thailand. [Google Scholar]
- Mazur, E. (1997). Peer instruction: A user’s manual. Prentice Hall. [Google Scholar]
- McDermott, L. C. (1984). Research on conceptual understanding in mechanics. Physics Today, 37(7), 24–32. [Google Scholar] [CrossRef] [Scilit]
- Planinić, M., Ivanjek, L., & Sušac, A. (2010). Rasch model based analysis of the force concept inventory. Physical Review Special Topics—Physics Education Research, 6(1), 010103. [Google Scholar] [CrossRef] [Scilit]
- Thornton, R. K., & Sokoloff, D. R. (1998). Assessing student learning of Newton’s laws: The force and motion conceptual evaluation. American Journal of Physics, 66(4), 338–352. [Google Scholar] [CrossRef] [Scilit]
- Vergnaud, G. (1991). La théorie des champs conceptuels. Recherches en Didactique des Mathématiques, 10(2–3), 133–170. [Google Scholar]
- Vergnaud, G. (1996). The theory of conceptual fields. In L. P. Steffe, & P. Nesher (Eds.), Theories of mathematical learning (pp. 219–239). Springer. [Google Scholar]
- Vergnaud, G. (2002). Qu’est-ce qu’apprendre? Conférence introductive au Colloque International de l’IUFM de l’Académie de Créteil. Available online: https://www.gerard-vergnaud.org/GVergnaud_2002_Qu-Est-Ce-QuApprendre_Colloque-IUFM-Creteil (accessed on 12 December 2025).
- Wang, J., & Bao, L. (2010). An investigation of student conceptual learning gains in physics: The role of item difficulty. Physical Review Special Topics—Physics Education Research, 6, 010105. [Google Scholar]
- Wang, Q., Ding, L., Cao, Y., Tian, Z., Wang, S., Tao, D., & Guo, L. (2023). Recursively summarizing enables long-term dialogue memory in large language models. arXiv. [Google Scholar] [CrossRef] [Scilit]
- West, C. G. (2023). AI and the FCI: Can ChatGPT project an understanding of introductory physics? arXiv. [Google Scholar] [CrossRef] [Scilit]
- Wheeler, S., & Scherr, R. E. (2023). ChatGPT reflects student misconceptions in physics. In Proceedings of the Physics Education Research Conference 2023 (pp. 386–390). American Association of Physics Teachers. [Google Scholar] [CrossRef] [Scilit]
- Zhong, W., Guo, L., Gao, Q., Ye, H., & Wang, Y. (2023). Enhancing large language models with long-term memory. arXiv. [Google Scholar] [CrossRef] [Scilit]






| Group | N | Mean | SD | Min | Max |
|---|---|---|---|---|---|
| ChatGPT | 11 | 452.8 | 65.8 | 389.57 | 585.45 |
| Humans | 14 | 537.13 | 103.77 | 389.57 | 733.85 |
| Item | b (Difficulty) | % Humans | % ChatGPT | Diff |
|---|---|---|---|---|
| 11 | 0.07 | 85.7 | 0.0 | 85.7 |
| 20 | −0.11 | 85.7 | 9.1 | 76.6 |
| 21 | −0.29 | 78.6 | 27.3 | 51.3 |
| 09 | −0.29 | 78.6 | 27.3 | 51.3 |
| 26 | 0.45 | 35.7 | 45.5 | −9.8 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Carvalho Junior, G.D.d.; Rezende Junior, M.F.; Lobet, M.; Carvalho, A.X.Z.d. Comparing Human and ChatGPT Performance on the Force Concept Inventory: An Item-Level Analysis Through Rasch Modeling and the Theory of Conceptual Fields. Educ. Sci. 2026, 16, 1154. https://doi.org/10.3390/educsci16071154
Carvalho Junior GDd, Rezende Junior MF, Lobet M, Carvalho AXZd. Comparing Human and ChatGPT Performance on the Force Concept Inventory: An Item-Level Analysis Through Rasch Modeling and the Theory of Conceptual Fields. Education Sciences. 2026; 16(7):1154. https://doi.org/10.3390/educsci16071154
Chicago/Turabian StyleCarvalho Junior, Gabriel Dias de, Mikael Frank Rezende Junior, Michaël Lobet, and Andressa Xavier Zinato de Carvalho. 2026. "Comparing Human and ChatGPT Performance on the Force Concept Inventory: An Item-Level Analysis Through Rasch Modeling and the Theory of Conceptual Fields" Education Sciences 16, no. 7: 1154. https://doi.org/10.3390/educsci16071154
APA StyleCarvalho Junior, G. D. d., Rezende Junior, M. F., Lobet, M., & Carvalho, A. X. Z. d. (2026). Comparing Human and ChatGPT Performance on the Force Concept Inventory: An Item-Level Analysis Through Rasch Modeling and the Theory of Conceptual Fields. Education Sciences, 16(7), 1154. https://doi.org/10.3390/educsci16071154

