Evaluation of Large Language Models in Generating Physical Exercise Rehabilitation Programs for Musculoskeletal Disorders Across Multiple Clinical Scenarios
Highlights
- Four major large language models (LLMs) generated highly structured and consistent physical exercise rehabilitation programs for musculoskeletal disorders, achieving moderate-to-high quality with a mean DISCERN score of 55.60.
- Readability assessments across six validated indices revealed that the LLM-generated responses consistently required reading levels significantly higher than standard patient comprehension recommendations.
- While LLMs demonstrate substantial potential as auxiliary tools in digital orthopedics, their lack of supplemental resources and high reading difficulty hinder independent patient use.
- Professional physician oversight remains indispensable to bridge readability gaps, refine exercise descriptions, and ensure patient safety in AI-driven health consultations.
Abstract
1. Introduction
2. Methods
2.1. Clinical Scenario Simulation
2.1.1. Scenario-1
2.1.2. Scenario-2
2.1.3. Scenario-3
2.2. Prompt Design
2.3. Question Input
2.4. Evaluation Procedures
2.4.1. Quality Assessment
2.4.2. Readability Evaluation
2.4.3. Statistical Analysis
3. Results
4. Discussion
Limitations
5. Conclusions
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Abbreviations
| LLM | Large Language Model |
| OA | Osteoarthritis |
| LBP | Low Back Pain |
| SP | Shoulder Pain |
| NP | Neck Pain |
| MSK | Musculoskeletal |
| PE | Physical Exercise |
| GPT-4 | Generative Pre-trained Transformer Version 4 |
| BMI | Body Mass Index |
| AI | Artificial Intelligence |
| ACSM | American College of Sports Medicine |
| FAQs | Frequently Asked Questions |
| CI | Confidence Interval |
| FKGL | Flesch–Kincaid Grade Level |
| FRES | Flesch Reading Ease Score |
| SMOG | Simplified Measure of Gobbledegook |
| GF | Gunning Fog |
| ARI | Automated Readability Index |
| CLI | Coleman–Liau Index |
| USDHHS | US Department of Health and Human Services |
| SD | Standard Deviation |
| SE | Standard Errors |
| ANOVA | Analysis of Variance |
References
- GBD 2021 Diseases and Injuries Collaborators. Global incidence, prevalence, years lived with disability (YLDs), disability-adjusted life-years (DALYs), and healthy life expectancy (HALE) for 371 diseases and injuries in 204 countries and territories and 811 subnational locations, 1990–2021: A systematic analysis for the Global Burden of Disease Study 2021. Lancet 2024, 403, 2133–2161. [Google Scholar] [CrossRef] [PubMed]
- Greggi, C.; Visconti, V.V.; Albanese, M.; Gasperini, B.; Chiavoghilefu, A.; Prezioso, C.; Persechino, B.; Iavicoli, S.; Gasbarra, E.; Iundusi, R.; et al. Work-Related Musculoskeletal Disorders: A Systematic Review and Meta-Analysis. J. Clin. Med. 2024, 13, 3964. [Google Scholar] [CrossRef] [PubMed]
- Institute of Medicine (US) Committee on Advancing Pain Research, Care, and Education. Relieving Pain in America: A Blueprint for Transforming Prevention, Care, Education, and Research; National Academies Press: Washington, DC, USA, 2011. Available online: http://www.ncbi.nlm.nih.gov/books/NBK91497/ (accessed on 8 July 2023).
- Safiri, S.; Kolahi, A.A.; Cross, M.; Carson-Chahhoud, K.; Almasi-Hashiani, A.; Kaufman, J.; Mansournia, M.A.; Sepidarkish, M.; Ashrafi-Asgarabad, A.; Hoy, D.; et al. Global, regional, and national burden of other musculoskeletal disorders 1990–2017: Results from the Global Burden of Disease Study 2017. Rheumatology 2021, 60, 855–865. [Google Scholar] [PubMed]
- Katz, J.N.; Arant, K.R.; Loeser, R.F. Diagnosis and Treatment of Hip and Knee Osteoarthritis: A Review. JAMA 2021, 325, 568. [Google Scholar] [CrossRef] [PubMed]
- Maestroni, L.; Read, P.; Bishop, C.; Papadopoulos, K.; Suchomel, T.J.; Comfort, P.; Turner, A. The Benefits of Strength Training on Musculoskeletal System Health: Practical Applications for Interdisciplinary Care. Sports Med. 2020, 50, 1431–1450. [Google Scholar] [CrossRef] [PubMed]
- Skou, S.T.; Roos, E.M. Physical therapy for patients with knee and hip osteoarthritis: Supervised, active treatment is current best practice. Clin. Exp. Rheumatol. 2019, 37, 112–117. [Google Scholar] [PubMed]
- Knezevic, N.N.; Candido, K.D.; Vlaeyen, J.W.S.; Van Zundert, J.; Cohen, S.P. Low back pain. Lancet 2021, 398, 78–92. [Google Scholar] [CrossRef] [PubMed]
- Henriksen, M.; Hansen, J.B.; Klokker, L.; Bliddal, H.; Christensen, R. Comparable effects of exercise and analgesics for pain secondary to knee osteoarthritis: A meta-analysis of trials included in Cochrane systematic reviews. J. Comp. Eff. Res. 2016, 5, 417–431. [Google Scholar] [CrossRef] [PubMed]
- Cohen, S.P.; Hooten, W.M. Advances in the diagnosis and management of neck pain. BMJ 2017, 358, j3221. [Google Scholar] [CrossRef] [PubMed]
- Steuri, R.; Sattelmayer, M.; Elsig, S.; Kolly, C.; Tal, A.; Taeymans, J.; Hilfiker, R. Effectiveness of conservative interventions including exercise, manual therapy and medical management in adults with shoulder impingement: A systematic review and meta-analysis of RCTs. Br. J. Sports Med. 2017, 51, 1340–1347. [Google Scholar] [CrossRef] [PubMed]
- Chen, X.; Giles, J.; Yao, Y.; Yip, W.; Meng, Q.; Berkman, L.; Chen, H.; Chen, X.; Feng, J.; Feng, Z.; et al. The path to healthy ageing in China: A Peking University–Lancet Commission. Lancet 2022, 400, 1967–2006. [Google Scholar] [CrossRef] [PubMed]
- Cieza, A.; Causey, K.; Kamenov, K.; Hanson, S.W.; Chatterji, S.; Vos, T. Global estimates of the need for rehabilitation based on the Global Burden of Disease study 2019: A systematic analysis for the Global Burden of Disease Study 2019. Lancet 2020, 396, 2006–2017. [Google Scholar] [CrossRef] [PubMed]
- Goh, S.-L.; Persson, M.S.M.; Stocks, J.; Hou, Y.; Lin, J.; Hall, M.C.; Doherty, M.; Zhang, W. Efficacy and potential determinants of exercise therapy in knee and hip osteoarthritis: A systematic review and meta-analysis. Ann. Phys. Rehabil. Med. 2019, 62, 356–365. [Google Scholar] [CrossRef] [PubMed]
- Tang, H.; Ng, J.H.K. Googling for a diagnosis—Use of Google as a diagnostic aid: Internet based study. BMJ 2006, 333, 1143–1145. [Google Scholar] [CrossRef] [PubMed]
- Jalees, R. Accuracy of Medical Information on the Internet-Scientific American Blog Network. 2012. Available online: https://blogs.scientificamerican.com/guest-blog/accuracy-of-medical-information-on-the-internet/ (accessed on 27 July 2023).
- Lee, P.; Bubeck, S.; Petro, J. Benefits, Limits, and Risks of GPT-4 as an AI Chatbot for Medicine. N. Engl. J. Med. 2023, 388, 1233–1239. [Google Scholar] [CrossRef] [PubMed]
- Qiang, S.; Zhang, H.; Liao, Y.; Zhang, Y.; Gu, Y.; Wang, Y.; Xu, Z.; Shi, H.; Han, N.; Yu, H. Application of Large Language Models in Stroke Rehabilitation Health Education: 2-Phase Study. J. Med. Internet Res. 2025, 27, e73226. [Google Scholar] [PubMed]
- Drost, A.; Jaarsma, E.H.; Ring, D.; Azarpey, A. Factors Associated With an LLM Arriving at the Same Diagnosis as a Musculoskeletal Specialist. J. Am. Acad. Orthop. Surg. 2026. [Google Scholar] [CrossRef] [PubMed]
- Liu, R.; Liu, Q.; Hu, Q.; Nan, R.; He, J.; Yang, J.; Zhang, J.; Yang, G.; Yang, Z.; Xiao, X.; et al. Tests of large language models’ medical competence and application for clinical decision support of musculoskeletal rehabilitation. Front. Digit. Health 2025, 7, 1719340. [Google Scholar] [PubMed]
- Bosbach, W.A.; Montazeri, E.; Senge, J.F.; Beisbart, C.; Mitrakovic, M.; Anderson, S.E.; Divjak, E.; Ivanac, G.; Grieser, T.; Weber, M.-A.; et al. Consensus-Level and Cluster-Adjusted Evaluation of a Large Language Model for Diagnostic Extraction from Musculoskeletal Radiology Reports. Diagnostics 2026, 16, 1590. [Google Scholar] [CrossRef] [PubMed]
- Kaarre, J.; Feldt, R.; Keeling, L.E.; Dadoo, S.; Zsidai, B.; Hughes, J.D.; Samuelsson, K.; Musahl, V. Exploring the potential of ChatGPT as a supplementary tool for providing orthopaedic information. Knee Surg. Sports Traumatol. Arthrosc. 2023, 31, 5190–5198. [Google Scholar] [CrossRef] [PubMed]
- Ma, Z.; Liu, Y.; Zhang, Z.; Chen, R.; Fan, H.; Cao, X.; Ni, L. Clinical applications of large language models in knee osteoarthritis: A systematic review. Front. Med. 2025, 12, 1670824. [Google Scholar] [CrossRef] [PubMed]
- Quinn, M.; Milner, J.D.; Schmitt, P.; Morrissey, P.; Lemme, N.; Marcaccio, S.; DeFroda, S.; Tabaddor, R.; Owens, B.D. Artificial Intelligence Large Language Models Address Anterior Cruciate Ligament Reconstruction: Superior Clarity and Completeness by Gemini Compared With ChatGPT-4 in Response to American Academy of Orthopaedic Surgeons Clinical Practice Guidelines. Arthroscopy 2025, 41, 2002–2008. [Google Scholar] [CrossRef] [PubMed]
- Zhang, Y.; Huang, T.; Liu, C.; Miller, A.N.; Yang, M.; Harris, I.A.; Sawaguchi, T.; Miclau, T.; Tian, M.; Chui, C.S.; et al. Comparative evaluation of large language models for hip fracture-related patient questions: DeepSeek-V3-FW, Gemini 2.0 Flash, and ChatGPT-4.5. Digit. Health 2026, 12, 20552076251412989. [Google Scholar] [CrossRef] [PubMed]
- Vos, T.; Abajobir, A.A.; Abate, K.H.; Abbafati, C.; Abbas, K.M.; Abd-Allah, F.; Abdulkader, R.S.; Abdulle, A.M.; Abebo, T.A.; Abera, S.F.; et al. Global, regional, and national incidence, prevalence, and years lived with disability for 328 diseases and injuries for 195 countries, 1990–2016: A systematic analysis for the Global Burden of Disease Study 2016. Lancet 2017, 390, 1211–1259. [Google Scholar] [CrossRef] [PubMed]
- Smith, E.; Hoy, D.G.; Cross, M.; Vos, T.; Naghavi, M.; Buchbinder, R.; Woolf, A.D.; March, L. The global burden of other musculoskeletal disorders: Estimates from the Global Burden of Disease 2010 study. Ann. Rheum. Dis. 2014, 73, 1462–1469. [Google Scholar] [CrossRef] [PubMed]
- Jin, Y.; Guo, C.; Abbasian, M.; Abbasifard, M.; Abbott, J.H.; Abdullahi, A.; Abedi, A.; Abidi, H.; Abolhassani, H.; Abu-Gharbieh, E.; et al. Global pattern, trend, and cross-country inequality of early musculoskeletal disorders from 1990 to 2019, with projection from 2020 to 2050. Open Access 2024, 5, 943–962. [Google Scholar] [CrossRef] [PubMed]
- Liaghat, B.; Pedersen, J.R.; Husted, R.S.; Pedersen, L.L.; Thorborg, K.; Juhl, C.B. Diagnosis, prevention and treatment of common shoulder injuries in sport: Grading the evidence—A statement paper commissioned by the Danish Society of Sports Physical Therapy (DSSF). Br. J. Sports Med. 2023, 57, 408–416. [Google Scholar] [PubMed]
- Piercy, K.L.; Troiano, R.P.; Ballard, R.M.; Carlson, S.A.; Fulton, J.E.; Galuska, D.A.; George, S.M.; Olson, R.D. The Physical Activity Guidelines for Americans. JAMA 2018, 320, 2020. [Google Scholar] [CrossRef] [PubMed]
- Piercy, K.L.; Troiano, R.P. Physical Activity Guidelines for Americans From the US Department of Health and Human Services: Cardiovascular Benefits and Recommendations. Circ. Cardiovasc. Qual. Outcomes 2018, 11, e005263. [Google Scholar] [CrossRef] [PubMed]
- Kohn, M.D.; Sassoon, A.A.; Fernando, N.D. Classifications in Brief: Kellgren-Lawrence Classification of Osteoarthritis. Clin. Orthop. Relat. Res. 2016, 474, 1886–1893. [Google Scholar] [CrossRef] [PubMed]
- Mishra, V.; Sarraju, A.; Kalwani, N.M.; Dexter, J.P. Evaluation of Prompts to Simplify Cardiovascular Disease Information Generated Using a Large Language Model: Cross-Sectional Study. J. Med. Internet Res. 2024, 26, e55388. [Google Scholar] [CrossRef] [PubMed]
- Sivarajkumar, S.; Kelley, M.; Samolyk-Mazzanti, A.; Visweswaran, S.; Wang, Y. An Empirical Evaluation of Prompting Strategies for Large Language Models in Zero-Shot Clinical Natural Language Processing: Algorithm Development and Validation Study. JMIR Med. Inform. 2024, 12, e55318. [Google Scholar] [CrossRef] [PubMed]
- Ott, S.; Hebenstreit, K.; Liévin, V.; Hother, C.E.; Moradi, M.; Mayrhauser, M.; Praas, R.; Winther, O.; Samwald, M. ThoughtSource: A central hub for large language model reasoning data. Sci. Data 2023, 10, 528. [Google Scholar] [CrossRef] [PubMed]
- Charnock, D.; Shepperd, S.; Needham, G.; Gann, R. DISCERN: An instrument for judging the quality of written consumer health information on treatment choices. J. Epidemiol. Community Health 1999, 53, 105–111. [Google Scholar] [CrossRef] [PubMed]
- Sun, F.; Yang, F.; Zheng, S. Evaluation of the Liver Disease Information in Baidu Encyclopedia and Wikipedia: Longitudinal Study. J. Med. Internet Res. 2021, 23, e17680. [Google Scholar] [CrossRef] [PubMed]
- Chiang, C.-H.; Lee, H. Can Large Language Models Be an Alternative to Human Evaluation? In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Association for Computational Linguistics: Stroudsburg, PA, USA, 2023. [Google Scholar]
- Koo, T.K.; Li, M.Y. A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research. J. Chiropr. Med. 2016, 15, 155–163. [Google Scholar] [CrossRef]
- Christmann, A.; Van Aelst, S. Robust estimation of Cronbach’s alpha. J. Multivar. Anal. 2006, 97, 1660–1674. [Google Scholar] [CrossRef]
- Cronbach, L.J. Coefficient alpha and the internal structure of tests. Psychometrika 1951, 16, 297–334. [Google Scholar] [CrossRef]
- Barnett, T.; Hoang, H.; Furlan, A. An analysis of the readability characteristics of oral health information literature available to the public in Tasmania, Australia. BMC Oral Health 2016, 16, 35. [Google Scholar] [CrossRef] [PubMed]
- Flesch, R. A new readability yardstick. J. Appl. Psychol. 1948, 32, 221–233. [Google Scholar] [CrossRef] [PubMed]
- Kincaid, J.P.; Fishburne, R.P., Jr.; Rogers, R.L.; Chissom, B.S. Derivation of New Readability Formulas (Automated Readability Index, Fog Count and Flesch Reading Ease Formula) for Navy Enlisted Personnel; Defense Technical Information Center: Fort Belvoir, VA, USA, 1975; Available online: https://journals.sagepub.com/doi/abs/10.1177/001872087001200505 (accessed on 9 September 2023).
- Gunning, R. The Technique of Clear Writing; McGraw-Hill: New York, NY, USA, 1952. [Google Scholar]
- Mc Laughlin, G.H. SMOG grading-a new readability formula. J. Read. 1969, 12, 639–646. [Google Scholar]
- Coleman, M.; Liau, T.L. A computer readability formula designed for machine scoring. J. Appl. Psychol. 1975, 60, 283. [Google Scholar] [CrossRef]
- Smith, E.A.; Kincaid, J.P. Derivation and Validation of the Automated Readability Index for Use with Technical Materials. Hum. Factors 1970, 12, 457–564. [Google Scholar] [CrossRef]
- Edmunds, M.R.; Barry, R.J.; Denniston, A.K. Readability Assessment of Online Ophthalmic Patient Information. JAMA Ophthalmol. 2013, 131, 1610. [Google Scholar] [CrossRef] [PubMed]
- Alkharusi, H. A descriptive analysis and interpretation of data from Likert scales in educational and psychological research. Indian J. Psychol. Educ. 2022, 12, 13–16. [Google Scholar]
- Haluza, D.; Naszay, M.; Stockinger, A.; Jungwirth, D. Digital Natives Versus Digital Immigrants: Influence of Online Health Information Seeking on the Doctor–Patient Relationship. Health Commun. 2017, 32, 1342–1349. [Google Scholar] [CrossRef] [PubMed]
- Journal of Medical Internet Research-Quality and Reliability of Liver Cancer–Related Short Chinese Videos on TikTok and Bilibili: Cross-Sectional Content Analysis Study. Available online: https://www.jmir.org/2023/1/e47210 (accessed on 8 November 2024).
- Kunze, K.N. Editorial Commentary: YouTube Videos Provide Poor-Quality Medical Information: Don’t Believe What You Watch! Arthroscopy 2020, 36, 3048–3049. [Google Scholar] [CrossRef] [PubMed]
- Mastrokostas, P.G.; Mastrokostas, L.E.; Emara, A.K.; Wellington, I.J.; Ginalis, E.; Houten, J.K.; Khalsa, A.S.; Saleh, A.; Razi, A.E.; Ng, M.K. GPT-4 as a Source of Patient Information for Anterior Cervical Discectomy and Fusion: A Comparative Analysis Against Google Web Search. Glob. Spine J. 2024, 14, 2389–2398. [Google Scholar] [CrossRef] [PubMed]
- Chung, P.; Fong, C.T.; Walters, A.M.; Aghaeepour, N.; Yetisgen, M.; O’reilly-Shah, V.N. Large Language Model Capabilities in Perioperative Risk Prediction and Prognostication. JAMA Surg. 2024, 159, 928–937. [Google Scholar] [CrossRef] [PubMed]
- Fink, A.; Rau, A.; Reisert, M.; Bamberg, F.; Russe, M.F. Retrieval-Augmented Generation with Large Language Models in Radiology: From Theory to Practice. Radiol. Artif. Intell. 2025, 7, e240790. [Google Scholar] [CrossRef] [PubMed]
- Blanchard, C.G.; Labrecque, M.S.; Ruckdeschel, J.C.; Blanchard, E.B. Information and decision-making preferences of hospitalized adult cancer patients. Soc. Sci. Med. 1988, 27, 1139–1145. [Google Scholar] [CrossRef] [PubMed]
- Greenfield, S. Expanding Patient Involvement in Care: Effects on Patient Outcomes. Ann. Intern. Med. 1985, 102, 520. [Google Scholar] [CrossRef] [PubMed]
- Kaba, R.; Sooriakumaran, P. The evolution of the doctor-patient relationship. Int. J. Surg. 2007, 5, 57–65. [Google Scholar] [CrossRef] [PubMed]
- Oh, Y.; Park, S.; Byun, H.K.; Cho, Y.; Lee, I.J.; Kim, J.S.; Ye, J.C. LLM-driven multimodal target volume contouring in radiation oncology. Nat. Commun. 2024, 15, 9186. [Google Scholar] [CrossRef] [PubMed]
- Zhou, J.; He, X.; Sun, L.; Xu, J.; Chen, X.; Chu, Y.; Zhou, L.; Liao, X.; Zhang, B.; Afvari, S.; et al. Pre-trained multimodal large language model enhances dermatological diagnosis using SkinGPT-4. Nat. Commun. 2024, 15, 5649. [Google Scholar] [CrossRef] [PubMed]
- Haver, H.L.; Lin, C.T.; Sirajuddin, A.; Yi, P.H.; Jeudy, J. Use of ChatGPT, GPT-4, and Bard to Improve Readability of ChatGPT’s Answers to Common Questions About Lung Cancer and Lung Cancer Screening. AJR Am. J. Roentgenol. 2023, 221, 701–704. [Google Scholar] [CrossRef] [PubMed]
- Doshi, R.; Amin, K.S.; Khosla, P.; Bajaj, S.; Chheang, S.; Forman, H.P. Quantitative Evaluation of Large Language Models to Streamline Radiology Report Impressions: A Multimodal Retrospective Analysis. Radiology 2024, 310, e231593. [Google Scholar] [CrossRef] [PubMed]
- Monteith, S.; Glenn, T.; Geddes, J.R.; Whybrow, P.C.; Achtyes, E.; Bauer, M. Artificial intelligence and increasing misinformation. Br. J. Psychiatry 2024, 224, 33–35. [Google Scholar] [CrossRef] [PubMed]
- Boyle, C.M. Difference Between Patients’ and Doctors’ Interpretation of Some Common Medical Terms. Br. Med. J. 1970, 2, 286–289. [Google Scholar] [CrossRef] [PubMed]
- Naddaf, M. ChatGPT generates fake data set to support scientific hypothesis. Nature 2023, 623, 895–896. [Google Scholar] [CrossRef] [PubMed]
- Yalamanchili, A.; Sengupta, B.; Song, J.; Lim, S.; Thomas, T.O.; Mittal, B.B.; Abazeed, M.E.; Teo, P.T. Quality of Large Language Model Responses to Radiation Oncology Patient Care Questions. JAMA Netw. Open 2024, 7, e244630. [Google Scholar] [CrossRef] [PubMed]





| ID | DISCERN Question |
|---|---|
| Q1 | Are the aims clear? |
| Q2 | Does it achieve its aims? |
| Q3 | Is it relevant? |
| Q4 | Is it clear what sources of information were used to compile the publication (other than the author or producer)? |
| Q5 | Is it clear when the information used or reported in the publication was produced? |
| Q6 | Is it balanced and unbiased? |
| Q7 | Does it provide details of additional sources of support and information? |
| Q8 | Does it refer to areas of uncertainty? |
| Q9 | Does it describe how each treatment works? |
| Q10 | Does it describe the benefits of each treatment? |
| Q11 | Does it describe the risks of each treatment? |
| Q12 | Does it describe what would happen if no treatment is used? |
| Q13 | Does it describe how the treatment choices affect overall quality of life? |
| Q14 | Is it clear that there may be more than one possible treatment choice? |
| Q15 | Does it provide support for shared decision-making? |
| Q16 | Based on the answers to all of the above questions, rate the overall quality of the publication as a source of information about treatment choices |
| Tool Name | Calculating Formula |
|---|---|
| FKGL | |
| FRES | |
| SMOG | |
| GF | |
| ARI | |
| CLI |
| FRES Readability Scale Score | FKGL, SMOF, CLI, GF or ARI Readability Scale Score | Age | USDHHS Reading Level |
|---|---|---|---|
| 90–100 | 0–1 | 6–7 yr old | Very Easy |
| 2–3 | 7–8 yr old | ||
| 3–4 | 8–9 yr old | ||
| 4–5 | 9–10 yr old | ||
| 5–6 | 10–11 yr old | ||
| 80–89 * | 6–7 * | 11–12 yr old * | Easy * |
| 70–79 | 7–8 | 12–13 yr old | Fairly Easy |
| 60–69 | 8–9 | 13–14 yr old | Average |
| 9–10 | 14–15 yr old | ||
| 50–59 | 10–11 | 15–16 yr old | Fairly Difficult |
| 11–12 | 16–17 yr old | ||
| 12–13 | 17–18 yr old | ||
| 30–49 | 13–17 | 18–22 yr old | Difficult |
| 0–29 | 17+ | 22+ yr old | Very difficult |
| Rater Identity | Mean | Standard Deviation | Range |
|---|---|---|---|
| Doctor-1 | 55.60 | 8.4 | 34 to 68 |
| Doctor-2 | 51.95 | 9.62 | 33 to 67 |
| Doctor-3 | 40.89 | 6.88 | 26 to 56 |
| Therapist-1 | 52.87 | 8.99 | 34 to 68 |
| Therapist-2 | 51.19 | 10.23 | 30 to 67 |
| GPT-4o | 55.82 | 5.16 | 43 to 67 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Fu, Y.; Li, H.; You, M.; Wang, L.; Liu, W.; Zhou, K.; Wang, L.; Chen, X.; Chen, G. Evaluation of Large Language Models in Generating Physical Exercise Rehabilitation Programs for Musculoskeletal Disorders Across Multiple Clinical Scenarios. Healthcare 2026, 14, 2389. https://doi.org/10.3390/healthcare14152389
Fu Y, Li H, You M, Wang L, Liu W, Zhou K, Wang L, Chen X, Chen G. Evaluation of Large Language Models in Generating Physical Exercise Rehabilitation Programs for Musculoskeletal Disorders Across Multiple Clinical Scenarios. Healthcare. 2026; 14(15):2389. https://doi.org/10.3390/healthcare14152389
Chicago/Turabian StyleFu, Yu, Hairui Li, Mingke You, Li Wang, Weizhi Liu, Kai Zhou, Lingcheng Wang, Xi Chen, and Gang Chen. 2026. "Evaluation of Large Language Models in Generating Physical Exercise Rehabilitation Programs for Musculoskeletal Disorders Across Multiple Clinical Scenarios" Healthcare 14, no. 15: 2389. https://doi.org/10.3390/healthcare14152389
APA StyleFu, Y., Li, H., You, M., Wang, L., Liu, W., Zhou, K., Wang, L., Chen, X., & Chen, G. (2026). Evaluation of Large Language Models in Generating Physical Exercise Rehabilitation Programs for Musculoskeletal Disorders Across Multiple Clinical Scenarios. Healthcare, 14(15), 2389. https://doi.org/10.3390/healthcare14152389
