Benchmarking Large Language Model Responses Against Surgical Clinical Practice Guidelines for Chronic Rhinosinusitis: The Importance of User Prompts
Abstract
1. Introduction
2. Materials and Methods
2.1. Data Collection
2.2. AI Response Generation
2.3. Physician Assessment
2.4. Readability Assessment
2.5. Statistical Analysis
3. Results
4. Discussion
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| CRS | Chronic Rhinosinusitis |
| LLMs | Large Language Models |
| FRE | Flesch Reading Ease |
| FKGL | Flesch-Kincaid Grade Level |
| IRR | Inter-rater Reliability |
| ICC | Intraclass Correlation Coefficient |
| AAO-HNS | American Academy of Otolaryngology—Head and Neck Surgery |
| AAO-HNSF | American Academy of Otolaryngology–Head and Neck Surgery Foundation |
| AMA | American Medical Association |
Appendix A
| Single Patient-Focused Prompt | |
|---|---|
| 1 | I’ve been dealing with sinus problems for a long time, and my doctor mentioned surgery might help. Can you explain what chronic rhinosinusitis is, how it’s diagnosed, when and what surgery might be needed, the possible benefits and risks, what recovery is like, and how I can take care of myself afterward? |
| Guideline Mapped Prompt | |
| 1 | What is chronic rhinosinusitis (CRS)? |
| 2 | How is the diagnosis of CRS confirmed? |
| 3 | What is sinus surgery? |
| 4 | How do I know if I need sinus surgery? |
| 5 | What are the benefits of verifying my diagnosis before surgery? |
| 6 | What happens during the assessment for surgery? |
| 7 | What are the risks of verifying my diagnosis or assessing my candidacy for surgery? |
| 8 | What makes me a good surgical candidate? |
| 9 | Can I play a role in deciding if surgery is right for me? |
| 10 | Are antibiotics helpful or not helpful for my chronic sinus disease? |
| 11 | Will surgery replace my need for most sinonasal medications? |
| 12 | Are there alternative treatments to surgery? |
| 13 | What is the difference between sinus dilation surgery and sinus surgery creating wide sinus openings? |
| 14 | How does creating wide sinus openings help my chronic sinus disease? |
| 15 | What should I expect after surgery? Should I expect pain? |
| 16 | What type of postoperative care and/or medications will I need, and for how long? |
| 17 | How many postoperative visits will I have, and what is the timing for these? |
| 18 | What will be covered or done at these postoperative visits? |
| 19 | What are the limitations after surgery? |
| 20 | How does the extent of surgery impact my healing after surgery? |
| 21 | Will my physician provide any resources about sinus surgery for me? |
| Questions | Rubric Scores |
|---|---|
| What is chronic rhinosinusitis (CRS?) | Signs and Symptoms (2 points): 0: No mention or incorrect explanation of chronic rhinosinusitis (CRS) symptoms or diagnostic timeframe. 1: Partial or vague description (e.g., mentions sinus congestion or drainage but omits duration, inflammation, or multiple symptom criteria). 2: Comprehensive and accurate explanation consistent with CPG—defines CRS as ≥12 weeks of ≥2 symptoms (nasal obstruction/congestion, discolored drainage, facial pain/pressure, or reduced sense of smell) |
| How is the diagnosis of CRS confirmed? | Diagnosis Confirmation 0: Missing or inaccurate discussion of how CRS is confirmed. 1: Mentions evaluation or imaging without specificity. 2: Clearly describes confirmation by objective evidence of sinonasal inflammation (nasal endoscopy or CT findings) and distinguishes CRS from allergy, acute sinusitis, or ear-related symptoms. |
| What is sinus surgery? | 0 points: No or incorrect explanation of sinus surgery. 1 point: Accurately explains endoscopic sinus surgery (ESS) as a telescope-guided, minimally invasive procedure performed through the nostrils to open blocked sinus pathways, clear infection or inflammation, and improve medication delivery. 2 points: Provides a comprehensive explanation including purpose (to improve breathing and reduce infections), benefits (better sinus drainage, reduced symptoms), and sets realistic expectations about recovery and outcomes. |
| How do I know if I need sinus surgery? | Need for Sinus Surgery (2 points): 0: No or inappropriate explanation of indications for surgery. 1: Mentions failure of medication or need for imaging without details. 2: Explains that surgery is considered when medical therapy (e.g., saline irrigation, intranasal/oral steroids, antibiotics in select cases) fails; identifies subtypes benefiting most (CRSwNP, AFRS, EMRS, fungal disease, bony obstruction); and emphasizes individualized timing and absence of a one-size-fits-all prerequisite regimen. |
| What are the benefits of verifying my diagnosis before surgery? | Benefits of Verifying Diagnosis Before Surgery (2 points): 0: No discussion of benefits. 1: Mentions excluding other diagnoses (migraines, asthma, etc) to guide treatment plan 2: Describes surgical planning benefit and assessing the success of surgery with a combination of symptoms and objective evidence |
| What happens during the assessment for surgery? | How to Verify Diagnosis Before Surgery (2 points): 0: No pre-operative assessment guidance. 1: Mentions CT or general evaluation only. 2: Details comprehensive assessment: symptom severity, quality-of-life impact, endoscopic findings, radiographic imaging, subtype identification (CRSwNP, EMRS, AFRS), and prior medical response. |
| What are the risks of verifying my diagnosis or assessing my candidacy for surgery? | Risks of Verifying Diagnosis Before Surgery (2 points): 0: No discussion of need to check, review, or repeat certain tests 1: Explains need to check, review, or repeat certain tests without mention of specifics. 2: Explains additional costs and minor risks of nasal endoscopy or imaging |
| What makes me a good surgical candidate? | Good Surgical Candidate Criteria (2 points): 0: No mention or incorrect explanation of who qualifies for sinus surgery. 1: Partial or vague criteria (e.g., mentions “failed medical therapy” without specifying what that means or which subtypes benefit most). 2: Comprehensive and guideline-consistent description—identifies candidates as adults with CRS who meet diagnostic criteria, have persistent symptoms or impaired quality of life despite appropriate medical therapy, and whose subtype or disease severity suggests surgery would provide greater benefit than continued medical management. |
| Can I play a role in deciding if surgery is right for me? | 0: No mention of patient input 1: Partial mention of patient input 2: Mentions input of patient is vital, asked about CRS impact on patient life, talks about doctor sharing process of decision making |
| Are antibiotics helpful or not helpful for my chronic sinus disease? | 0: No mention of abx use 1: Mentions antibiotic use but not specifically when indicated 2: Refers to pathophysiology of CRS and why abx may not be indicated, discusses that abx for bacterial infection and CRS is not necessarily triggered by bacteria. Refers to CRS time course, relates side effects of antibiotics and dangers of overuse, reviews signs of active infection and clinician performing exam to determine |
| Will surgery replace my need for most sinonasal medications? | 0: Does not answer question 1: Mentions that it is part of treatment 2: Mentions that it is not a replacement for medications and part of treatment and allows topicals to work better |
| Are there alternative treatments to surgery? | 0: Describes no alternatives 1: Describes some alternative treatments 2: Describes alternatives and when surgery is considered |
| What is the difference between sinus dilation surgery and sinus surgery creating wide sinus openings? | 0: Does not define the differences between surgeries 1: Defines only one of the surgeries 2: Defines both surgeries |
| How does creating wide sinus openings help my chronic sinus disease? | 0: Refers to only 1 way 1: Refers to only 2–3 out of the 4 ways 2: Describes all 4 benefits |
| What should I expect after surgery? Should I expect pain? | 0: No postoperative expectations 1: Gives only some expectations and some suggestions for dealing with pain 2: Gives full expectations and describes level of pain, recommends possible pain meds and when narcotics indicated, describes some symptoms that may occur and when to resume nasal irrigations and sprays |
| What type of postoperative care and/or medications will I need, and for how long? | 0: Doesn’t describe any care or meds 1: Describes only meds or describes only timeline 2: Gives nasal care and timeline |
| How many postoperative visits will I have, and what is the timing for these? | 0: Doesn’t answer how many visits 1: Does not give time course of how regular the visits will be 2: Conveys patients will be seen regularly and at specific intervals |
| What will be covered or done at these postoperative visits? | 0: Does not review what will happen at postop visits 1: Mentions that visits may involve nasal endoscopy and debridement 2: Mentions that visits may involve nasal endoscopy and debridement and defines, talks about taking OTC pain med before appointments and assessing symptoms and surgical recovery, modify meds, perform physical exam |
| What are the limitations after surgery? | 0: No specifics on limitations after surgery 1: Give some examples of how long maybe be out of work or when to stop taking stronger pain med 2: Also avoids heavy physical activity, and avoiding sneezing, and not going underwater |
| How does the extent of surgery impact my healing after surgery? | 0: Does not explain extent of surgery impact 1: Goes into some detail of process of healing 2: Describes process of healing and what may take longer vs. shorter |
| Will my physician provide any resources about sinus surgery for me? | 0: Provides no resources 1: Provides some resources 2: Provides educational materials in paper/electronic format, including restrictions, important meds, signs and symptoms for urgent eval during post op |
References
- Fokkens, W.J.; Lund, V.J.; Hopkins, C.; Hellings, P.W.; Kern, R.; Reitsma, S.; Toppila-Salmi, S.; Bernal-Sprekelsen, M.; Mullol, J.; Alobid, I.; et al. European Position Paper on Rhinosinusitis and Nasal Polyps 2020. Rhinology 2020, 58, 1–464. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ariyaratne, S.; Jenko, N.; Mark Davies, A.; Iyengar, K.P.; Botchu, R. Could ChatGPT Pass the UK Radiology Fellowship Examinations? Acad. Radiol. 2024, 31, 2178–2182. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Siam, M.K.; Faruk, M.J.H.; Cheng, J.Q.; Gu, H. Fusion-Augmented Large Language Models: Boosting Diagnostic Trustworthiness via Model Consensus. In Proceedings of the 2025 IEEE EMBS International Conference on Biomedical and Health Informatics (BHI); IEEE: New York, NY, USA, 2025; pp. 1–7. [Google Scholar] [CrossRef] [Scilit]
- Hirani, H.; Saboo, B.; Modi, A.; Modi, P.; Samajdar, S.S.; Saboo, B.; Chawla, M.; Maheshwari, A.; Gupta, A.; Parikh, R.; et al. Clinical Assessment of Large Language Models: A Comprehensive Multi-domain Performance Study for Healthcare Applications. Int. J. Diabetes Technol. 2025, 4, 159–165. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Moise, A.; Tatar, L.; Sela, N.; da Silva, S.D.; Kouz, J.; Tamilia, M.; Hier, M.P.; Forest, V.-I.; Payne, R.J. Thyroid Nodule Experts Evaluating ChatGPT’s Assessment of Thyroid Nodules Classified by the Bethesda System for Reporting Thyroid Cytopathology. J. Otolaryngol.-Head. Neck Surg. 2025, 54, 19160216251387617. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Abdul Sami, M.; Abdul Samad, M.; Parekh, K.; Suthar, P.P. Comparative Accuracy of ChatGPT 4.0 and Google Gemini in Answering Pediatric Radiology Text-Based Questions. Cureus 2024, 16, e70897. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Draelos, R.L.; Afreen, S.; Blasko, B.; Brazile, T.L.; Chase, N.; Desai, D.P.; Evert, J.; Gardner, H.L.; Herrmann, L.; House, A.V.; et al. Large language models provide unsafe answers to patient-posed medical questions. arXiv 2025, arXiv:2507.18905. [Google Scholar] [CrossRef] [Scilit]
- US Patients are Turning to ChatGPT to Navigate Healthcare. Available online: https://www.fiercehealthcare.com/ai-and-machine-learning/40m-people-use-chatgpt-answer-healthcare-questions-openai-says (accessed on 17 February 2026).
- Zhang, D.; Earp, B.E.; Kilgallen, E.E.; Blazar, P. Readability of Online Hand Surgery Patient Educational Materials: Evaluating the Trend Since 2008. J. Hand Surg. Am. 2022, 47, 186.e1–186.e8. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Shin, J.J.; Wilson, M.; McKenna, M.; Rosenfeld, R.; Ammon, K.; Crosby, D.; Fuchs, J.M.; Hensler, J.B.; Illing, E.A.; Lam, K.; et al. Clinical Practice Guideline: Surgical Management of Chronic Rhinosinusitis. Otolaryngol. Head Neck Surg. 2025, 172, S1–S47. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, L.-W.; Miller, M.J.; Schmitt, M.R.; Wen, F.K. Assessing readability formula differences with written health information materials: Application, results, and recommendations. Res. Soc. Adm. Pharm. 2013, 9, 503–516. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Eltorai, A.E.; Ghanian, S.; Adams, C.A., Jr.; Born, C.T.; Daniels, A.H. Readability of patient education materials on the american association for surgery of trauma website. Arch. Trauma Res. 2014, 3, e18161. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
- Carl, N.; Haggenmüller, S.; Wies, C.; Nguyen, L.; Winterstein, J.T.; Hetz, M.J.; Mangold, M.H.; Hartung, F.O.; Grüne, B.; Holland-Letz, T.; et al. Evaluating interactions of patients with large language models for medical information. BJU Int. 2025, 135, 1010–1017. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Bean, A.M.; Payne, R.E.; Parsons, G.; Kirk, H.R.; Ciro, J.; Mosquera-Gómez, R.; M, S.H.; Ekanayaka, A.S.; Tarassenko, L.; Rocher, L.; et al. Reliability of LLMs as medical assistants for the general public: A randomized preregistered study. Nat. Med. 2026, 32, 609–615. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Geracitano, J.; Anderson, B.; Coffel, M.; Rosenzweig, M.; Dorn, S.D.; Khairat, S.; Conklin, J. The Accuracy of ChatGPT in Answering FAQs, Making Clinical Recommendations, and Categorizing Patient Symptoms: A Literature Review. Adv. Health Inf. Sci. Pract. 2025, 1, VXUL2925. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Megafu, M.; Guerrero, O.; Yendluri, A.; Parsons, B.O.; Galatz, L.M.; Li, X.; Kelly, J.D.; Parisien, R.L. ChatGPT and Gemini Are Not Consistently Concordant with the 2020 American Academy of Orthopaedic Surgeons Clinical Practice Guidelines When Evaluating Rotator Cuff Injury. Arthroscopy 2025, 41, 2753–2757. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhang, Y.; Yang, D.; Shi, Y.; Liu, Y. Performance of Large Language Models in Lung Cancer Clinical Decision-Making: A Comparative Analysis Based on DeepSeek, Grok, and GPT. Cureus 2025, 17, e99026. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Alarifi, M. Appropriateness of Thyroid Nodule Cancer Risk Assessment and Management Recommendations Provided by Large Language Models. J. Imaging Inform. Med. 2025, 38, 4324–4335. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Aliyeva, A.; Sari, E.; Alaskarov, E.; Nasirov, R. Enhancing Postoperative Cochlear Implant Care with ChatGPT-4: A Study on Artificial Intelligence (AI)-Assisted Patient Education and Support. Cureus 2024, 16, e53897. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- McLaughlin, N.D.; Srinivas, A.N.; Lowe, Z.F.; Botterbush, K.S.; Patel, M.S.; Avila, M.J. Large Language Model Hallucinations in Spine Surgery: A Comparative Analysis of Clinician vs Patient-Level Prompts. Neurosurg. Pract. 2026, 7, e000244. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chelli, M.; Descamps, J.; Lavoué, V.; Trojani, C.; Azar, M.; Deckert, M.; Raynier, J.-L.; Clowez, G.; Boileau, P.; Ruetsch-Chelli, C. Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis. J. Med. Internet Res. 2024, 26, e53164. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Shah, H.A.; Househ, M. Chain of Thought Strategy for Smaller LLMs for Medical Reasoning. Stud. Health Technol. Inform. 2025, 327, 783–787. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chen, G.; Lin, C.; Kim, J.H.; Du, F.; Luo, Z.; Shin, Y.S.; Li, X. Readability and information quality of LLM-Generated HIV content: A methodological content evaluation. BMC Infect. Dis. 2025, 25, 1624. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Mudrik, A.; Nadkarni, G.N.; Efros, O.; Soffer, S.; Klang, E. Prompt engineering in large language models for patient education: A systematic review. medRxiv 2025. [Google Scholar] [CrossRef] [Scilit]
- Lautrup, A.D.; Hyrup, T.; Schneider-Kamp, A.; Dahl, M.; Lindholt, J.S.; Schneider-Kamp, P. Heart-to-heart with ChatGPT: The impact of patients consulting AI for cardiovascular health advice. Open Heart 2023, 10, e002455. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhou, M.; Pan, Y.; Zhang, Y.; Song, X.; Zhou, Y. Evaluating AI-generated patient education materials for spinal surgeries: Comparative analysis of readability and DISCERN quality across ChatGPT and DeepSeek models. Int. J. Med. Inform. 2025, 198, 105871. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Marcaccini, G.; Seth, I.; Xie, Y.; Susini, P.; Pozzi, M.; Cuomo, R.; Rozen, W.M. Breaking bones, breaking barriers: ChatGPT, DeepSeek, and Gemini in hand fracture management. J. Clin. Med. 2025, 14, 1983. [Google Scholar] [CrossRef] [Scilit] [PubMed]



Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Lad, H.; Kwon, E.; Chadha, A.; Haimowitz, S.Z.; Gold, B.S.; Kaye, R.; Hsueh, W.D. Benchmarking Large Language Model Responses Against Surgical Clinical Practice Guidelines for Chronic Rhinosinusitis: The Importance of User Prompts. J. Otorhinolaryngol. Hear. Balanc. Med. 2026, 7, 28. https://doi.org/10.3390/ohbm7020028
Lad H, Kwon E, Chadha A, Haimowitz SZ, Gold BS, Kaye R, Hsueh WD. Benchmarking Large Language Model Responses Against Surgical Clinical Practice Guidelines for Chronic Rhinosinusitis: The Importance of User Prompts. Journal of Otorhinolaryngology, Hearing and Balance Medicine. 2026; 7(2):28. https://doi.org/10.3390/ohbm7020028
Chicago/Turabian StyleLad, Hetal, Emily Kwon, Ayushi Chadha, Sean Z. Haimowitz, Brandon S. Gold, Rachel Kaye, and Wayne D. Hsueh. 2026. "Benchmarking Large Language Model Responses Against Surgical Clinical Practice Guidelines for Chronic Rhinosinusitis: The Importance of User Prompts" Journal of Otorhinolaryngology, Hearing and Balance Medicine 7, no. 2: 28. https://doi.org/10.3390/ohbm7020028
APA StyleLad, H., Kwon, E., Chadha, A., Haimowitz, S. Z., Gold, B. S., Kaye, R., & Hsueh, W. D. (2026). Benchmarking Large Language Model Responses Against Surgical Clinical Practice Guidelines for Chronic Rhinosinusitis: The Importance of User Prompts. Journal of Otorhinolaryngology, Hearing and Balance Medicine, 7(2), 28. https://doi.org/10.3390/ohbm7020028

