Modification and Validation of the System Causability Scale Using AI-Based Therapeutic Recommendations for Urological Cancer Patients: A Basis for the Development of a Prospective Comparative Study
Abstract
1. Introduction
2. Materials and Methods
2.1. Planned Prospective Trial
2.2. Selection of Two Appropriate LLMs for Comparison with the MTB (1)
2.3. Development of Standardized Prompts for Data Input on GUC Patients and the Creation of a Uniform Recommendation Matrix for Both LLMs and the MTB to Facilitate Blinded Assessment (2)
2.4. Modification and Validation of the Newly Developed mSCS Using a Cohort of 40 Patients with Varying Organ-Specific GUCs (3)
2.5. Biometric Sample Size Planning for the Prospective Trial, Preceded by a Moderated Delphi Process with the Entire Study Team to Establish What Level of Difference in the mSCS, Derived from Preliminary Study Results, Would Still Be Considered Non-Inferior for LLMs Compared to the MTB (4)
2.6. Precise Documentation and Listing of Statistical Methods to Validate the mSCS and Compare Results Between the Groups (MTB vs. LLM) (5)
3. Results
3.1. Selection of Two Appropriate LLMs for Comparison with the MTB (1)
3.2. Development of Standardized Prompts for Data Input on Urological Tumor Patients and the Creation of a Uniform Recommendation Matrix for Both LLMs and the MTB to Facilitate Blinded Assessment (2)
- (1)
- Preferred therapy recommendation (if available);
- (2)
- Therapy alternatives;
- (3)
- Justification of the recommendations;
- (4)
- Supportive measures/supplementary therapies;
- (5)
- Further information/explanations.
3.3. Modification and Validation of the Newly Developed mSCS Using a Cohort of 40 GUC Patients with Varying Organ-Specific Cancers (3)
| Item | SCS | mSCS |
|---|---|---|
| 1 | I found that the recommendation included all relevant known causal factors with sufficient precision and granularity. | I found that the recommendation included all relevant patient-specific factors (individual patient data such as individual tumor stages, previous treatments, and specific health conditions) with sufficient precision and granularity. |
| 2 | I understood the explanations within the context of my work. | I found the quality and representativeness of the recommendations, particularly in relation to oncological scenarios, sufficient. |
| 3 | I could change the level of detail on demand. | I found that all reasonable treatment alternatives were specified. |
| 4 | I did not need support to understand the explanations. | I did not need support to understand the explanations. |
| 5 | I found the explanations helped me to understand causality | I found that the recommendation was explained and made transparent. |
| 6 | I was able to use the explanations with my knowledge base. | I found the recommendation to be consistent with current clinical guidelines. |
| 7 | I did not find inconsistencies between explanations. | I did not find inconsistencies between explanations/recommendations. |
| 8 | I think that most people would learn to understand the explanations very quickly. | I think that most healthcare professionals would learn to understand the explanations very quickly. |
| 9 | I did not need more references in the explanations, e.g., medical guidelines and regulations. | I found the recommendation demonstrates access to the latest research and clinical guidelines. |
| 10 | I received the explanations in a timely and efficient manner | I found the quality of interaction (ease of use and accessibility) sufficient. |
3.4. Biometric Sample Size Planning for the Prospective Trial, Preceded by a Moderated Delphi Process with the Entire Study Team to Establish What Level of Difference in the mSCS, Derived from Preliminary Study Results, Would Still Be Considered Non-Inferior for LLMs Compared to the MTB (4)
3.5. Validation of the mSCS and Comparison Between the Groups (MTB vs. LLM) (5)
3.5.1. Interrater Reliability
3.5.2. Agreement Between SCS and mSCS
3.5.3. Internal Consistency
3.5.4. Evaluation of Clinical Applicability of the mSCS Compared to the SCS
4. Discussion
Limitations
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Dave, T.; Athaluri, S.A.; Singh, S. ChatGPT in medicine: An overview of its applications, advantages, limitations, future prospects, and ethical considerations. Front. Artif. Intell. 2023, 6, 1169595. [Google Scholar] [CrossRef] [Scilit]
- Ray, P.P. ChatGPT: A comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope. Internet Things Cyber-Phys. Syst. 2023, 3, 121–154. [Google Scholar] [CrossRef] [Scilit]
- Rajpurkar, P.; Chen, E.; Banerjee, O.; Topol, E.J. AI in health and medicine. Nat. Med. 2022, 28, 31–38. [Google Scholar] [CrossRef] [Scilit]
- Thirunavukarasu, A.J.; Ting, D.S.J.; Elangovan, K.; Gutierrez, L.; Tan, T.F.; Ting, D.S.W. Large language models in medicine. Nat. Med. 2023, 29, 1930–1940. [Google Scholar] [CrossRef] [Scilit]
- Kowalewski, K.-F.; Rodler, S. Large Language Models in der Wissenschaft. [Large language models in science]. Die Urol. 2024, 63, 860–866. [Google Scholar] [CrossRef] [Scilit]
- OpenAI. Introducing ChatGPT. 30 November 2022. Available online: https://openai.com/blog/chatgpt (accessed on 22 September 2024).
- Eppler, M.; Ganjavi, C.; Ramacciotti, L.S.; Piazza, P.; Rodler, S.; Checcucci, E.; Rivas, J.G.; Kowalewski, K.F.; Belenchón, I.R.; Puliatti, S.; et al. Awareness and Use of ChatGPT and Large Language Models: A Prospective Cross-sectional Global Survey in Urology. Eur. Urol. 2024, 85, 146–153. [Google Scholar] [CrossRef] [Scilit]
- Pillay, B.; Wootten, A.C.; Crowe, H.; Corcoran, N.; Tran, B.; Bowden, P.; Crowe, J.; Costello, A.J. The impact of multidisciplinary team meetings on patient assessment, management and outcomes in oncology settings: A systematic review of the literature. Cancer Treat. Rev. 2016, 42, 56–72. [Google Scholar] [CrossRef] [Scilit]
- Taylor, C.; Munro, A.J.; Glynne-Jones, R.; Griffith, C.; Trevatt, P.; Richards, M.; Ramirez, A.J. Multidisciplinary team working in cancer: What is the evidence? BMJ 2010, 340, c951. [Google Scholar] [CrossRef] [Scilit]
- Perez-Gracia, J.L.; Awada, A.; Calvo, E.; Amaral, T.; Arkenau, H.-T.; Gruenwald, V.; Bodoky, G.; Lolkema, M.P.; Nicola, M.D.; Penel, N.; et al. ESMO Clinical Research Observatory (ECRO): Improving the efficiency of clinical research through rationalisation of bureaucracy. ESMO Open 2020, 5, e000662. [Google Scholar] [CrossRef] [Scilit]
- Levin, G.; Gotlieb, W.; Ramirez, P.; Meyer, R.; Brezinov, Y. ChatGPT in a gynaecologic oncology multidisciplinary team tumour board: A feasibility study. BJOG Int. J. Obstet. Gynaecol. 2024. [Google Scholar] [CrossRef] [Scilit]
- Schmidl, B.; Hütten, T.; Pigorsch, S.; Stögbauer, F.; Hoch, C.C.; Hussain, T.; Wollenberg, B.; Wirth, M. Assessing the use of the novel tool Claude 3 in comparison to ChatGPT 4.0 as an artificial intelligence tool in the diagnosis and therapy of primary head and neck cancer cases. Eur. Arch. Otorhinolaryngol. 2024, 281, 6099–6109. [Google Scholar] [CrossRef] [Scilit]
- Stalp, J.L.; Denecke, A.; Jentschke, M.; Hillemanns, P.; Klapdor, R. Quality of ChatGPT-Generated Therapy Recommendations for Breast Cancer Treatment in Gynecology. Curr. Oncol. 2024, 31, 3845–3854. [Google Scholar] [CrossRef] [Scilit]
- Schmidl, B.; Hütten, T.; Pigorsch, S.; Stögbauer, F.; Hoch, C.C.; Hussain, T.; Wollenberg, B.; Wirth, M. Assessing the role of advanced artificial intelligence as a tool in multidisciplinary tumor board decision-making for primary head and neck cancer cases. Front. Oncol. 2024, 14, 1353031. [Google Scholar] [CrossRef] [Scilit]
- Aghamaliyev, U.; Karimbayli, J.; Giessen-Jung, C.; Matthias, I.; Unger, K.; Andrade, D.; Hofmann, F.O.; Weniger, M.; Angele, M.K.; Westphalen, C.B.; et al. ChatGPT’s Gastrointestinal Tumor Board Tango: A limping dance partner? Eur. J. Cancer 2024, 205, 114100. [Google Scholar] [CrossRef] [Scilit]
- Benary, M.; Wang, X.D.; Schmidt, M.; Soll, D.; Hilfenhaus, G.; Nassir, M.; Sigler, C.; Knödler, M.; Keller, U.; Beule, D.; et al. Leveraging Large Language Models for Decision Support in Personalized Oncology. JAMA Netw. Open 2023, 6, e2343689. [Google Scholar] [CrossRef] [Scilit]
- Griewing, S.; Gremke, N.; Wagner, U.; Lingenfelder, M.; Kuhn, S.; Boekhoff, J. Challenging ChatGPT 3.5 in Senology—An Assessment of Concordance with Breast Cancer Tumor Board Decision Making. J. Pers. Med. 2023, 13, 1502. [Google Scholar] [CrossRef] [Scilit]
- Vela Ulloa, J.; King Valenzuela, S.; Riquoir Altamirano, C.; Urrejola Schmied, G. Artificial intelligence-based decision-making: Can ChatGPT replace a multidisciplinary tumour board? Br. J. Surg. 2023, 110, 1543–1544. [Google Scholar] [CrossRef] [Scilit]
- Lukac, S.; Dayan, D.; Fink, V.; Leinert, E.; Hartkopf, A.; Veselinovic, K.; Janni, W.; Rack, B.; Pfister, K.; Heitmeir, B.; et al. Evaluating ChatGPT as an adjunct for the multidisciplinary tumor board decision-making in primary breast cancer cases. Arch. Gynecol. Obstet. 2023, 308, 1831–1844. [Google Scholar] [CrossRef] [Scilit]
- Delourme, S.; Redjdal, A.; Bouaud, J.; Seroussi, B. Measured Performance and Healthcare Professional Perception of Large Language Models Used as Clinical Decision Support Systems: A Scoping Review. Stud. Health Technol. Inform. 2024, 316, 841–845. [Google Scholar] [CrossRef] [Scilit]
- Sorin, V.; Klang, E.; Sklair-Levy, M.; Cohen, I.; Zippel, D.B.; Balint Lahat, N.; Konen, E.; Barash, Y. Large language model (ChatGPT) as a support tool for breast tumor board. NPJ Breast Cancer 2023, 9, 44. [Google Scholar] [CrossRef] [Scilit]
- Holzinger, A.; Carrington, A.; Müller, H. Measuring the Quality of Explanations: The System Causability Scale (SCS): Comparing Human and Machine Explanations. Künstliche Intell. 2020, 34, 193–198. [Google Scholar] [CrossRef] [Scilit]
- Cohen, J. Weighted kappa: Nominal scale agreement with provision for scaled disagreement or partial credit. Psychol. Bull. 1968, 70, 213–220. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Cohen, J. A Coefficient of Agreement for Nominal Scales. Educ. Psychol. Meas. 1960, 20, 37–46. [Google Scholar] [CrossRef] [Scilit]
- Shrout, P.E.; Fleiss, J.L. Intraclass correlations: Uses in assessing rater reliability. Psychol. Bull. 1979, 86, 420–428. [Google Scholar] [CrossRef]
- Landis, J.R.; Koch, G.G. The Measurement of Observer Agreement for Categorical Data. Biometrics 1977, 33, 159. [Google Scholar] [CrossRef] [Scilit]
- Koo, T.K.; Li, M.Y. A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research. J. Chiropr. Med. 2016, 15, 155–163. [Google Scholar] [CrossRef] [Scilit]
- Cronbach, L.J. Coefficient alpha and the internal structure of tests. Psychometrika 1951, 16, 297–334. [Google Scholar] [CrossRef] [Scilit]
- Taber, K.S. The Use of Cronbach’s Alpha When Developing and Reporting Research Instruments in Science Education. Res. Sci. Educ. 2018, 48, 1273–1296. [Google Scholar] [CrossRef] [Scilit]
- Wright, F.C.; de Vito, C.; Langer, B.; Hunter, A. Multidisciplinary cancer conferences: A systematic review and development of practice standards. Eur. J. Cancer 2007, 43, 1002–1010. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Huang, R.S.; Mihalache, A.; Nafees, A.; Hasan, A.; Ye, X.Y.; Liu, Z.; Leighl, N.B.; Raman, S. The impact of multidisciplinary cancer conferences on overall survival: A meta-analysis. J. Natl. Cancer Inst. 2024, 116, 356–369. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Berardi, R.; Morgese, F.; Rinaldi, S.; Torniai, M.; Mentrasti, G.; Scortichini, L.; Giampieri, R. Benefits and Limitations of a Multidisciplinary Approach in Cancer Patient Management. Cancer Manag. Res. 2020, 12, 9363–9374. [Google Scholar] [CrossRef] [Scilit] [PubMed]
| Prompt-Original Input in German | Prompt-English Translation |
|---|---|
| Formuliere eine stichpunktartige Therapieempfehlung für den folgenden Patientenfall. Die Vortherapien und andere relevante Befunde sind im Fall enthalten. Beschränke dich hierbei nicht nur auf Medikamente, sondern beschreibe alle möglichen Therapieoptionen. Bitte benenne Therapien und Medikamente, falls du eine medikamentöse Therapie vorschlägst, konkret. Versuche, deine Empfehlung anhand der in Deutschland zugelassenen und leitliniengerechten Therapien zu treffen. Bitte benenne zudem explizit die aus deiner Sicht beste Therapieoption für den individuellen Patienten. Begrenze mit Deinen Antworten auf maximal 80 Wörter und orientiere Dich in ihnen an folgender Struktur: (1.) Präferierte Therapie-empfehlung (falls vorhanden), (2.) Therapie-alternativen, (3.) Begründung der Empfeh-lungen, (4.) Supportivmaßnahmen/ergänzende Therapien, (5.) Weiterführende Informationen/Erklärungen. Patientenfall: | Formulate a key point-based treatment recommendation for the following patient case. The previous therapies and other relevant findings are included in the case. Do not limit yourself to medication, but describe all possible treatment options. Please name therapies and medications specifically if you are suggesting drug therapy. Try to make your recommendation based on the therapies approved in Germany and in line with the guidelines. Please also explicitly state what you consider to be the best treatment option for the individual patient. Limit your answers to a maximum of 80 words and base them on the following structure: (1) Preferred therapy recommendation (if available), (2) Therapy alternatives, (3) Justification of the recommendations, (4) Supportive measures/supplementary therapies, (5) Further information/explanations. Patient case: |
| (a) | ||||
|---|---|---|---|---|
| Interrater Reliability SCS | Interrater Reliability mSCS | |||
| MTB | LLM | MTB | LLM | |
| K (p-Value) | K (p-Value) | K (p-Value) | K (p-Value) | |
| Item 1 | 0.80 (<0.001) | 0.70 (<0.001) | 0.90 (<0.001) | 0.70 (<0.001) |
| Item 2 | 0.95 (<0.001) | 0.75 (<0.001) | 1.00 (<0.001) | 0.85 (<0.001) |
| Item 3 | 0.70 (<0.001) | 0.75 (<0.001) | 0.80 (<0.001) | 0.65 (<0.001) |
| Item 4 | 1.00 (<0.001) | 0.90 (<0.001) | 1.00 (<0.001) | 0.95 (<0.001) |
| Item 5 | 0.90 (<0.001) | 0.70 (<0.001) | 0.75 (<0.001) | 0.70 (<0.001) |
| Item 6 | 0.95 (<0.001) | 0.65 (<0.001) | 1.00 (<0.001) | 0.80 (<0.001) |
| Item 7 | 0.90 (<0.001) | 0.70 (<0.001) | 1.00 (<0.001) | 0.80 (<0.001) |
| Item 8 | 0.85 (<0.001) | 0.80 (<0.001) | 1.00 (<0.001) | 0.85 (<0.001) |
| Item 9 | 0.90 (<0.001) | 0.85 (<0.001) | 1.00 (<0.001) | 0.95 (<0.001) |
| Item 10 | 1.00 (<0.001) | 0.75 (<0.001) | 1.00 (<0.001) | 0.80 (<0.001) |
| All Items | 0.90 (<0.001) | 0.74 (<0.001) | 0.95 (<0.001) | 0.81 (<0.001) |
| (b) | ||||
| Interrater Reliability SCS | Interrater Reliability mSCS | |||
| MTB | LLM | MTB | LLM | |
| ICC (p-Value) | ICC (p-Value) | ICC (p-Value) | ICC (p-Value) | |
| Item 1 | 0.89 (<0.001) | 0.83 (<0.001) | 0.95 (<0.001) | 0.83 (<0.001) |
| Item 2 | 0.98 (<0.001) | 0.86 (<0.001) | 1.00 (<0.001) | 0.92 (<0.001) |
| Item 3 | 0.83 (<0.001) | 0.86 (<0.001) | 0.89 (<0.001) | 0.79 (<0.001) |
| Item 4 | 1.00 (<0.001) | 0.95 (<0.001) | 1.00 (<0.001) | 0.98 (<0.001) |
| Item 5 | 0.95 (<0.001) | 0.83 (<0.001) | 0.86 (<0.001) | 0.83 (<0.001) |
| Item 6 | 0.98 (<0.001) | 0.79 (<0.001) | 1.00 (<0.001) | 0.89 (<0.001) |
| Item 7 | 0.95 (<0.001) | 0.83 (<0.001) | 1.00 (<0.001) | 0.89 (<0.001) |
| Item 8 | 0.92 (<0.001) | 0.89 (<0.001) | 1.00 (<0.001) | 0.92 (<0.001) |
| Item 9 | 0.95 (<0.001) | 0.92 (<0.001) | 1.00 (<0.001) | 0.98 (<0.001) |
| Item 10 | 1.00 (<0.001) | 0.86 (<0.001) | 1.00 (<0.001) | 0.89 (<0.001) |
| All Items | 0.95 (<0.001) | 0.85 (<0.001) | 0.97 (<0.001) | 0.89 (<0.001) |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2024 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
Share and Cite
Rinderknecht, E.; von Winning, D.; Kravchuk, A.; Schäfer, C.; Schnabel, M.J.; Siepmann, S.; Mayr, R.; Grassinger, J.; Goßler, C.; Pohl, F.; et al. Modification and Validation of the System Causability Scale Using AI-Based Therapeutic Recommendations for Urological Cancer Patients: A Basis for the Development of a Prospective Comparative Study. Curr. Oncol. 2024, 31, 7061-7073. https://doi.org/10.3390/curroncol31110520
Rinderknecht E, von Winning D, Kravchuk A, Schäfer C, Schnabel MJ, Siepmann S, Mayr R, Grassinger J, Goßler C, Pohl F, et al. Modification and Validation of the System Causability Scale Using AI-Based Therapeutic Recommendations for Urological Cancer Patients: A Basis for the Development of a Prospective Comparative Study. Current Oncology. 2024; 31(11):7061-7073. https://doi.org/10.3390/curroncol31110520
Chicago/Turabian StyleRinderknecht, Emily, Dominik von Winning, Anton Kravchuk, Christof Schäfer, Marco J. Schnabel, Stephan Siepmann, Roman Mayr, Jochen Grassinger, Christopher Goßler, Fabian Pohl, and et al. 2024. "Modification and Validation of the System Causability Scale Using AI-Based Therapeutic Recommendations for Urological Cancer Patients: A Basis for the Development of a Prospective Comparative Study" Current Oncology 31, no. 11: 7061-7073. https://doi.org/10.3390/curroncol31110520
APA StyleRinderknecht, E., von Winning, D., Kravchuk, A., Schäfer, C., Schnabel, M. J., Siepmann, S., Mayr, R., Grassinger, J., Goßler, C., Pohl, F., Siska, P. J., Zeman, F., Breyer, J., Schmelzer, A., Gilfrich, C., Brookman-May, S. D., Burger, M., Haas, M., & May, M. (2024). Modification and Validation of the System Causability Scale Using AI-Based Therapeutic Recommendations for Urological Cancer Patients: A Basis for the Development of a Prospective Comparative Study. Current Oncology, 31(11), 7061-7073. https://doi.org/10.3390/curroncol31110520

