1. Introduction
Testicular cancer (TC) is the most frequently diagnosed cancer among males aged 15–40 years [
1]. In men, it accounts for 1–2% of all tumours and, overall, is the 20th most common tumour [
2]. Like many patients, young men and their families will conduct their own research to better understand a medical condition. Around 30–50% of patients use the internet to access this information when diagnosed with cancer [
3]. A survey of patients with TC in Canada (
n = 44) found that 82% searched for information online about their condition and 94% believed that it increased their understanding [
4]. In this same study, common search terms included “treatment” (n = 17, 53%) and “survivorship” (n = 11, 34%) [
4]. However, little is known about the quality of information provided to people when enquiring about these topics online. Yeo et al. [
5] (pp. 363–368) evaluated the top 100 websites for TC and found that less than half of these disclosed the author (n = 44) and only 46 had been updated within the last two years. Reliability of this information was also variable, with only 42 websites citing a source, making it difficult for patients to judge the quality of the resources provided [
5].
In addition, distress is common after a diagnosis of cancer, and it can be difficult to engage patients due to a variety of factors. Providing patients with good-quality information and support may reduce this distress, aiming to improve engagement in care models. A study of patients recently diagnosed with TC (n = 39) assessed their distress after the implementation of a web-based tool, Nuts and Bolts, which is a Movember initiative designed to provide support for patients with TC [
6]. The tool provided information to patients about TC and connected them with TC survivors to facilitate one-on-one peer support [
6]. There was no significant reduction in distress (
p = 0.85) when patients used this web-based tool at the time of orchidectomy (early) or one week later (delayed) [
6]. However, patients who used the tool had lower median anxiety, depression and distress scores after orchidectomy, remaining stable from thereon [
6]. The key finding from this study was that providing patients with online support and good-quality information can reduce distress after the diagnosis of TC. This suggests that online resources could be used as adjuncts to reduce distress levels for patients with TC, if optimised appropriately. Although, the information provided must be judiciously evaluated by healthcare professionals before implementation in clinical practice.
Interestingly, patients may have lower expectations for the quality, understandability, readability and actionability of medical information. This is emphasised by patients searching for treatment options online, even if there is some inaccuracy, demonstrating a desire to be well-informed [
7]. As a result, we must consider a patient’s health literacy when developing medical information, so that it is understandable, accessible and actionable. Health literacy can be defined as the ability to “obtain, process and understand basic health information and services, required to make appropriate health decisions” [
8]. Patients without the necessary skills to comprehend information may struggle to engage successfully in their healthcare.
Regardless of the quality, it may seem that patients with TC tend to engage more actively with online information compared with paper-based material. Smith et al. [
9] (pp. 229–236) evaluated the data quality of online and postal questionnaires for patients with TC (n = 292) and found that online forms were returned at significantly faster rates than postal forms (mean [online] = 17.9 days, mean [postal] = 31.2,
p = 0.001). There were also significantly fewer reminders sent to patients for the online forms, who had not yet completed the questionnaires (mean [online] = 1.1, mean [postal] = 2.2,
p < 0.001) [
9]. Even if there are faults in online resources for health information, patients with TC at least show a preference to engage with web-based information.
Recently, the rise in interest in Artificial Intelligence (AI) has seen the widespread uptake of AI chatbots (AIC) which can quickly synthesise online information and respond to questions. The result is a more personalised online experience, with the ability to answer health queries and even provide medical advice. Yet this information provided to patients by AICs can be dubious and needs to be carefully scrutinised [
10]. Information about some urological cancers provided by AICs has been shown to lack readability, comprehensibility and actionability [
11]. Yet, information about TC provided by AICs has not been individually evaluated. This study aims to assess AICs as an information source about TC for patients, based on their quality, understandability, readability and actionability.
2. Materials and Methods
Google Trends and the Cancer Council Australia website were used to identify the most asked questions relating to TC. These questions addressed multiple aspects of TC, including the causes, diagnosis and treatment options (
Figure 1). The Google Trends search was conducted for a time interval of 1 January 2004 to present (15 April 2025), specifically using “all categories” and “web-search” in the search modifications tab. This method was deemed to most appropriately capture the questions asked by the general population over the past 20 years. The Cancer Council Australia webpage on TC was used to evaluate more structured questions from the point of view of a healthcare professional. Together, fourteen questions were selected to adequately capture the commonly asked questions related to TC without duplication. These were selected in collaboration with senior investigators B.T. (senior medical oncologist), N.M.C. (senior urologist) and N.S. (urology fellow), who frequently manage patients with TC. Those of a similar nature (e.g., question 1 and 4) were included to reflect specific patient concerns, as questions about TC may be asked in different ways. In May 2025, the questions were inputted into four different AICs: ChatSonic (Writesonic Inc., San Francisco, CA, USA), Bing AI (Microsoft Corporation, Redmond, WA, USA), ChatGPT version 4.0 (OpenAI, San Francisco, CA, USA) and Perplexity AI (Perplexity AI, Inc., San Francisco, CA, USA). The AICs produced responses that were collected under their default settings, and chat search history was cleared between inputs. A separate private browsing tab was used for each AIC when submitting the questions.
Evaluation of responses was conducted using three validated instruments, which have been similarly used in previous studies [
11]. The primary outcomes of this study were to determine the quality, readability, understandability and actionability of information provided by AICs about TC. The DISCERN score is used to assess information quality (1 = poor to 5 = high), which comprises sixteen Yes/No questions for the critical appraisal of written health information [
12]. For each question, a rating scale of 1–5 is given by the assessor (1 = “No” to 5 = “Yes”), representing how well the criterion is fulfilled by the publication [
12]. The Patient Education Materials Assessment Tool for Printable Materials (PEMAT) score is used to evaluate understandability and actionability [
13]. Understandability is defined as the likelihood of the reader to understand and explain the key messages of the information, whereas actionability is defined as the likelihood that a person knows how to act on that information provided [
14]. The PEMAT score is calculated as a percentage (%) of the criteria met for the material, and a threshold of 70% is used in PEMAT scoring tools to assess whether information is sufficiently understandable and/or actionable [
13]. These instruments both have clear instructions for investigators to assess the printed information accurately without requiring expertise on the topic of the information provided. Two investigators, trained in the use of these scoring tools, independently marked the responses of AICs (H.L and B.D). The scoring was overseen and reviewed by senior investigators N.S and N.M.C prior to analysis. Given that the scores recorded using DISCERN and PEMAT tools are subjective, inter-rater reliability (IRR) was evaluated to maintain scientific rigour. The Flesch–Kincaid readability score is used to assess reading complexity, associating it with an approximate education level one must need to interpret the text [
15]. A higher readability score indicates that the material is easier to read, with a lower score suggesting it is more complex and difficult to understand [
15]. The Flesch–Kincaid readability scores with their associated education levels comprise the following: 0–30 (college graduate), 30–50 (college student), 50–60 (10th–12th grade), 60–70 (8th–9th grade), 70–80 (7th grade), 80–90 (6th grade) and 90–100 (5th grade) [
15]. Additionally, the word count for each response was recorded and additionally used to determine the correlation with readability.
Findings were summarised using descriptive statistics, using median (interquartile range [IQR]) (
Table 1). Correlation between word count and readability was assessed using Pearson’s Correlation Coefficient (R), using a 95% confidence interval (CI) (
Table 2). Statistical analysis was undertaken to assess the IRR for DISCERN and PEMAT scores, using the calculation of Cohen’s Kappa (K), with a statistical significance level
p < 0.05 (
Table 3). All analyses were performed using Excel (Microsoft Corporation, Redmond, CA, USA) and SPSS version 30.0 (SPSS Inc., IBM Corp., Armonk, NY, USA).
3. Results
Across AICs, the overall responses recorded were brief, with a median word count of 61.5 (IQR 41.3–91.3). ChatSonic produced the longest responses (median 86, IQR 47.3–117) and the shortest were provided by Perplexity (median 32, IQR 27.5–41.8). Answers provided by AICs were written at a level of a college student, extrapolated from a median Flesch–Kincaid readability score of 34.1 (IQR 26.0–52.2). The exception was ChatSonic, which had a median readability score of a college graduate (29.8). The quality of information was uniformly low for all AICs (median 1, IQR 1–4) and slightly worse for Perplexity (median 1, IQR 1–3). Generally, the understandability of information was moderate across all four platforms, with a median PEMAT of 58.3 (IQR 50.0–66.7), and the highest scores were provided by ChatGPT (median 62.5, IQR 56.3–66.7). Together, actionability of responses was very poor (median 0, IQR 0–25), and individually, all AICs recorded a median PEMAT-Actionability score of 0%. Few responses provided by AICs scored highly, with occasional perfect scores for certain questions. However, our study found that AICs do not consistently provide high-quality, useful and actionable material about TC. K values to evaluate IRR for DISCERN, PEMAT-Actionability and PEMAT-Understandability were 0.822, 0.767 and 0.630, respectively (p < 0.001). These findings were statistically significant and indicate a substantial agreement between scores for PEMAT-Actionability and PEMAT-Understandability and almost perfect agreement for DISCERN.
4. Discussion
AICs represent a new wave of technology in healthcare that is widely adopted and requested by patients [
16]. However, given its rapid uptake, the medical information and advice provided by AICs requires meticulous interrogation prior to endorsement from healthcare professionals. This study evaluated responses provided by AICs in relation to frequently asked questions about TC, demonstrating that the information was moderately understandable yet lacked quality, readability and actionability. These findings reflect the cautious advocacy in the literature for the implementation of AICs in clinical practice, which have found low accuracy in providing information that aligns with major international guidelines [
17]. However, there is support for AICs to provide general information about health conditions for public awareness, rather than to suggest specific recommendations for patients [
17].
Moreover, there is a limited capacity of AICs to produce easy-to-read information about TC. This may ostracise certain people who lack the skills required to interpret answers provided by these models. At first, complex written information may not seem to affect patients with TC, who are often young adults and presumed to have higher reading comprehension than other demographics. However, a more nuanced interpretation recognises that a person’s reading level may be influenced by many factors other than age, such as race, ethnicity and language barriers [
18,
19]. Patient information that does not consider all these variables may render it inaccessible for certain populations.
In our study, all AICs produced brief responses that were written at a reading level of a college student, except for ChatSonic, which was associated with the reading level of a college graduate. ChatGPT had the highest median readability score (39.9), whilst ChatSonic produced the lowest (29.8). This was reflected by ChatSonic having a moderate negative correlation between word count and readability scores (R = −0.501, p = 0.068). In fact, ChatGPT 4.0, Bing AI and the combined results of AICs all shared a similar pattern, suggesting that a lower word count correlated to improve readability. However, the opposite was true for Perplexity, which had a moderate positive correlation between word count and readability (R = 0.392, p = 0.166). Comparing the responses to the other AICs, Perplexity produced sentences with shorter length and simple structures, minimising the use of complex grammatical tools. This may explain the difference in the trend between word count and readability scores when evaluating Perplexity, yet all results were still statistically insignificant (p > 0.05).
There was a poor quality of information across all platforms, indicated by a low median DISCERN score of 1 (IQR 1–4). This suggests a limited utility for patients when attempting to make informed decisions about TC. The broad IQR indicates that there is some ability for AICs to produce high-quality information, but it is not consistently delivered for all questions. When compared with other social media platforms like TikTok and YouTube, AICs do deliver material that is more objective and less likely to contain misinformation [
11,
20]. In our study, responses lacked sufficient coverage of many aspects to TC management, including treatment alternatives, considerations about making decisions and the implications these would have on a patient’s quality of life. The lack of references also impacted the scores, particularly for ChatGPT 4.0, which did not provide any references for its answers, despite a more detailed response. When listed, references were largely from reputable sources with valid links (URLs) from major, validated health information platforms, including the Mayo Clinic, National Health Service (NHS) and Cancer Council Australia. The reliability of sources has been documented in previous studies, which have studied AICs’ output for patient information on urological malignancies [
21]. The quality of information is particularly important for TC because it largely affects a young patient population, who must consider important management decisions for their condition, such as surgery, chemotherapy and fertility preservation [
22]. Additionally, this is a population that is especially likely to search and rely on online sources for information on their diagnosis and treatment options.
Understandability ranged from 50 to 62.5%, reflecting a moderate clarity. Most responses used language that was straightforward but often failed to clearly define medical jargon. Responses that scored poorly lacked structure and other strategies to simplify responses, such as dot points or headings. No images or schematics were utilised by AICs, relying on text only to explain concepts. ChatGPT 4.0 performed highest in understandability with a median of 62.5 (IQR 56.3–66.7), and Perplexity performed poorest with a median of 50.0 (IQR 50.0–58.3). AICs that performed better in understandability ratings used basic language and features like dot points and headings to improve comprehension. The actionability of responses was universally very poor and, to some extent, inactionable (median 0, IQR 0–25). Many answers failed to specify any guidance or next steps that a patient could take when consulting a healthcare professional. Interestingly, AICs failed to incorporate actionable content, even for questions about sex life, early detection and recurrence of cancer, which presumably should prompt action from patients. This implies AICs need to include actionable suggestions and behavioural prompts, which is a hallmark of effective patient education materials [
23]. If AICs included even general prompts, such as consulting a healthcare professional to discuss further, it could improve actionability.
This analysis demonstrates that AICs respond to frequently asked questions about TC with reasonable answers for educated readers who can comprehend the responses. Factors such as long sentences, medical jargon and lack of headings, tables or other visual aids may likely have contributed to this. The readability of AICs was above recommended reading levels from multiple healthcare institutions like the Centers for Disease Control and Prevention (CDC), which suggest health information for patients should not exceed a reading level of 6–8th grade [
14]. This may isolate patients with low health literacy levels from gaining suitable insight from responses when seeking to learn more about a medical condition or symptoms.
Apart from our study, patient information produced by AICs has been evaluated for other conditions in both urology and other specialties. Pan et al. [
24] (pp. 1437–1440) evaluated 100 responses of four different AICs (ChatGPT 3.5, Perplexity, ChatSonic and BingAi) for common queries related to skin, lung, breast, colorectal and prostate cancer. Similar to our findings, there was a moderate median PEMAT-Understandability score (66.7%), poor median PEMAT-Actionability (20.0%) and responses that were produced at a Flesch–Kincaid reading level of a college student [
24]. However, the quality of information produced in this study was higher, with a median DISCERN score of 5 [
24]. Another study evaluating patient information produced by AICs for interstitial cystitis found that the information produced had sufficient understandability with a median PEMAT-Understandability score of 75 (IQR 66.7–83.3), exceeding the 70% threshold [
25]. However, there was still low quality of information with a median DISCERN score of 3 (IQR 2–3) [
25] and a low PEMAT-Actionability score of 40 (IQR 20–60) [
25]. Across the literature, AICs tend to produce general information about medical conditions with variable quality and low actionability that is written above typical patient comprehension levels. Our findings mirrored those in other studies, although we found the quality of information produced by AICs for TC was quite low in comparison. To further investigate this, information related to TC could be evaluated with newer versions of AICs, which would be useful to compare results over time. This recognises that AICs undergo frequent model updates and therefore may improve the quality of information produced.
Limitations of this study include evaluating a small number of questions (n = 14) at a specific time point (15 April 2025). This snapshot may not be reflective of responses produced by AICs beyond this date. Evaluation in this study was limited to questions and responses in the English language. Future studies that incorporate a variety of languages would be beneficial to further understand the quality, readability and understandability in culturally and linguistically diverse populations. Another feature not considered was the accuracy of information produced, in relation to international guidelines or reputable patient information resources for TC, such as the Healthy Males website [
26]. This is an important area to evaluate for information produced by AICs to ensure that the current recommendations for certain medical conditions are provided to patients. Studies assessing the accuracy of outputs are critical to undertake prior to implementing AICs in clinical practice. Moreover, this study was limited to text-only questions without considering other types of inputs, such as image-based questions or other multimedia. Evaluating the latest updates of AICs with a variety of question types, whilst incorporating prompt engineering strategies to boost actionability, will be important in future research. Additional studies comparing AICs to online forums such as those in Reddit, which frequently host question and answer-based discussions with health experts, may also provide insight into the broad spectrum of tools at a patient’s disposal for in health information [
27].