Skip to Content
  • Article
  • Open Access

6 August 2026

Artificial Intelligence Chatbots as Patient Information Sources in Penile Cancer: A Multi-Platform Evaluation

,
,
and
1
Department of Surgery, The University of Melbourne, Grattan Street, Parkville, VIC 3052, Australia
2
Department of Urology, St Vincent’s Hospital Melbourne, 41 Victoria Parade, Fitzroy, VIC 3065, Australia
3
Department of Urology, Austin Health, 145 Studley Road, Heidelberg, VIC 3084, Australia
*
Author to whom correspondence should be addressed.

Abstract

Background/Objectives: Penile cancer carries a disproportionate burden of stigma and delayed presentation, with affected men increasingly turning to artificial intelligence (AI) chatbots as an anonymous information source. This study evaluated the quality, readability, understandability, actionability and clinical accuracy of AI chatbot responses to standardised penile cancer patient queries across six publicly available platforms. Methods: Fourteen standardised questions were submitted to ChatGPT-4o, Gemini, Perplexity, Microsoft Copilot, Claude, and DeepSeek, generating 84 responses. The responses were evaluated using DISCERN (information quality, scored 16–80), Patient Education Materials Assessment Tool for Printable Materials (PEMAT-P) (understandability and actionability, scored 0–100%), and Flesch–Kincaid grade level (reading complexity, recommended threshold ≤Grade 8). Between-platform and between-domain comparisons used the Kruskal–Wallis test with Dunn’s post hoc analysis. Clinical accuracy was assessed against the 2026 European Association of Urology-American Society of Clinical Oncology (EAU-ASCO) Collaborative Guidelines on Penile Cancer using a three-point ordinal scale across 10 guideline-scorable questions (maximum 20 points per platform); the four questions not addressed by the guidelines were scored separately for factual accuracy against authoritative external evidence. Results: No significant between-platform differences were identified for DISCERN (p = 0.796) or PEMAT-P understandability (p = 0.147). All platforms exceeded the 70% understandability adequacy threshold. Perplexity and Copilot demonstrated significantly higher actionability than all other platforms (100% vs. 75%, p < 0.001). No platforms achieved the recommended Grade 8 reading threshold, with Claude generating significantly more complex responses than ChatGPT-4o and Perplexity (grade 11.4 vs. 8.65 and 8.60, p < 0.05). Significant variation was identified across clinical domains for DISCERN (p < 0.001), Flesch–Kincaid grade level (p < 0.001), and word count (p < 0.001), with treatment-related questions achieving the highest information quality but also generating the most complex responses. Clinical accuracy scores ranged from 13/20 (ChatGPT-4o) to 16/20 (Copilot). Two critical errors were identified: Claude and DeepSeek both recommended bleomycin-containing chemotherapy regimens, directly contradicting the EAU-ASCO Strong recommendation against bleomycin due to pulmonary toxicity risk. Responses to the four survivorship and quality-of-life questions were factually accurate against external evidence in 23 of 24 cases. Conclusions: AI chatbot responses to penile cancer patient queries are broadly understandable but consistently fail to meet recommended readability thresholds and provide limited actionable guidance. Two platforms recommended bleomycin-containing regimens against an EAU-ASCO Strong recommendation; no platforms achieved full guideline concordance. Urologists should counsel patients on the limitations of AI chatbots as a health information source.

1. Introduction

Penile cancer is a rare malignancy with an incidence of 0.84 cases per 100,000 person-years; however, it is psychologically devastating for those who receive a diagnosis [1,2]. Unlike more prevalent genitourinary cancers, penile cancer carries a unique burden of stigma and embarrassment that frequently delays both presentation and diagnosis, with documented intervals between symptom onset and clinical review exceeding six months [3]. This combination of rarity and social sensitivity creates a structural gap in patient education where traditional resources are sparse, clinical consultations are brief, and patients who are reluctant to ask their doctors increasingly turn to artificial intelligence (AI) chatbots to fill that void.
The proliferation of AI chatbots, powered by large language models (LLMs), has fundamentally altered the landscape of health information seeking. Platforms such as ChatGPT are now capable of generating detailed, conversational responses to complex medical queries at any hour and without the friction of a clinical encounter. In a landmark cross-sectional study, AI chatbot responses to patient questions posted on public forums outperformed physician responses on measures of both quality and empathy [4]. Specialist preference for AI chatbot responses over physician answers has similarly been documented in oncological contexts [5]. Despite this rapid adoption, the reliability of AI-generated information for patient education remains an area of active and necessary scrutiny. Evaluations across urological conditions such as urological malignancies, kidney stones, benign prostatic hyperplasia, interstitial cystitis, and Wilms tumour have consistently demonstrated that AI chatbot responses, while broadly accurate, are written above recommended health literacy levels and lack actionable guidance [6,7,8,9,10]. Beyond information quality, the potential for AI chatbots to recommend interventions contrary to guideline-endorsed care represents a distinct safety concern that has received limited systematic evaluation.
Notably absent from this growing evidence base is any evaluation of AI chatbot performance in the context of penile cancer. Given the condition’s rarity and the documented propensity of affected men to seek information outside formal healthcare settings, understanding the quality of what AI platforms deliver is of direct clinical relevance. This study therefore aims to evaluate the quality, understandability, actionability, and readability of AI chatbot responses to penile cancer patient queries across six publicly available platforms: ChatGPT-4o, Gemini, Perplexity, Microsoft Copilot, Claude, and DeepSeek. Secondary objectives were to compare performance across platforms, to examine whether information quality varied according to the nature of the patient query, and to assess the clinical accuracy of responses against current European Association of Urology-American Society of Clinical Oncology (EAU-ASCO) guidelines.

2. Methods

This was a cross-sectional observational study requiring no ethical approval. Six AI chatbot platforms were evaluated: ChatGPT-4o (GPT-5.3 Instant; OpenAI, San Francisco, CA, USA), Gemini 3 Flash (Google LLC, Mountain View, CA, USA), Perplexity AI Sonar (Perplexity AI, Inc., San Francisco, CA, USA), Microsoft Copilot (Microsoft Corporation, Redmond, WA, USA), Claude sonnet 4.5 (Anthropic PBC, San Francisco, CA, USA), and DeepSeek-V3.2 (DeepSeek, Hangzhou, Zhejiang, China). All platforms were accessed via free, publicly available interfaces without paid subscriptions.

2.1. Question Generation

Fourteen standardised questions were developed through a structured process. Google Trends was searched using the terms “penile cancer”, “penis cancer”, and “cancer of the penis” from January 2004 to present to identify frequently searched patient queries. Patient-facing resources from Cancer Council Australia, Cancer Research UK, and the European Association of Urology (patients.uroweb.org) were reviewed to identify clinician-curated topics. A candidate question list was compiled ensuring coverage of key clinical domains, including aetiology, diagnosis and staging, treatment, treatment consequences, prognosis, and psychosocial impact, and was finalised following review by a senior urologist. Questions were classified into seven domains for secondary analysis: general (Q1–2), aetiology (Q3–4), diagnosis (Q5–6), treatment (Q7–9), treatment consequences (Q10–12), prognosis (Q13), and psychosocial (Q14). A full list of the questions can be found in Table 1.
Table 1. Standardised patient questions.

2.2. Data Collection

Between 1 and 14 April 2026, all fourteen questions were submitted to each platform independently, producing 84 responses. Each question was entered into a fresh private browsing session with chat history cleared. Questions were copied exactly as written without rephrasing or follow-up prompting. The model version was recorded for each platform. Responses were copied verbatim into a secure spreadsheet prior to scoring.

2.3. Scoring

Three instruments were applied to all 84 responses. DISCERN (16 items, scored 1–5 per item, total range of 16–80) assessed information quality [11]. Patient Education Materials Assessment Tool for Printable Materials (PEMAT-P) assessed understandability and actionability separately, each expressed as a percentage (≥70% considered adequate) [12]. The full, validated versions of both instruments were used. For PEMAT-P, items not applicable to text-only material (for example, those concerning visual aids) were scored ‘not applicable’ in accordance with the instrument’s standard scoring instructions and excluded from the denominator, so that no modification of the instruments was required. Flesch–Kincaid Grade Level assessed readability (recommended threshold ≤Grade 8). Word count was also recorded. Two investigators independently scored all responses using DISCERN and PEMAT-P, blinded to the platform. Flesch–Kincaid grade level and word count were calculated using Readable.com.
Inter-rater reliability for DISCERN total scores and PEMAT-P scores was assessed using the Intraclass Correlation Coefficient (ICC; two-way agreement). Individual DISCERN item agreement was assessed using weighted Kappa. ICC was interpreted using Koo and Li [13] and Kappa using Landis and Koch [14].

2.4. Statistical Analysis

Results are reported as median (interquartile range, IQR). Between-platform and between-domain comparisons used the Kruskal–Wallis test with Dunn’s post hoc test (Bonferroni correction) where significant. Pearson correlation assessed relationships between word count and DISCERN total, PEMAT-P understandability, and Flesch–Kincaid grade level. Statistical significance was p < 0.05. All analyses were performed in R version 4.5.0 (R Foundation for Statistical Computing, Vienna, Austria).

2.5. Clinical Accuracy Assessment

Clinical accuracy of all 84 responses was assessed against the 2026 EAU-ASCO Collaborative Guidelines on Penile Cancer by two expert urologist reviewers [15]. Each response was scored on a three-point ordinal scale: 0 (inconsistent; response contains a statement that directly contradicts an explicit EAU-ASCO recommendation, defined as either violating a Strong “Do not offer” recommendation, or actively recommending an intervention against which the guideline issues a Strong recommendation), 1 (partially consistent; guideline topic addressed but response is incomplete, imprecise, or omits clinically important recommended content), or 2 (consistent; key recommendations accurately conveyed). Four questions (Table 1; Q10, Q11, Q13, and Q14) were designated not applicable (N/A), as the EAU-ASCO guidelines do not provide substantive clinical content against which responses could be scored for these topics, yielding 60 scorable responses across 10 questions with a maximum possible score of 20 per platform. Because these four questions address survivorship and quality-of-life topics of high relevance to patients, the same two reviewers additionally assessed each of the corresponding 24 responses for factual accuracy against authoritative external evidence rather than guideline concordance: Surveillance, Epidemiology, and End Results (SEER) population survival data for Q13 [16], and peer-reviewed functional-outcome [17] and psycho-oncology [18] literature for Q10, Q11, and Q14. Responses were scored on the same three-point ordinal scale (0 = contains a statement factually inconsistent with authoritative evidence; 1 = broadly appropriate but incomplete or imprecise; 2 = factually accurate and consistent with authoritative evidence). This secondary factual-accuracy score was analysed separately from the guideline-concordance score and is reported in Supplementary Table S3. Both reviewers scored all responses independently, blinded to platform identity and to each other’s scores. Scoring discrepancies were resolved by discussion to consensus, and the consensus score was used in all analyses. Inter-rater agreement prior to adjudication was quantified using weighted Cohen’s Kappa for the guideline-concordance scores, interpreted according to Landis and Koch [14]; for the secondary factual-accuracy scores, which were concentrated in the highest-scoring category, agreement was quantified using raw percentage agreement and Gwet’s AC1 [19], which is more robust than Kappa under skewed marginal distributions. Responses scoring 0 were catalogued as critical errors and characterised by the specific guideline contradiction identified.

3. Results

3.1. Inter-Rater Reliability

The ICC for DISCERN total scores was 0.988 (95% confidence interval (CI) 0.949–0.995, p < 0.001), indicating excellent agreement. The ICCs for PEMAT-P understandability and actionability scores were 0.820 (95% CI 0.659–0.898) and 0.836 (95% CI 0.746–0.894) respectively, both indicating good agreement. Full item-level weighted Kappa results are provided in Supplementary Table S1. Agreement between the two reviewers for the clinical accuracy ordinal scores was almost perfect (weighted κ = 0.89, p < 0.001), with the reviewers initially concordant on 56 of 60 scorable responses (93.3%); all discrepancies were resolved by consensus.

3.2. Platform Comparisons

Median scores for all instruments by platform are presented in Table 2. No significant differences between platforms were identified for DISCERN total score (χ2(5) = 2.370, p = 0.796) or PEMAT-P understandability (χ2(5) = 8.179, p = 0.147). All platforms exceeded the PEMAT-P understandability adequacy threshold of 70%.
Table 2. Primary results by platform.
A significant difference in PEMAT-P actionability was identified between platforms (χ2(5) = 66.648, p < 0.001). Perplexity and Copilot both scored 100.0% across all responses, significantly higher than ChatGPT-4o, Gemini, Claude, and DeepSeek (all p < 0.001, Dunn’s post hoc Bonferroni-corrected). No significant differences were identified between Perplexity and Copilot (p = 1.000), nor among the remaining four platforms (all p = 1.000). All platforms met the 70% adequacy threshold.
A significant difference in Flesch–Kincaid grade level was identified between platforms (χ2(5) = 16.492, p = 0.006). No platform achieved the recommended Grade 8 threshold on median score. Claude produced significantly more complex responses than ChatGPT-4o (p = 0.017) and Perplexity (p = 0.031). No other pairwise comparisons were significant.
A significant difference in word count was identified between platforms (χ2(5) = 23.951, p < 0.001). Perplexity produced significantly shorter responses than DeepSeek (p = 0.002) and Gemini (p < 0.001). Gemini produced significantly longer responses than ChatGPT-4o (p = 0.029). Copilot produced significantly longer responses than Perplexity (p = 0.038). No significant differences were identified between other platform pairs.

3.3. Correlation Analysis

A moderate positive correlation was identified between word count and DISCERN total score (r = 0.404, 95% CI 0.208–0.569, p < 0.001). A moderate negative correlation was identified between word count and PEMAT-P understandability score (r = −0.431, 95% CI −0.591 to −0.239, p < 0.001). No significant correlations were identified between word count and Flesch–Kincaid grade level (r = 0.089, 95% CI −0.128 to 0.298, p = 0.421). Longer responses were therefore associated with higher information quality scores and lower understandability scores simultaneously.

3.4. Domain Analysis

Median scores by clinical domain across all instruments are presented in Table 3. Significant differences were identified across domains for DISCERN total score (χ2(6) = 70.356, p < 0.001), PEMAT-P understandability (χ2(6) = 21.692, p = 0.001), Flesch–Kincaid grade level (χ2(6) = 32.502, p < 0.001), and word count (χ2(6) = 25.337, p < 0.001). No significant differences in PEMAT-P actionability were identified across domains (χ2(6) = 1.344, p = 0.969).
Table 3. Domain analysis.
In post hoc analysis, the treatment and treatment consequences domains scored significantly higher on DISCERN than the aetiology, diagnosis, and general domains (all p < 0.001). The prognosis and treatment consequences domains scored significantly higher on PEMAT-P understandability than diagnosis (p = 0.005 and p = 0.019 respectively), with prognosis also scoring significantly higher than treatment (p = 0.028). The aetiology and treatment domains produced significantly more complex responses on Flesch–Kincaid grade level than the diagnosis and general domains (all p ≤ 0.006). Psychosocial generated significantly longer responses than general (p = 0.009) and prognosis (p = 0.009), and treatment consequences generated significantly longer responses than general (p = 0.022) and prognosis (p = 0.028).

3.5. Clinical Accuracy

Scores ranged from 13/20 (ChatGPT-4o) to 16/20 (Copilot), with Gemini and Claude each scoring 15/20 and Perplexity and DeepSeek each scoring 14/20. The overall aggregate score across all platforms was 87/120 (73%). No platforms achieved full guideline concordance across all scorable questions (Supplementary Table S2).
Two critical errors (score 0) were identified, both involving Q7 (treatment options for penile cancer). Claude recommended BMP (bleomycin, methotrexate, and cisplatin), and DeepSeek listed bleomycin among systemic chemotherapy agents, directly contradicting the EAU-ASCO Strong recommendation against bleomycin due to pulmonary toxicity risk. No other platforms recommended bleomycin.
The most consistent omission across all six platforms was low socioeconomic status as an European Association of Urology (EAU)-listed risk factor for penile cancer, with every platform scoring 1 for Q3. Staging imprecision was identified in five of the six platforms (ChatGPT-4o, Perplexity, Copilot, Claude, and DeepSeek), which described corpus spongiosum (T2) and corpus cavernosum (T3) invasion as a single, undifferentiated Stage II without distinguishing Stage IIA (T1b or T2) from Stage IIB (T3); only Gemini correctly separated the two sub-stages. Non-guideline treatments not addressed by the EAU-ASCO guidelines were recommended by four of six platforms: Mohs micrographic surgery (Gemini, Copilot, Claude, and DeepSeek), photodynamic therapy (Copilot), and cryotherapy (Copilot and DeepSeek). These responses were scored 1 (partially consistent), as they did not directly contradict an explicit EAU-ASCO recommendation but introduced content not endorsed by the guideline. The EAU-ASCO Strong recommendation restricting minimally invasive inguinal lymph node dissection to clinical trial settings for cN1–2 disease was absent from four of six platform responses to Q9.
One platform (Gemini) attributed non-guideline treatment recommendations (Mohs micrographic surgery) directly to the EAU-ASCO guidelines, representing a confabulated source attribution distinct from simple omission or imprecision.
On the secondary factual-accuracy assessment of the four survivorship and quality-of-life questions (Q10, Q11, Q13, and Q14; 24 responses), 23 of 24 responses scored 2 (factually accurate) and one scored 1, with no factually inconsistent statements (score 0) identified on any platform. The single partially consistent response (DeepSeek, Q13) reported stage-specific survival estimates that diverged from SEER population data; survival statistics on the remaining platforms were consistent with these figures, and responses addressing urinary function, sexual function, and emotional coping were judged consistent with established clinical and psycho-oncological practice. Inter-rater agreement was high (raw agreement, 91.7%; Gwet’s AC1 = 0.90), and the two discordant responses were resolved by consensus (Supplementary Table S3).

4. Discussion

This study evaluated the quality, readability, understandability, actionability and clinical accuracy of AI chatbot responses to fourteen standardised patient questions about penile cancer across six publicly available platforms. To our knowledge, this is the first study to evaluate AI chatbot information quality specifically in the context of penile cancer, addressing a notable gap in the growing literature on AI-generated health information in urology.
Overall information quality as measured by DISCERN was comparable across all six platforms, with no statistically significant between-platform differences. This finding is consistent with prior evaluations of AI chatbots in urological oncology, including Musheyev et al.’s, who similarly found no significant platform differences in response quality for prostate, bladder, kidney, and testicular cancer queries [20]. The uniformity of performance across platforms, including the newer DeepSeek model, suggests that limitations in AI-generated health information quality are systemic features of LLM architecture rather than platform-specific shortcomings.
All platforms exceeded the 70% PEMAT-P understandability adequacy threshold, indicating that AI chatbot responses were broadly comprehensible to patients. However, actionability was significantly lower for four of six platforms. Perplexity and Copilot consistently achieved 100% actionability across all responses, significantly outperforming ChatGPT-4o, Gemini, Claude, and DeepSeek. This difference likely reflects the search-augmented architecture and interface design of these two platforms. Unlike the predominantly parametric, conversational models, Perplexity and Copilot are built around live web retrieval and present responses with inline citations and source links; these citations function as the tangible tools and concrete next steps that the PEMAT-P actionability domain rewards, whereas the more discursive prose of the other platforms tended to omit them. Their structuring of responses around signposted resources and suggested follow-up queries further scaffolds user action. Higher actionability of this kind reflects the provision of links and prompts, however, rather than any guarantee of the quality or guideline concordance of the linked content. This is of particular clinical significance in penile cancer, where the interval between symptom onset and clinical review frequently exceeds six months [3]. In this context, a patient consulting an AI chatbot prior to or in lieu of seeking medical attention requires not only accurate information but explicit guidance on what to do next. The failure of four platforms to consistently provide this guidance represents a meaningful gap in the patient information journey at precisely the point where actionability is the most critical. One interpretation of this finding is that lower actionability may partly reflect a deliberate design choice to avoid directive recommendations that could encourage patients to act without professional input; such caution would be appropriate where it concerns specific treatment decisions. However, PEMAT-P actionability does not require directive medical advice but credits safe guidance such as advising prompt clinical review or outlining concrete next steps. That the two highest-scoring platforms achieved 100% actionability through this form of signposting, rather than by encouraging unsupervised action, indicates that high actionability and appropriate clinical deference are not mutually exclusive. The lower-scoring platforms therefore appear to have missed opportunities to scaffold safe next steps, most importantly the recommendation to seek timely specialist review, which is the most valuable action in a disease characterised by delayed presentation.
No platform achieved the Centers for Disease Control and Prevention (CDC)-recommended Grade 8 reading threshold on median Flesch–Kincaid score, with Claude producing significantly more complex responses than ChatGPT-4o and Perplexity. This finding mirrors consistently reported results across the broader AI chatbot health information literature [7,20,21] and is of particular relevance in penile cancer, where affected men are disproportionately older [22], may have lower health literacy, and are already navigating significant psychological distress at the time of information seeking [2].
A notable finding from the correlation analyses was the moderate positive correlation between word count and DISCERN total score alongside a simultaneous moderate negative correlation between word count and PEMAT-P understandability. Longer responses were associated with higher information quality but lower understandability. This trade-off suggests that when AI chatbots produce more comprehensive responses covering treatment options, risks, and benefits, thereby scoring higher on DISCERN, they do so in language that is less accessible to patients. This represents a fundamental tension in AI-generated health communication that has not been previously described in the urological AI chatbot literature.
The domain analysis revealed significant variation in information quality across clinical question types. The treatment and treatment consequences domains achieved the highest DISCERN scores, while the aetiology and treatment domains generated the most complex responses by Flesch–Kincaid grade level. Psychosocial questions generated substantially longer responses than any other domain, suggesting that AI chatbots elaborate considerably when addressing emotional content.
Both Claude and DeepSeek recommended bleomycin-containing chemotherapy regimens for Q7, directly contradicting an EAU-ASCO Strong recommendation [15]. The clinical accuracy assessment revealed that no platform achieved full guideline concordance, with partial consistency being the predominant pattern across scorable questions. AI chatbots consistently identified broad treatment frameworks while systematically omitting EAU-specific procedural restrictions, staging distinctions, and dose thresholds that differentiate guideline-concordant care from a generic overview. A patient using either platform to inform a conversation about systemic therapy could present with a treatment expectation that is not only guideline-discordant but potentially harmful. Beyond these contradictions, non-guideline treatments not addressed by the EAU-ASCO guidelines, including Mohs micrographic surgery, photodynamic therapy, and cryotherapy, were recommended by four of six platforms, indicating that AI chatbots frequently introduce content lacking guideline endorsement rather than restricting recommendations to evidence-based options. Notably, one platform (Gemini) attributed these non-guideline recommendations directly to the EAU-ASCO guidelines, demonstrating that AI-generated source attribution can be not only absent but actively confabulated, a finding that compounds the transparency deficits discussed below.
Importantly, no platform across any question provided dated information, reflecting a fundamental currency limitation of generative AI. Responses are produced without timestamps or indication of when the underlying information was last verified. Attribution and authorship criteria present similar inherent challenges, as AI-generated responses are not associated with identifiable authors or declared conflicts of interest. Critically, this transparency deficit was uniform across all six platforms regardless of their relative performance on other instruments, indicating it is a systemic architectural feature of generative AI rather than a platform-specific shortcoming. No prior AI chatbot evaluation in urological oncology has foregrounded this distinction.
These findings carry particular implications for the population most affected by penile cancer, which is concentrated among men of lower socioeconomic and educational status. The readability and actionability limitations identified here are likely amplified in groups with lower health literacy, raising the possibility that AI tools in their current form are the least accessible to those who could benefit the most and may widen rather than narrow existing information inequities, a risk compounded by limited access to devices, reliable internet, and digital literacy. Realising real-world benefit for this population would require equity-conscious design and implementation: constraining output to plain-language reading levels, offering multilingual and voice-based interfaces that lower literacy and typing barriers, and embedding these tools within trusted settings such as primary care, sexual health, and community services rather than relying on independent patient access. Given the propensity for delayed presentation in this disease, such tools should also foreground risk factor information and consistently prompt timely clinical review; the universal omission of low socioeconomic status as a listed risk factor across all platforms underscores that current outputs do not yet meet this need.
Several limitations of this study warrant acknowledgement. Data were collected from a single time point within a two-week window (April 2026), and given the rapid iterative development of large language models, findings may not reflect the performance of updated model versions. This study evaluated single, text-only responses to standardised queries and did not assess the impact of prompt engineering, multi-turn dialogue, or patient satisfaction with responses. This single-prompt design, while enabling standardised and reproducible comparison across platforms against a fixed guideline benchmark, does not capture the interactive, conversational capabilities that are increasingly recognised as a defining strength of these tools; evaluating how platforms perform across multi-turn clinical conversations in penile cancer, including their handling of follow-up and clarifying questions, represents an important and more novel direction for future work. The DISCERN and PEMAT-P instruments were originally developed for static printed health materials and were not designed for AI-generated conversational responses, which may limit their discriminatory validity in this context. We retained the full, validated instruments rather than abbreviated versions with non-applicable items removed, both to preserve comparability with the broader patient-education and AI chatbot evaluation literature and because such modified versions have not been independently validated; the items most affected by the text-only format, particularly source attribution and information currency, scored uniformly low across platforms, and we regard these as informative findings about current LLM patient-education output rather than artefacts to be removed. The three-point guideline-concordance scale was purpose-developed for this study and has not been formally validated, as no validated instruments exist for scoring the concordance of free-text AI responses against a specific clinical guideline; to mitigate subjectivity, each score was anchored to explicit criteria derived directly from the guideline’s recommendations, with the lowest score being reserved for direct contradiction of an explicit Strong recommendation, and the high inter-rater agreement between the two reviewers (weighted κ = 0.89) supports the reproducibility of the approach. In addition, the guideline-concordance score is, by construction, weighted toward the diagnostic, staging, and treatment domains, where the EAU-ASCO guidelines provide explicit recommendations, and does not capture survivorship and quality-of-life accuracy to the same extent. We partially addressed this through a secondary factual-accuracy assessment of the relevant questions against authoritative external evidence; however, no validated guideline standard exists against which to anchor accuracy for these highly patient-relevant topics. Finally, with fourteen questions across six platforms, the sample size for domain-level comparisons was modest, and results should be interpreted accordingly.

5. Conclusions

This study provides the first evaluation of AI chatbot information quality for penile cancer across six publicly available platforms. While responses were broadly understandable and information quality was comparable across platforms, no AI chatbot achieved recommended reading level thresholds, and four of six platforms demonstrated suboptimal actionability. Significant variation in performance was identified across clinical question domains, with treatment-related questions achieving the highest information quality scores but also generating the most complex and least accessible responses. Critical clinical accuracy errors were identified for two platforms, with no platform achieving full guideline concordance.
These findings have direct clinical relevance for a condition in which delayed presentation is common and patients frequently seek information outside formal healthcare settings. Urologists managing penile cancer should proactively counsel patients on the limitations of AI chatbots as a health information source, particularly regarding readability, actionability, and the absence of source verification. Future iterations of AI platforms should prioritise delivering information at appropriate health literacy levels with explicit actionable guidance. Subsequent research should assess whether AI chatbot responses meet the direct information needs of penile cancer patients and evaluate the impact of prompt engineering strategies on both response quality and clinical accuracy.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/siuj7040047/s1. Table S1: DISCERN item-level weighted Kappa; Table S2: Clinical accuracy of AI chatbot responses by question against EAU-ASCO Penile Cancer Guidelines (2026); Table S3: Secondary factual-assessment of survivorship and quality-of-life questions.

Author Contributions

Conceptualisation, K.O., S.H., D.W. and K.S.; Methodology, K.O., S.H., D.W. and K.S.; Validation, K.O., S.H., D.W. and K.S.; Formal Analysis, K.O., S.H., D.W. and K.S.; Investigation, K.O., S.H., D.W. and K.S.; Data Curation, K.O., S.H., D.W. and K.S.; Writing—Original Draft Preparation, K.O., S.H., D.W. and K.S.; Writing—Review and Editing, K.O., S.H., D.W. and K.S.; Visualisation, K.O., S.H., D.W. and K.S.; Supervision, K.S.; Project Administration, K.O. and K.S. All authors have read and agreed to the published version of the manuscript.

Funding

This research study received no external funding.

Institutional Review Board Statement

Ethical review and approval were waived for this study as it analysed publicly available, non-identifiable information generated by artificial intelligence chatbot platforms and did not involve human participants, patient data, or animal subjects. Formal institutional ethics approval was not required in accordance with local institutional guidelines and the National Health and Medical Research Council (NHMRC) National Statement on Ethical Conduct in Human Research.

Data Availability Statement

The data presented in this study are available within the article and Supplementary Materials. Raw chatbot responses and scoring data are available from the corresponding author upon reasonable request.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Montes Cardona, C.E.; García-Perdomo, H.A. Incidence of Penile Cancer Worldwide: Systematic Review and Meta-Analysis. Rev. Panam. Salud Publica 2017, 41, e117. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  2. Maddineni, S.B.; Lau, M.M.; Sangar, V.K. Identifying the Needs of Penile Cancer Sufferers: A Systematic Review of the Quality of Life, Psychosexual and Psychosocial Literature in Penile Cancer. BMC Urol. 2009, 9, 8. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  3. Skeppner, E.; Andersson, S.-O.; Johansson, J.-E.; Windahl, T. Initial Symptoms and Delay in Patients with Penile Carcinoma. Scand. J. Urol. Nephrol. 2012, 46, 319–325. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  4. Ayers, J.W.; Poliak, A.; Dredze, M.; Leas, E.C.; Zhu, Z.; Kelley, J.B.; Faix, D.J.; Goodman, A.M.; Longhurst, C.A.; Hogarth, M.; et al. Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum. JAMA Intern. Med. 2023, 183, 589–596. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Singhal, K.; Tu, T.; Gottweis, J.; Sayres, R.; Wulczyn, E.; Amin, M.; Hou, L.; Clark, K.; Pfohl, S.R.; Cole-Lewis, H.; et al. Toward Expert-Level Medical Question Answering with Large Language Models. Nat. Med. 2025, 31, 943–950. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  6. Ozgor, F.; Caglar, U.; Halis, A.; Cakir, H.; Aksu, U.C.; Ayranci, A.; Sarilar, O. Urological Cancers and ChatGPT: Assessing the Quality of Information and Possible Risks for Patients. Clin. Genitourin. Cancer 2024, 22, 454–457.e4. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  7. Santucci, J.; Stapleton, P.; Ibrahim, J.; Johns-Putra, L.; Elmer, S.; Sathianathen, N. Quality of Patient Information on Interstitial Cystitis from Artificial Intelligence Chatbots. BJU Int. 2026, 138, S57–S63. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  8. Stapleton, P.; Santucci, J.; Cundy, T.P.; Sathianathen, N. Quality of Information on Wilms Tumor From Artificial Intelligence Chatbots: What Are Your Patients and Their Families Reading? Urology 2025, 198, 130–134. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. Musheyev, D.; Pan, A.; Kabarriti, A.E.; Loeb, S.; Borin, J.F. Quality of Information About Kidney Stones from Artificial Intelligence Chatbots. J. Endourol. 2024, 38, 1056–1061. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  10. Warren, C.J.; Payne, N.G.; Edmonds, V.S.; Voleti, S.S.; Choudry, M.M.; Punjani, N.; Abdul-Muhsin, H.M.; Humphreys, M.R. Quality of Chatbot Information Related to Benign Prostatic Hyperplasia. Prostate 2025, 85, 175–180. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  11. Charnock, D.; Shepperd, S.; Needham, G.; Gann, R. DISCERN: An Instrument for Judging the Quality of Written Consumer Health Information on Treatment Choices. J. Epidemiol. Community Health 1999, 53, 105–111. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  12. Agency for Healthcare Research and Quality. The Patient Education Materials Assessment Tool (PEMAT) and User’s Guide. Available online: https://www.ahrq.gov/health-literacy/patient-education/pemat-p.html (accessed on 4 May 2026).
  13. Koo, T.K.; Li, M.Y. A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research. J. Chiropr. Med. 2016, 15, 155–163. [Google Scholar] [PubMed]
  14. Landis, J.R.; Koch, G.G. The Measurement of Observer Agreement for Categorical Data. Biometrics 1977, 33, 159–174. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  15. European Association of Urology. EAU-ASCO Collaborative Guidelines on Penile Cancer. 2026. Available online: https://uroweb.org/guidelines/penile-cancer (accessed on 19 June 2026).
  16. Surveillance Research Program, National Cancer Institute. SEER*Explorer: An Interactive Website for SEER Cancer Statistics. Surveillance Research Program, National Cancer Institute. Available online: https://seer.cancer.gov/statistics-network/explorer/ (accessed on 19 June 2026).
  17. Croghan, S.M.; Cullen, I.M.; Raheem, O. Functional Outcomes and Health-Related Quality of Life Following Penile Cancer Surgery: A Comprehensive Review. Sex. Med. Rev. 2023, 11, 441–459. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  18. Torres Irizarry, V.M.; Paster, I.C.; Ogbuji, V.; Gomez, D.M.; Mccormick, K.; Chipollini, J. Improving Quality of Life and Psychosocial Health for Penile Cancer Survivors: A Narrative Review. Cancers 2024, 16, 1309. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  19. Gwet, K.L. Computing Inter-Rater Reliability and Its Variance in the Presence of High Agreement. Br. J. Math. Stat. Psychol. 2008, 61, 29–48. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  20. Musheyev, D.; Pan, A.; Loeb, S.; Kabarriti, A.E. How Well Do Artificial Intelligence Chatbots Respond to the Top Search Queries About Urological Malignancies? Eur. Urol. 2024, 85, 13–16. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  21. Rothchild, E.; Jung, G.; Ricci, J.A. Evaluating ChatGPT’s Ability to Address Frequently Asked Questions in Gender-Affirmation Surgery. J. Homosex. 2026, 73, 142–153. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  22. Douglawi, A.; Masterson, T.A. Penile Cancer Epidemiology and Risk Factors: A Contemporary Review. Curr. Opin. Urol. 2019, 29, 145–149. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.