Review Reports
- Jianshan Cheng
Reviewer 1: Anonymous Reviewer 2: Anonymous Reviewer 3: Anonymous
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsThe manuscript is on the topic," Can Large Language Models Support University Counseling?Evidence from Counseling Alliance, Disclosure Willingness, and Risk Recognition," and inteeresting , there are some suggestions:
1. In the introduction, authors wrote the statement,"AI systems may provide
empathic responses while failing to accurately assess the severity of risk or recommend
appropriate intervention strategies," can a refrenece be provided in the manuscript.
2. In the introduction, it is mentioned that trends of using the LLMs is increasing, "ChatGPT demonstrate remarkable capabilities ", a trending statistical figure from the 2020 to 2026 can be given to strengthen the sentences.
3. In the literature section, there should be some supportive diagrams to enhance the interest of readers.
4. Section3 is giving Participants and recruitment procedure and a total of 400 questionnaires were distributed. how these questions were selected, any past experience, validation of correctness.
5. Foe section 3.3, procedure , a flow chart should be provided, for quick understanding.
6. In section 3.5, how the value of p < .05 was selected, any procedure.
7. Section 4 is presenting, results, 388 valid responses were retained for analysis, some validated proofs of statements must be submitted as supplementary material.
8. Any graphical statistical results can be provided, there are tables that are not sufficient.
9. Section 5.5 is practical implications, three observations are discussed, can 1-2 more practical implications can be given.
10. Conclusion section is long and can be revised for more quality.
Author Response
The manuscript is on the topic," Can Large Language Models Support University Counseling?Evidence from Counseling Alliance, Disclosure Willingness, and Risk Recognition," and interesting , there are some suggestions:
- In the introduction, authors wrote the statement,"AI systems may provide
emphathic responses while failing to accurately assess the severity of risk or recommend
appropriate intervention strategies," can a reference be provided in the manuscript.
Response:
Thank you for this valuable suggestion. We agree that this statement should be supported by empirical evidence. Accordingly, we have added recent references demonstrating that although advanced large language models often generate emphathic responses, they may still fail to reliably recognize crisis severity or provide appropriate intervention recommendations during high-risk mental health interactions. Specifically, we have cited Santos et al. (2026) and Stamatis et al. (2026) to support this statement.
- In the introduction, it is mentioned that trends of using the LLMs is increasing, "ChatGPT demonstrate remarkable capabilities ", a trending statistical figure from the 2020 to 2026 can be given to strengthen the sentences.
Response:
Thank you for this insightful suggestion. To better demonstrate the rapid adoption of large language models, we have added recent statistics describing the global growth of ChatGPT and LLM applications. The revised Introduction now notes that ChatGPT reached over 100 million monthly active users within two months of release and exceeded approximately 800 million weekly active users by 2025, illustrating the unprecedented expansion of LLM-based conversational systems. These statistics strengthen the rationale for investigating the application of LLMs in university counseling.
3.In the literature section, there should be some supportive diagrams to enhance the interest of readers.
Response:
Thank you for this helpful recommendation. Following the reviewer's suggestion, we have added a conceptual framework figure at the end of the Literature Review. The figure visually summarizes the hypothesized relationships among the LLM-based counseling condition, counseling alliance, disclosure willingness, and risk recognition. We believe this figure improves the readability of the theoretical framework and helps readers better understand the proposed research model.
- Section3 is giving Participants and recruitment procedure and a total of 400 questionnaires were distributed. how these questions were selected, any past experience, validation of correctness.
Response:
Thank you for this valuable comment. We agree that additional information regarding questionnaire development and validation improves the methodological transparency of the study. Accordingly, we have expanded Sections 3.2 and 3.4 to clarify the participant recruitment procedure, the rationale for the sample size, and the development and validation of the measurement instruments.
Specifically, we now explain that the counseling alliance scale was adapted from the established Working Alliance Inventory (WAI), while the disclosure willingness scale was developed based on previous digital mental health and self-disclosure research following standard scale-development procedures. Prior to the formal study, the questionnaire was pilot-tested with university students to evaluate item clarity and content validity, and minor wording revisions were made based on participant feedback. Reliability and construct validity were subsequently confirmed through Cronbach's alpha, composite reliability (CR), average variance extracted (AVE), and confirmatory factor analysis (CFA).
- Foe section 3.3, procedure , a flow chart should be provided, for quick understanding.
Response:
Thank you for this helpful suggestion. Following the reviewer's recommendation, we have added a flowchart illustrating the overall experimental procedure. The figure summarizes participant recruitment, random assignment, intervention procedures, post-test assessment, and statistical analysis, thereby improving the readability of the methodology.
- In section 3.5, how the value of p < .05 was selected, any procedure.
Response:
Thank you for this comment. We have clarified the statistical criteria adopted in this study. Following conventional practice in psychological and behavioral research, statistical significance was evaluated using a two-tailed significance level of α = .05. This threshold balances the risks of Type I and Type II errors and is widely recommended in behavioral science methodology. The corresponding explanation has been added to Section 3.5.
- Section 4 is presenting, results, 388 valid responses were retained for analysis, some validated proofs of statements must be submitted as supplementary material.
Response:
Thank you for this valuable comment. We appreciate the reviewer's concern regarding the transparency of the reported results. The statement that 388 valid responses were retained was derived from the predefined data screening procedures described in the Methods section. Specifically, all returned questionnaires were screened for completeness, response consistency, duplicate submissions, and response quality before statistical analyses were conducted. Only questionnaires meeting all predefined inclusion criteria were retained for analysis. To improve clarity, we have revised the manuscript to provide a more explicit description of the data screening and validation procedures in the Methods and Results sections, thereby making the basis for the final sample size transparent within the manuscript itself.
- Any graphical statistical results can be provided, there are tables that are not sufficient.
Response:
Thank you for this valuable suggestion. We agree that graphical presentations can improve the readability and interpretation of the statistical findings. In the revised manuscript, we have added additional graphical representations of the major statistical results to complement the existing tables. Specifically, we have included (1) a comparison figure illustrating differences between the LLM-based intelligent agent group and the control group across the main outcome variables (counseling alliance, disclosure willingness, and risk recognition), and (2) a conceptual mediation path figure presenting the indirect effect of counseling alliance between the experimental condition and disclosure willingness.
These graphical results provide a more intuitive understanding of the magnitude and direction of the observed effects while maintaining consistency with the statistical analyses reported in Tables 3–5. The revised figures have been added to the Results section.
- Section 5.5 is practical implications, three observations are discussed, can 1-2 more practical implications can be given.
Response:
Thank you for this constructive suggestion. We agree that the practical implications section can be further strengthened by expanding the discussion beyond the current three implications. In the revised manuscript, we have added two additional practical implications focusing on (1) AI system governance and counselor training, and (2) personalized implementation strategies for different student populations.
The expanded discussion emphasizes that successful adoption of LLM-based counseling agents requires not only technological capability but also institutional governance, ethical safeguards, and integration with existing counseling systems.
- Conclusion section is long and can be revised for more quality.
Response:
Thank you for this helpful comment. We agree that the previous conclusion included some repetition of the discussion section. In the revised manuscript, we have substantially shortened and refined the conclusion by focusing on the core findings, theoretical contribution, practical significance, and future direction. The revised conclusion avoids excessive repetition and provides a more concise synthesis of the study contribution.
Reviewer 2 Report
Comments and Suggestions for AuthorsComments for Authors
This manuscript examines the potential role of large language model (LLM)-based conversational agents in university counselling by investigating counselling alliance, disclosure willingness, and risk recognition among Chinese university students. The topic is highly timely and relevant, and the randomised experimental design represents a notable strength. The manuscript is well organised, clearly written, and grounded in an extensive and up-to-date review of the rapidly growing literature on AI-supported mental health interventions. The findings contribute to the emerging discussion of hybrid human–AI counselling models by emphasising both the potential benefits and safety limitations of LLM-based systems.
Several issues to be addressed before publication:
- Clarify the experimental intervention - The manuscript does not provide sufficient detail regarding the LLM-based counselling condition. The specific model, prompts, conversation protocol, duration of interactions, and degree of standardisation should be described more thoroughly to improve reproducibility.
- Strengthen ecological validity - The study relies on simulated counselling interactions and vignette-based risk recognition tasks. While appropriate for an initial experiment, these conditions differ substantially from authentic counselling sessions. The discussion should place greater emphasis on this limitation and avoid implying direct applicability to real clinical practice.
- Moderate interpretation of counselling alliance - The manuscript occasionally suggests that participants developed a counselling alliance comparable to therapeutic relationships with human counsellors. Because the alliance was assessed immediately following a brief experimental interaction, the findings are better interpreted as perceived interaction quality rather than evidence of an established therapeutic alliance. The discussion should adopt more cautious language.
- Interpret risk-recognition findings more cautiously - Although statistically significant, improvements in high-risk sensitivity were relatively modest compared with the large effects observed for alliance and disclosure willingness. Greater emphasis should be placed on practical significance rather than statistical significance, particularly given the importance of clinical safety.
- Address generalizability - The sample consisted exclusively of Chinese university students. Cultural factors related to help-seeking, stigma, and AI acceptance may limit the applicability of the findings to other educational or cultural contexts. These issues deserve greater discussion.
- Participants - Provide additional information on participant recruitment and response rates.
- Ensure consistent terminology when referring to counselling alliance, therapeutic alliance, and working alliance.
Author Response
Several issues to be addressed before publication:
1.Clarify the experimental intervention - The manuscript does not provide sufficient detail regarding the LLM-based counselling condition. The specific model, prompts, conversation protocol, duration of interactions, and degree of standardisation should be described more thoroughly to improve reproducibility.
Response:
Thank you for this important suggestion. We agree that a more detailed description of the LLM-based counseling intervention is necessary to improve methodological transparency and reproducibility. In the revised manuscript, we have expanded the description of the experimental intervention by specifying the LLM model used, prompt design principles, conversation protocol, interaction duration, and standardization procedures.
Specifically, we have clarified that the intervention used a GPT-based large language model configured with a standardized counseling-oriented system prompt. The prompt instructed the model to adopt an empathic, supportive, and non-diagnostic communication style, including reflective listening, open-ended questioning, emotional validation, and appropriate referral recommendations when risk-related content emerged. We have also added details regarding the interaction sequence, participant instructions, and procedures used to ensure consistency across participants.
These revisions have been incorporated into Section 3.3 (Procedure) of the manuscript.
2.Strengthen ecological validity - The study relies on simulated counselling interactions and vignette-based risk recognition tasks. While appropriate for an initial experiment, these conditions differ substantially from authentic counselling sessions. The discussion should place greater emphasis on this limitation and avoid implying direct applicability to real clinical practice.
Response:
Thank you for highlighting this important limitation. We agree that simulated counseling interactions and vignette-based assessments cannot fully represent the complexity of authentic counseling practice. In the revised manuscript, we have strengthened the discussion of ecological validity by explicitly acknowledging that the findings represent responses to controlled experimental interactions rather than evidence of effectiveness in real-world clinical counseling.
We have also revised the interpretation of practical implications to avoid suggesting immediate clinical deployment. Instead, we emphasize that LLM-based counseling agents should currently be considered experimental supportive technologies requiring further validation through longitudinal studies and field-based evaluations.
The relevant revisions have been made in Sections 5.3, 5.5, and 5.6.
3.Moderate interpretation of counselling alliance - The manuscript occasionally suggests that participants developed a counselling alliance comparable to therapeutic relationships with human counsellors. Because the alliance was assessed immediately following a brief experimental interaction, the findings are better interpreted as perceived interaction quality rather than evidence of an established therapeutic alliance. The discussion should adopt more cautious language.
Response:
Thank you for this insightful observation. We agree that the original manuscript used overly broad language when interpreting counseling alliance. The present study assessed participants’ immediate perceptions following a brief AI-mediated interaction, rather than a longitudinal therapeutic relationship comparable to human counseling.
Accordingly, we have revised the manuscript throughout by replacing stronger expressions suggesting the formation of a therapeutic alliance with more cautious terminology, such as “perceived counseling alliance,” “initial relational perception,” and “interactional quality.” We have also clarified this conceptual boundary in the Discussion section.
4.Interpret risk-recognition findings more cautiously - Although statistically significant, improvements in high-risk sensitivity were relatively modest compared with the large effects observed for alliance and disclosure willingness. Greater emphasis should be placed on practical significance rather than statistical significance, particularly given the importance of clinical safety.
Response:
Thank you for this important comment. We agree that statistical significance alone does not necessarily indicate practical importance, especially in the context of psychological risk detection where false negatives may have serious consequences. In the revised manuscript, we have moderated our interpretation of the risk-recognition findings by emphasizing that the improvement in high-risk sensitivity was statistically significant but relatively small compared with the stronger effects observed for perceived counseling alliance and disclosure willingness.
We have revised the Discussion and Practical Implications sections to clarify that the LLM-based agent may support preliminary identification of psychological concerns but should not be considered a reliable standalone crisis detection system. We now emphasize the importance of effect magnitude, safety considerations, and human oversight when interpreting AI-assisted risk recognition.
5.Address generalizability - The sample consisted exclusively of Chinese university students. Cultural factors related to help-seeking, stigma, and AI acceptance may limit the applicability of the findings to other educational or cultural contexts. These issues deserve greater discussion.
Response:
Thank you for raising this important issue. We agree that cultural and contextual factors may influence how students perceive AI-supported psychological services. The revised manuscript now provides a more explicit discussion of the cultural specificity of the sample, particularly regarding help-seeking stigma, attitudes toward psychological counseling, and acceptance of AI technologies in Chinese university settings.
We have expanded the limitations section to clarify that the findings should not be directly generalized to other cultural or educational contexts without further validation. Future research directions regarding cross-cultural comparisons and diverse populations have also been added.
6.Participants - Provide additional information on participant recruitment and response rates.
Response:
Thank you for this suggestion. We agree that additional information regarding recruitment procedures and response rates would improve transparency. In the revised manuscript, we have expanded the Participants section by providing details regarding recruitment channels, eligibility criteria, the number of individuals invited, completed responses, excluded responses, and the final analytical sample.
7.Ensure consistent terminology when referring to counselling alliance, therapeutic alliance, and working alliance.
Response:
Thank you for identifying this terminology issue. We agree that the previous manuscript used the terms “counseling alliance,” “therapeutic alliance,” and “working alliance” somewhat interchangeably, which may create conceptual ambiguity.
In the revised manuscript, we have standardized terminology throughout the manuscript. Because the present study examined participants’ immediate perceptions following a brief AI-mediated interaction rather than an established therapeutic relationship, we now consistently use the term “perceived counseling alliance” to describe the measured construct. The terms “therapeutic alliance” and “working alliance” are retained only when referring to previous theoretical literature or human counseling research.
Reviewer 3 Report
Comments and Suggestions for AuthorsOverall, this manuscript addresses a timely and meaningful topic by examining the potential role of LLM-based intelligent agents in university counseling. The study is well organized, the research design is appropriate, and the findings provide useful implications for both research and practice. In particular, the discussion, theoretical implications, and practical implications are generally well developed and logically presented. I therefore believe that the manuscript has publication potential after the following issues have been addressed.
1. The introduction is generally well written and provides a convincing justification for the study. I particularly appreciate that the authors establish the research context using the current mental health challenges and counseling environment of Chinese university students. However, the theoretical rationale for examining the relationships among the proposed variables could be strengthened further. While the conceptual summary improves the overall logic of the research model, the manuscript would benefit from a clearer explanation of why these variables should be examined within a single integrated framework.
2. The methodology requires additional clarification. The manuscript does not explicitly report the response format used for the self-report measures (e.g., five- or seven-point Likert scale). In addition, the disclosure willingness scale was developed for this study, yet the manuscript provides limited information regarding the scale development process. A brief description of item development and evidence supporting the content validity of the newly developed measure would improve methodological transparency.
3. The authors state that confirmatory factor analysis (CFA) was conducted and evaluated using χ²/df, CFI, TLI, RMSEA, and SRMR. However, the actual model fit indices are not reported in the manuscript. These results should be presented so that readers can evaluate the adequacy of the measurement model.
4. The discussion and implication sections are well organized and provide meaningful interpretations of the findings. However, the strong emphasis on the Chinese university context presented in the introduction becomes less visible in these sections. Reconnecting the findings to the cultural and institutional characteristics of Chinese higher education would further strengthen the coherence of the manuscript and better highlight its contribution.
Overall, I found this manuscript to be interesting and worthwhile. Addressing the points above would further improve its theoretical clarity, methodological transparency, and overall contribution.
Author Response
- The introduction is generally well written and provides a convincing justification for the study. I particularly appreciate that the authors establish the research context using the current mental health challenges and counseling environment of Chinese university students. However, the theoretical rationale for examining the relationships among the proposed variables could be strengthened further. While the conceptual summary improves the overall logic of the research model, the manuscript would benefit from a clearer explanation of why these variables should be examined within a single integrated framework.
Response:
Thank you for this insightful comment. We agree that the theoretical integration among the proposed variables required further clarification. In the revised manuscript, we strengthened the theoretical rationale for examining perceived counseling alliance, disclosure willingness, and risk recognition within a unified framework.
Specifically, we revised the Introduction and Literature Review sections to clarify that these three constructs represent complementary dimensions of AI-supported university counseling: (1) perceived counseling alliance reflects the relational quality of human–AI interaction, (2) disclosure willingness represents the psychological mechanism through which users engage with counseling support, and (3) risk recognition captures the safety boundary of AI applications in mental health contexts.
We further explained that LLM-based counseling should not be evaluated solely based on conversational fluency or perceived usefulness, but rather through an integrated framework combining relational effectiveness, engagement processes, and safety considerations. Accordingly, the revised manuscript presents LLM-based counseling as a hybrid support model in which relational connection facilitates disclosure, while risk recognition determines the appropriate boundary of AI involvement.
These revisions have been incorporated into the Introduction section and the newly strengthened “Summary of the Conceptual Logic” subsection.
- The methodology requires additional clarification. The manuscript does not explicitly report the response format used for the self-report measures (e.g., five- or seven-point Likert scale). In addition, the disclosure willingness scale was developed for this study, yet the manuscript provides limited information regarding the scale development process. A brief description of item development and evidence supporting the content validity of the newly developed measure would improve methodological transparency.
Response:
Thank you for this valuable suggestion. We agree that additional methodological details were necessary to improve transparency.
In the revised manuscript, we have clarified that all self-report measures were assessed using a five-point Likert scale ranging from 1 = strongly disagree to 5 = strongly agree.
Furthermore, we expanded the description of the disclosure willingness scale development process. Specifically, we added information regarding the theoretical basis of item generation, including previous research on online counseling and chatbot-mediated self-disclosure. We also clarified that the initial item pool was reviewed by experts and pilot-tested among undergraduate students before the formal study. Item wording was refined based on feedback regarding clarity and contextual appropriateness.
These revisions have been added to the Measures subsection to provide a clearer justification for the newly developed disclosure willingness measure and strengthen methodological transparency.
- The authors state that confirmatory factor analysis (CFA) was conducted and evaluated using χ²/df, CFI, TLI, RMSEA, and SRMR. However, the actual model fit indices are not reported in the manuscript. These results should be presented so that readers can evaluate the adequacy of the measurement model.
Response:
Thank you for identifying this omission. We agree that reporting CFA model fit indices is necessary for evaluating the adequacy of the measurement model.
In the revised manuscript, we have added the CFA results, including χ²/df, CFI, TLI, RMSEA, and SRMR values. These indices demonstrate that the measurement model achieved an acceptable fit to the data.
The relevant results have been added to the Measurement Quality subsection and Table 1 has been revised accordingly.
- The discussion and implication sections are well organized and provide meaningful interpretations of the findings. However, the strong emphasis on the Chinese university context presented in the introduction becomes less visible in these sections. Reconnecting the findings to the cultural and institutional characteristics of Chinese higher education would further strengthen the coherence of the manuscript and better highlight its contribution.
Response:
Thank you for this important suggestion. We agree that the Chinese higher education context should be more explicitly integrated into the interpretation of the findings.
In the revised Discussion and Practical Implications sections, we strengthened the connection between our findings and the characteristics of Chinese university counseling environments. Specifically, we added discussion regarding culturally influenced barriers to psychological help-seeking, including stigma concerns, fear of social evaluation, privacy concerns, and limited utilization of formal counseling services among Chinese university students.
We further clarified that the value of LLM-based counseling agents in this context lies not in replacing professional counselors, but in providing a culturally appropriate low-threshold pathway that encourages initial engagement and disclosure before students seek formal psychological support.
These revisions have been incorporated into Sections 5.2 and 5.5 of the revised manuscript.