Automated Working Alliance Assessment in Psychological Counseling Using Gemini and XGBoost
Abstract
1. Introduction
- (1)
- A generative data augmentation strategy is proposed, through which the PsyCase corpus is constructed with heterogeneous multilingual psychotherapy case reports to fine-tune BERTs. This approach successfully delineates speaker roles within complex, multilingual narrative texts across Chinese and English reports, establishing a structural foundation for automated alliance assessment.
- (2)
- A computational framework for the automated assessment of therapeutic alliance from complex case reports is proposed and validated, demonstrating performance superior to existing benchmarks.
2. Working Alliance Assessment Framework Based on Gemini and XGBoost
2.1. Data Preparation
2.2. Speaker Role Delineation
2.3. Alliance Rating Annotation with Expert-Defined Rubrics
2.4. Working Alliance Assessment
2.4.1. Hybrid Feature Engineering
2.4.2. Predictive Modeling and Optimization
3. Experiments
3.1. Identifying Dialogues and Speakers in a Case Report
3.1.1. Datasets and Baselines
- Format. To ensure that the PsyCase dataset reflects the structural characteristics and informational richness of real-world psychotherapy records, we adopt three standardized documentation formats commonly used by clinical psychologists and counselors: SOAP (Subjective–Objective–Assessment–Plan) [20], BIRP (Behavior–Intervention–Response–Plan), and DAP (Data–Assessment–Plan) [21]. To enhance the structural diversity, two informal formats are also incorporated—narrative and timeline.
- Content. Generation prompts specify a broad range of clinical and psychological focuses, including intervention procedures, metaphor identification and interpretation, emotional tracking and annotation, cognitive-behavioral (CBT) analysis, psychodynamic analysis, person-centered process focus, behavioral pattern and function analysis, and cultural context analysis [22].
- Style. Four distinct writing styles are defined: academic, instructional, client-oriented, and reflective-practice. These three dimensions can be flexibly combined to generate a diverse set of case reports with coherent structure and professional expression.
3.1.2. Experimental Settings and Evaluation Metrics
3.1.3. Experimental Results and Comparison
3.2. Working Alliance Assessment
3.2.1. Dataset
3.2.2. Experimental Settings and Evaluation Metrics
3.2.3. Performance Comparison
3.2.4. Ablation Study
3.2.5. Linguistic Feature Analysis
3.2.6. Generalization Capability Analysis
3.3. Validation of LLM-Generated Alliance Scores
4. Discussion
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Data Availability Statement
Conflicts of Interest
Abbreviations
| WAI | Working Alliance Inventory |
| TF-IDF | Term Frequency-Inverse Document Frequency |
| NLP | Natural Language Processing |
| BERT | Bidirectional Encoder Representations from Transformers |
| LLMs | Large Language Models |
| SOAP | Subjective–Objective–Assessment–Plan |
| BIRP | Behavior–Intervention–Response–Plan |
| DAP | Data–Assessment–Plan |
| CBT | Cognitive-Behavioral Therapy |
| FCNN | Fully Connected Neural Network |
| POS | Part-of-Speech |
| LSM | Linguistic Style Matching |
| LSDR | Linguistic Style Dominance Ratio |
| RFECV | Recursive Feature Elimination with Cross-Validation |
| XGBoost | eXtreme Gradient Boosting |
| MAE | Mean Absolute Error |
| RMSE | Root Mean Square Error |
| ICC | Intraclass Correlation Coefficient |
| SVR | Support Vector Regression |
| SD | Standard Deviation |
References
- Horvath, A.O.; Luborsky, L. The role of the therapeutic alliance in psychotherapy. J. Consult. Clin. Psychol. 1993, 61, 561. [Google Scholar] [CrossRef] [PubMed]
- Bordin, E.S. The generalizability of the psychoanalytic concept of the working alliance. Psychother. Theory Res. Pract. 1979, 16, 252–260. [Google Scholar] [CrossRef] [Scilit]
- Horvath, A.O. An Exploratory Study of the Working Alliance: Its Measurement and Relationship to Therapy Outcome. Ph.D. Thesis, University of British Columbia, Vancouver, BC, Canada, 1981. [Google Scholar]
- Goldberg, S.B.; Flemotomos, N.; Martinez, V.R.; Tanana, M.J.; Kuo, P.B.; Pace, B.T.; Villatte, J.L.; Georgiou, P.G.; Van Epps, J.; Imel, Z.E.; et al. Machine learning and natural language processing in psychotherapy research: Alliance as example use case. J. Couns. Psychol. 2020, 67, 438–448. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Salton, G.; Buckley, C. Term-weighting approaches in automatic text retrieval. Inf. Process. Manag. 1988, 24, 513–523. [Google Scholar] [CrossRef] [Scilit]
- Zhou, Y.; Chen, X.Y.; Liu, D.; Pan, Y.L.; Hou, Y.F.; Gao, T.T.; Peng, F.; Wang, X.C.; Zhang, X.Y. Predicting first session working alliances using deep learning algorithms: A proof-of-concept study for personalized psychotherapy. Psychother. Res. 2022, 32, 1100–1109. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Xie, H.; Chen, Y.; Xing, X.; Lin, J.; Xu, X. PsyDT: Using LLMs to construct the digital twin of psychological counselor with personalized counseling style for psychological counseling. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, Vienna, Austria, 27 July–1 August 2025; Volume 1, pp. 1081–1115. [Google Scholar] [CrossRef] [Scilit]
- Chaszczewicz, A.; Shah, R.S.; Louie, R.; Arnow, B.A.; Kraut, R.; Yang, D. Multi-level feedback generation with large language models for empowering novice peer counselors. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, Bangkok, Thailand, 11–16 August 2024; Volume 1, pp. 4130–4161. [Google Scholar] [CrossRef] [Scilit]
- Chen, K.; Sun, Z.; Wen, Y.; Lian, H.; Gao, Y.; Li, Y. Psy-Insight: Explainable multi-turn bilingual dataset for mental health counseling. arXiv 2025, arXiv:2503.03607. [Google Scholar]
- Zhang, C.; Li, R.; Tan, M.; Yang, M.; Zhu, J.; Yang, D.; Zhao, J.; Ye, G.; Li, C.; Hu, X. CPsyCoun: A report-based multi-turn dialogue reconstruction and evaluation framework for Chinese psychological counseling. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2024, Bangkok, Thailand, 11–16 August 2024; pp. 13947–13966. [Google Scholar] [CrossRef] [Scilit]
- Guo, D.; Yang, D.; Zhang, H.; Song, J.; Wang, P.; Zhu, Q.; Xu, R.; Zhang, R.; Ma, S.; Bi, X.; et al. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature 2025, 645, 633–638. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Luo, R.; Xu, J.; Zhang, Y.; Zhang, Z.; Ren, X.; Sun, X. PKUSEG: A toolkit for multi-domain Chinese word segmentation. arXiv 2019, arXiv:1906.11455. [Google Scholar]
- Honnibal, M.; Montani, I.; Van Landeghem, S.; Boyd, A. spaCy: Industrial-Strength Natural Language Processing in Python. 2017. Available online: https://github.com/explosion/spaCy (accessed on 3 March 2025).
- Doorn, A.V.; Porcerelli, J.; Müller-Frommeyer, L.C. Language style matching in psychotherapy: An implicit aspect of alliance. J. Couns. Psychol. 2020, 67, 509–522. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Pedregosa, F.; Varoquaux, G.; Gramfort, A.; Michel, V.; Thirion, B.; Grisel, O.; Blondel, M.; Prettenhofer, P.; Weiss, R.; Dubourg, V.; et al. Scikit-learn: Machine learning in Python. J. Mach. Learn. Res. 2011, 12, 2825–2830. [Google Scholar]
- Shannon, C.E. A mathematical theory of communication. Bell Syst. Tech. J. 1948, 27, 379–423. [Google Scholar] [CrossRef] [Scilit]
- Kraskov, A.; Stögbauer, H.; Grassberger, P. Estimating mutual information. Phys. Rev. E 2004, 69, 066138. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chen, J.; Xiao, S.; Zhang, P.; Luo, K.; Lian, D.; Liu, Z. M3-Embedding: Multi-linguality, multi-functionality, multi-granularity text embeddings through self-knowledge distillation. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2024, Bangkok, Thailand, 11–16 August 2024; pp. 2318–2335. [Google Scholar] [CrossRef] [Scilit]
- Chen, T.; Guestrin, C. XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, 13–17 August 2016; pp. 785–794. [Google Scholar] [CrossRef] [Scilit]
- Cameron, S.; Turtle-Song, I. Learning to write case notes using the SOAP format. J. Couns. Dev. 2002, 80, 286–292. [Google Scholar] [CrossRef] [Scilit]
- Reiter, M.D. A Therapist’s Guide to Writing in Psychotherapy: Assessment, Documentation, and Intervention, 1st ed.; Routledge: New York, NY, USA, 2023. [Google Scholar] [CrossRef] [Scilit]
- Norcross, J.C.; Goldfried, M.R. (Eds.) Handbook of Psychotherapy Integration, 3rd ed.; Oxford University Press: New York, NY, USA, 2019. [Google Scholar] [CrossRef] [Scilit]
- Akiba, T.; Sano, S.; Yanase, T.; Ohta, T.; Koyama, M. Optuna: A next-generation hyperparameter optimization framework. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, Anchorage, AK, USA, 4–8 August 2019; pp. 2623–2631. [Google Scholar] [CrossRef] [Scilit]
- Martinez, V.R.; Flemotomos, N.; Ardulov, V.; Somandepalli, K.; Goldberg, S.B.; Imel, Z.E.; Atkins, D.C.; Narayanan, S. Identifying therapist and client personae for therapeutic alliance estimation. In Proceedings of the Interspeech 2019, Graz, Austria, 15–19 September 2019; pp. 1901–1905. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- OpenAI. GPT-4o system card. arXiv 2024, arXiv:2410.21276. [Google Scholar]
- Aafjes-Van Doorn, K.; Cicconet, M.; Cohn, J.F.; Aafjes, M. Predicting working alliance in psychotherapy: A multi-modal machine learning approach. Psychother. Res. 2025, 35, 256–270. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Qiu, H.; Lan, Z. PsyDial: A large-scale long-term conversational dataset for mental health support. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, Vienna, Austria, 27 July–1 August 2025; Volume 1, pp. 21624–21655. [Google Scholar] [CrossRef] [Scilit]





| No. | Subdimension | Description |
|---|---|---|
| 1 | Goal Clarity | The client clearly expressed the goals of the counseling. |
| 2 | Goal Summary | The therapist summarized and confirmed the client’s goals. |
| 3 | Goal Adjustment | Both parties discussed and adjusted the counseling goals together. |
| 4 | Client Participation | The client actively participated in the counseling process. |
| 5 | Therapist Support | The therapist provided effective guidance and support. |
| 6 | Effective Communication | Both parties were able to communicate effectively. |
| 7 | Positive Emotion | The client expressed positive emotions, such as hope, trust, and gratitude. |
| 8 | Therapist Empathy | The therapist expressed empathy and understanding. |
| 9 | Atmosphere Label | The conversation atmosphere was relaxed, natural, and sincere. |
| 10 | Self Disclosure | The client was willing to disclose themselves and share inner feelings. |
| 11 | Active Listening | The therapist demonstrated active listening and responsive engagement. |
| (A) | |||
| Feature Category | Feature Name | Description | Purpose |
| Pronoun Usage | i_pron_ratio | The proportion of first-person singular pronouns (e.g., ‘I’, ‘me’, ‘my’, ‘myself’) to the total word count per utterance. | To capture self-focused expression within each utterance. |
| we_pron_ratio | The proportion of first-person plural pronouns (e.g., ‘we’, ‘us’, ‘our’, ‘ourselves’) to the total word count per utterance. | To capture collective or collaborative expression within each utterance. | |
| Linguistic Style & Cognitive Process | non_fluency_ratio | The proportion of disfluency markers (e.g., ‘um’, ‘uh’, ‘er’, ‘hmm’) to the total word count per utterance. | To reflect hesitation, uncertainty, or disruptions in verbal fluency. |
| lsm_{category}_ratio | The proportion of words belonging to specific Part-of-Speech (POS) categories (e.g., verbs, adjectives, adverbs, conjunctions) per utterance. | To characterize the linguistic style and lexical composition of each utterance. | |
| (B) | |||
| Feature Name | Description | Purpose | |
| role-level statistics (_mean, _std, _sum) | For each utterance-level feature, the mean, standard deviation, and sum are calculated across all utterances for a specific role (client/therapist) within a session. | To capture role-specific linguistic tendencies, variability, and accumulated usage patterns within a session. | |
| talk_turn_ratio | The ratio of the total number of turns taken by the client to the total number of turns taken by the therapist in a session. | To measure the balance of conversational participation between the client and the therapist. | |
| cross-role ratios (ratio_{feature}) | The ratio of the client’s mean value for a given linguistic feature to the therapist’s mean value for the same feature (client_mean/therapist_mean). | To quantify linguistic asymmetry between the client and the therapist for the same feature. | |
| Dataset (Language) | K | |||||
|---|---|---|---|---|---|---|
| PsyCase (cn) | 73 | 1.292 | 0.733 | 43.3% | 0.684 | 47.1% |
| Psy-Insight (cn) | 38 | 1.279 | 0.780 | 39.0% | 0.758 | 40.7% |
| Dataset (Language) | Case-Reports | Counselor-Utterances | Client-Utterances | Other-Utterances |
|---|---|---|---|---|
| PsyCase (cn) | 400 | 7279 | 7279 | 17,013 |
| Psy-Insight (cn) | 431 | 2911 | 2865 | 1292 |
| CPsyCounR(cn) | 3127 | 2853 | 2901 | 32,907 |
| Unified Corpus (cn) | 3958 | 13,043 | 13,045 | 51,212 |
| PsyCase (en) | 400 | 1902 | 1731 | 19,115 |
| Psy-Insight (en) | 520 | 3097 | 3111 | 3090 |
| Unified Corpus (en) | 920 | 4999 | 4842 | 22,205 |
| Model | Precision Macro | F1 Macro | Accuracy |
|---|---|---|---|
| XGBoost (cn) | 0.94 | 0.93 | 0.95 |
| BERT (cn) | 0.97 | 0.97 | 0.98 |
| FCNN (cn) | 0.94 | 0.94 | 0.96 |
| Logistic Regression (cn) | 0.96 | 0.95 | 0.96 |
| XGBoost (en) | 0.91 | 0.91 | 0.95 |
| BERT (en) | 0.94 | 0.94 | 0.97 |
| FCNN (en) | 0.92 | 0.92 | 0.96 |
| Logistic Regression (en) | 0.91 | 0.91 | 0.95 |
| Category | Total | Therapist | Client |
|---|---|---|---|
| Cases (cn) | 75 | - | - |
| Sessions (cn) | 431 | - | - |
| Turn (cn) | 5776 | 2911 | 2865 |
| Avg.Utts./Case (cn) | 77.01 | 38.81 | 38.2 |
| Avg.Utts./Session (cn) | 13.4 | 6.75 | 6.65 |
| Cases (en) | 114 | - | - |
| Sessions (en) | 520 | - | - |
| Turn (en) | 6208 | 3097 | 3111 |
| Avg.Utts./Case (en) | 54.46 | 27.15 | 27.31 |
| Avg.Utts./Session (en) | 11.94 | 5.95 | 5.99 |
| Model | MAE [95% CI] | RMSE [95% CI] | Pearson’s r (p < 0.05) | ICC |
|---|---|---|---|---|
| XGBoost (cn) | 0.51 [0.42, 0.61] | 0.61 [0.48, 0.74] | 0.71 | 0.68 |
| BERT-RNN (cn) | 0.59 [0.44, 0.75] | 0.79 [0.59, 0.99] | 0.50 | 0.47 |
| ElasticNet (cn) | 0.54 [0.42, 0.68] | 0.71 [0.53, 0.87] | 0.58 | 0.57 |
| TF-IDF_Ridge Regression (cn) † | 0.67 [0.53, 0.81] | 0.83 [0.64, 1.02] | 0.32 | 0.25 |
| SVR (cn) | 0.63 [0.49, 0.77] | 0.80 [0.61, 0.97] | 0.38 | 0.26 |
| GPT-4o-mini (cn) † | 1.14 [0.94, 1.37] | 1.35 [1.12, 1.56] | 0.45 | 0.25 |
| XGBoost (en) | 0.45 [0.33, 0.58] | 0.67 [0.44, 0.86] | 0.76 | 0.71 |
| BERT-RNN (en) † | 0.67 [0.51, 0.83] | 0.89 [0.69, 1.07] | 0.55 | 0.52 |
| ElasticNet (en) | 0.53 [0.39, 0.69] | 0.77 [0.52, 1.00] | 0.67 | 0.66 |
| TF-IDF_Ridge Regression (en) † | 0.64 [0.49, 0.82] | 0.90 [0.62, 1.16] | 0.54 | 0.46 |
| SVR (en) † | 0.72 [0.55, 0.89] | 0.95 [0.68, 1.20] | 0.38 | 0.18 |
| GPT-4o-mini (en) † | 0.96 [0.80, 1.12] | 1.13 [0.95, 1.28] | 0.74 | 0.57 |
| Feature | MAE [95% CI] | RMSE [95% CI] | Pearson’s r (p < 0.05) | ICC |
|---|---|---|---|---|
| Semantic Only (cn) | 0.52 [0.40, 0.65] | 0.67 [0.52, 0.82] | 0.65 | 0.56 |
| Ling Only Full (cn) | 0.54 [0.41, 0.69] | 0.70 [0.49, 0.90] | 0.63 | 0.60 |
| Ling Only Selected (cn) | 0.52 [0.38, 0.66] | 0.68 [0.48, 0.88] | 0.63 | 0.57 |
| Hybrid (cn) | 0.51 [0.41, 0.62] | 0.61 [0.48, 0.74] | 0.71 | 0.68 |
| Semantic Only (en) † | 0.52 [0.40, 0.66] | 0.74 [0.53, 0.94] | 0.70 | 0.61 |
| Ling Only Full (en) | 0.51 [0.39, 0.63] | 0.69 [0.52, 0.85] | 0.73 | 0.69 |
| Ling Only Selected (en) | 0.49 [0.38, 0.62] | 0.68 [0.49, 0.84] | 0.76 | 0.69 |
| Hybrid (en) | 0.45 [0.33, 0.58] | 0.67 [0.44, 0.86] | 0.76 | 0.71 |
| Rank | Feature | MI |
|---|---|---|
| 1 | client_lsm_Pronoun_ratio_sum | 0.221 |
| 2 | client_lsm_Verb_ratio_sum | 0.214 |
| 3 | client_I_pron_ratio_sum | 0.191 |
| 4 | client_lsm_Noun_ratio_mean | 0.178 |
| 5 | client_non_fluency_ratio_std | 0.177 |
| 6 | client_lsm_Intj_ratio_sum | 0.170 |
| 7 | therapist_lsm_Verb_ratio_sum | 0.165 |
| 8 | client_I_pron_ratio_std | 0.150 |
| 9 | client_lsm_Negation_ratio_mean | 0.138 |
| 10 | client_lsm_Adv_ratio_sum | 0.138 |
| Rank | Feature | MI |
|---|---|---|
| 1 | therapist_lsm_pronoun_ratio_sum | 0.269 |
| 2 | client_lsm_verb_ratio_sum | 0.191 |
| 3 | therapist_lsm_verb_ratio_sum | 0.176 |
| 4 | client_lsm_pronoun_ratio_sum | 0.169 |
| 5 | therapist_lsm_adj_ratio_std | 0.167 |
| 6 | client_lsm_adp_ratio_sum | 0.165 |
| 7 | client_lsm_conj_ratio_sum | 0.160 |
| 8 | therapist_lsm_adv_ratio_std | 0.156 |
| 9 | therapist_lsm_adv_ratio_sum | 0.142 |
| 10 | client_lsm_verb_ratio_std | 0.140 |
| Aspect | Psy-Insight | PsyDial |
|---|---|---|
| Data Source | Books and blogs | Real counseling platform |
| Language | Chinese and English | Chinese only |
| Scale | 431 CN sessions; 520 EN sessions | 300 sampled sessions |
| Dialogue structure | Shorter session dialogues | Long-term dialogues; 37.8 turns/session |
| Model | MAE [95% CI] | RMSE [95% CI] | Pearson’s r (p < 0.05) | Spearman ρ (p < 0.05) |
|---|---|---|---|---|
| XGBoost | 0.38 [0.35, 0.41] | 0.46 [0.43, 0.50] | 0.35 | 0.35 |
| BERT-RNN | 0.40 [0.36, 0.43] | 0.51 [0.47, 0.56] | n.s. | n.s. |
| ElasticNet † | 0.79 [0.71, 0.87] | 1.06 [0.96, 1.17] | 0.25 | 0.28 |
| TF-IDF_Ridge Regression † | 0.66 [0.61, 0.69] | 0.75 [0.71, 0.78] | n.s. | n.s. |
| SVR † | 0.85 [0.81, 0.89] | 0.94 [0.90, 0.98] | n.s. | 0.16 |
| GPT-4o-mini † | 0.44 [0.39, 0.49] | 0.59 [0.54, 0.65] | n.s. | n.s. |
| Rater Group | Mean | SD | Min | Max | MAE | Pearson’s r |
|---|---|---|---|---|---|---|
| Trained Students | 5.37 | 0.95 | 1.55 | 7.00 | 0.80 | 0.59 |
| Gemini-2.5-Flash | 4.91 | 1.03 | 1.27 | 7.00 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Li, Y.; Sun, N.; Mai, Z.; Li, D.; Fu, G.; Yang, X. Automated Working Alliance Assessment in Psychological Counseling Using Gemini and XGBoost. Entropy 2026, 28, 699. https://doi.org/10.3390/e28060699
Li Y, Sun N, Mai Z, Li D, Fu G, Yang X. Automated Working Alliance Assessment in Psychological Counseling Using Gemini and XGBoost. Entropy. 2026; 28(6):699. https://doi.org/10.3390/e28060699
Chicago/Turabian StyleLi, Yuexi, Ningtao Sun, Zhuoxi Mai, Dalin Li, Guifang Fu, and Xueling Yang. 2026. "Automated Working Alliance Assessment in Psychological Counseling Using Gemini and XGBoost" Entropy 28, no. 6: 699. https://doi.org/10.3390/e28060699
APA StyleLi, Y., Sun, N., Mai, Z., Li, D., Fu, G., & Yang, X. (2026). Automated Working Alliance Assessment in Psychological Counseling Using Gemini and XGBoost. Entropy, 28(6), 699. https://doi.org/10.3390/e28060699

