A Focused Survey of Generative AI-Based Music Therapy Systems: Recent Progress and Open Challenges
Abstract
1. Introduction
- RQ1 (Conceptual perspective): How have generative AI-based music generation techniques been conceptualized and utilized in music therapy contexts, and how have therapeutic considerations, such as emotional or physiological regulation, been discussed in relation to generative mechanisms?
- RQ2 (System-level perspective): How do AI-augmented music therapy systems incorporate common components-including sensing, affective state modeling, generative music modules, and adaptive feedback loops-and how are these components implemented across experimental and real-world settings?
- RQ3 (Future outlook): What major open challenges remain at the intersection of generative music, affect-aware adaptation, and digital health, and what research directions are required to enable scalable and personalized generative AI-based music therapy?
2. Music Therapy and Digital Technology-Based Interventions
2.1. Digital Transformation of Music Therapy Interventions
- Large-scale personalization: GMAI enables the creation of music that reflects individual preferences, cultural background, and therapeutic needs. This capability supports more fine-grained personalization even in contexts where direct therapist involvement is limited.
- Dynamic adaptability: By leveraging real-time physiological and behavioral data-such as heart rate, affective state estimates, or sleep-related indicators-generative systems can produce context-sensitive music that responds to a user’s current condition. This dynamic adaptation has the potential to enhance intervention effectiveness in applications such as stress management, mood regulation, and neurorehabilitation.
- Accessibility and scalability: By reducing reliance on continuous access to trained therapists or pre-curated musical content, AI-based music intervention systems can improve accessibility to music therapy in resource-constrained or remote environments.
2.2. Generative AI for Music
3. System Design Framework for AI-Based Music Therapy
3.1. Multimodal Acquisition Layer
3.1.1. Physiological Signals
3.1.2. Behavioral and Kinematic Signals
3.1.3. Visual and Facial Cues
3.1.4. Acoustic and Speech-Related Features
3.1.5. Self-Report and Interaction-Based Inputs
3.2. Affective Representation and Modeling Layer
3.3. AI-Assisted Music Synthesis Core
3.4. Feedback and Adaptive Optimization Layer
3.5. System Integration and Implementation
4. Current Status and Future Research Directions for Generative-AI Based Music Therapy
4.1. Literature Search and Case Study Selection
4.2. Open Challenges in AI-Based Music Therapy Systems
4.2.1. Advanced Personalization via Biofeedback Loops
4.2.2. Neurorehabilitation and Motor Function Recovery
4.2.3. Human-in-the-Loop and Co-Creative Paradigms
4.2.4. Expansion of Therapeutic Domains and Clinical Validation
4.3. Generative AI-Specific Research Gaps and Opportunities
4.3.1. Goal-Oriented and Therapy-Aware GMAI Models
4.3.2. User Interface, Interaction Design, and Translational Barriers
4.3.3. Digital Twin-Driven Simulation for Generative Music AI Design
5. Conclusions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Agres, K.R.; Schaefer, R.S.; Volk, A.; Van Hooren, S.; Holzapfel, A.; Dalla Bella, S.; Müller, M.; De Witte, M.; Herremans, D.; Ramirez Melendez, R.; et al. Music, computing, and health: A roadmap for the current and future roles of music technology for health care and well-being. Music Sci. 2021, 4, 2059204321997709. [Google Scholar] [CrossRef] [Scilit]
- Zaatar, M.T.; Alhakim, K.; Enayeh, M.; Tamer, R. The transformative power of music: Insights into neuroplasticity, health, and disease. Brain Behav. Immun.-Health 2024, 35, 100716. [Google Scholar] [CrossRef] [Scilit]
- Williams, D.; Hodge, V.J.; Wu, C.Y. On the use of ai for generation of functional music to improve mental health. Front. Artif. Intell. 2020, 3, 497864. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hillecke, T.; Nickel, A.; Bolay, H.V. Scientific perspectives on music therapy. Ann. N. Y. Acad. Sci. 2005, 1060, 271–282. [Google Scholar] [CrossRef] [Scilit]
- Hong, Y.J.; Han, J.; Ryu, H. The effects of synthesizing music using AI for preoperative management of Patients’ anxiety. Appl. Sci. 2022, 12, 8089. [Google Scholar] [CrossRef] [Scilit]
- Rodwin, A.H.; Shimizu, R.; Travis, R., Jr.; James, K.J.; Banya, M.; Munson, M.R. A systematic review of music-based interventions to improve treatment engagement and mental health outcomes for adolescents and young adults. Child Adolesc. Soc. Work J. 2023, 40, 537–566. [Google Scholar] [CrossRef] [Scilit]
- Ridder, H.M.O.; Stige, B.; Qvale, L.G.; Gold, C. Individual music therapy for agitation in dementia: An exploratory randomized controlled trial. Aging Ment. Health 2013, 17, 667–678. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhou, Z.; Zhao, X.; Yang, Q.; Zhou, T.; Feng, Y.; Chen, Y.; Chen, Z.; Deng, C. A randomized controlled trial of the efficacy of music therapy on the social skills of children with autism spectrum disorder. Res. Dev. Disabil. 2025, 158, 104942. [Google Scholar] [CrossRef] [Scilit]
- Feng, Y.; Wang, M. Effect of music therapy on emotional resilience, well-being, and employability: A quantitative investigation of mediation and moderation. BMC Psychol. 2025, 13, 47. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Jiao, D. Advancing personalized digital therapeutics: Integrating music therapy, brainwave entrainment methods, and AI-driven biofeedback. Front. Digit. Health 2025, 7, 1552396. [Google Scholar] [CrossRef] [Scilit]
- Mondanaro, J. Challenges to music therapy programming: A case study of innovation, burden, and resilience in United States hospitals. Music Med. 2019, 11, 115–126. [Google Scholar] [CrossRef] [Scilit]
- Baglione, A.N.; Clemens, M.P.; Maestre, J.F.; Min, A.; Dahl, L.; Shih, P.C. Understanding the technological practices and needs of music therapists. Proc. ACM Hum.-Comput. Interact. 2021, 5, 33. [Google Scholar] [CrossRef] [Scilit]
- Civit, M.; Civit-Masot, J.; Cuadrado, F.; Escalona, M.J. A systematic review of artificial intelligence-based music generation: Scope, applications, and future trends. Expert Syst. Appl. 2022, 209, 118190. [Google Scholar] [CrossRef] [Scilit]
- Dash, A.; Agres, K. AI-based affective music generation systems: A review of methods and challenges. ACM Comput. Surv. 2024, 56, 287. [Google Scholar] [CrossRef] [Scilit]
- Wei, L.; Yu, Y.; Qin, Y.; Zhang, S. From Tools to Creators: A Review on the Development and Application of Artificial Intelligence Music Generation. Information 2025, 16, 656. [Google Scholar] [CrossRef] [Scilit]
- Das, R.; Singh, T.D. Multimodal sentiment analysis: A survey of methods, trends, and challenges. ACM Comput. Surv. 2023, 55, 270. [Google Scholar] [CrossRef] [Scilit]
- Clements-Cortés, A.; Pranjić, M.; Knott, D.; Mercadal-Brotons, M.; Fuller, A.; Kelly, L.; Selvarajah, I.; Vaudreuil, R. International music therapists’ perceptions and experiences in telehealth music therapy provision. Int. J. Environ. Res. Public Health 2023, 20, 5580. [Google Scholar] [CrossRef] [Scilit]
- Copet, J.; Kreuk, F.; Gat, I.; Remez, T.; Kant, D.; Synnaeve, G.; Adi, Y.; Défossez, A. Simple and controllable music generation. Adv. Neural Inf. Process. Syst. 2023, 36, 47704–47720. [Google Scholar]
- Poddar, S.; Wan, Y.; Ivison, H.; Gupta, A.; Jaques, N. Personalizing reinforcement learning from human feedback with variational preference learning. Adv. Neural Inf. Process. Syst. 2024, 37, 52516–52544. [Google Scholar]
- Raglio, A. Applications of technology innovations in music therapy practice. Nord. J. Music Ther. 2025, 34, 31–41. [Google Scholar] [CrossRef] [Scilit]
- Herremans, D.; Chuan, C.H.; Chew, E. A functional taxonomy of music generation systems. ACM Comput. Surv. (CSUR) 2017, 50, 69. [Google Scholar] [CrossRef] [Scilit]
- Liu, C.H.; Ting, C.K. Computational intelligence in music composition: A survey. IEEE Trans. Emerg. Top. Comput. Intell. 2016, 1, 2–15. [Google Scholar] [CrossRef] [Scilit]
- Hernandez-Olivan, C.; Beltran, J.R. Music composition with deep learning: A review. In Advances in Speech and Music Technology: Computational Aspects and Applications; Springer: New York, NY, USA, 2022; pp. 25–50. [Google Scholar]
- Huang, C.Z.A.; Vaswani, A.; Uszkoreit, J.; Shazeer, N.; Simon, I.; Hawthorne, C.; Dai, A.M.; Hoffman, M.D.; Dinculescu, M.; Eck, D. Music transformer. In Proceedings of the International Conference on Learning Representations, New Orleans, LA, USA, 6–9 May 2019. [Google Scholar]
- Van Den Oord, A.; Dieleman, S.; Zen, H.; Simonyan, K.; Vinyals, O.; Graves, A.; Kalchbrenner, N.; Senior, A.; Kavukcuoglu, K. Wavenet: A generative model for raw audio. arXiv 2016, arXiv:1609.03499. [Google Scholar] [CrossRef] [Scilit]
- Donahue, C.; McAuley, J.; Puckette, M. Adversarial audio synthesis. arXiv 2018, arXiv:1802.04208. [Google Scholar]
- Dhariwal, P.; Jun, H.; Payne, C.; Kim, J.W.; Radford, A.; Sutskever, I. Jukebox: A generative model for music. arXiv 2020, arXiv:2005.00341. [Google Scholar] [CrossRef] [Scilit]
- Kong, Z.; Ping, W.; Huang, J.; Zhao, K.; Catanzaro, B. Diffwave: A versatile diffusion model for audio synthesis. arXiv 2020, arXiv:2009.09761. [Google Scholar]
- Saeed, A.; Grangier, D.; Zeghidour, N. Contrastive learning of general-purpose audio representations. In Proceedings of the ICASSP 2021—2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); IEEE: New York, NY, USA, 2021; pp. 3875–3879. [Google Scholar]
- Niizumi, D.; Takeuchi, D.; Ohishi, Y.; Harada, N.; Kashino, K. BYOL for audio: Exploring pre-trained general-purpose audio representations. IEEE/ACM Trans. Audio Speech Lang. Process. 2022, 31, 137–151. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Y.; Xia, G.; Levy, M.; Dixon, S. COSMIC: A conversational interface for human-AI music co-creation. In Proceedings of the International Conference on New Interfaces for Musical Expression, Shanghai, China, 15–18 June 2021. [Google Scholar]
- Lin, W.; Li, C. Review of studies on emotion recognition and judgment based on physiological signals. Appl. Sci. 2023, 13, 2573. [Google Scholar] [CrossRef] [Scilit]
- Mohammed, M.H.; Kadhim, M.N.; Al-Shammary, D.; Ibaida, A. EEG-Based Emotion Detection Using Roberts Similarity and PSO Feature Selection. IEEE Access 2025, 13, 79353–79366. [Google Scholar] [CrossRef] [Scilit]
- Wang, Z.; Zou, Y.; Liu, J.; Peng, W.; Li, M.; Zou, Z. Heart rate variability in mental disorders: An umbrella review of meta-analyses. Transl. Psychiatry 2025, 15, 104. [Google Scholar] [CrossRef] [Scilit]
- Kim, S.R.; Zhan, Y.; Davis, N.; Bellamkonda, S.; Gillan, L.; Hakola, E.; Hiltunen, J.; Javey, A. Electrodermal activity as a proxy for sweat rate monitoring during physical and mental activities. Nat. Electron. 2025, 8, 353–361. [Google Scholar] [CrossRef] [Scilit]
- Al-Nafjan, A.; Aldayel, M. Anxiety detection system based on galvanic skin response signals. Appl. Sci. 2024, 14, 10788. [Google Scholar] [CrossRef] [Scilit]
- Karg, M.; Samadani, A.A.; Gorbet, R.; Kühnlenz, K.; Hoey, J.; Kulić, D. Body movements for affective expression: A survey of automatic recognition and generation. IEEE Trans. Affect. Comput. 2013, 4, 341–359. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Song, W.; Tao, W.; Liotta, A.; Yang, D.; Li, X.; Gao, S.; Sun, Y.; Ge, W.; Zhang, W.; et al. A systematic review on affective computing: Emotion models, databases, and recent advances. Inf. Fusion 2022, 83, 19–52. [Google Scholar] [CrossRef] [Scilit]
- Wani, T.M.; Gunawan, T.S.; Qadri, S.A.A.; Kartiwi, M.; Ambikairajah, E. A comprehensive review of speech emotion recognition systems. IEEE Access 2021, 9, 47795–47814. [Google Scholar] [CrossRef] [Scilit]
- Bradley, M.M.; Lang, P.J. Measuring emotion: The self-assessment manikin and the semantic differential. J. Behav. Ther. Exp. Psychiatry 1994, 25, 49–59. [Google Scholar] [CrossRef] [Scilit]
- Geetha, A.; Mala, T.; Priyanka, D.; Uma, E. Multimodal emotion recognition with deep learning: Advancements, challenges, and future directions. Inf. Fusion 2024, 105, 102218. [Google Scholar]
- Eerola, T.; Vuoskoski, J.K. A comparison of the discrete and dimensional models of emotion in music. Psychol. Music 2011, 39, 18–49. [Google Scholar] [CrossRef] [Scilit]
- Poria, S.; Cambria, E.; Bajpai, R.; Hussain, A. A review of affective computing: From unimodal analysis to multimodal fusion. Inf. Fusion 2017, 37, 98–125. [Google Scholar] [CrossRef] [Scilit]
- Kim, J.; André, E. Emotion recognition based on physiological changes in music listening. IEEE Trans. Pattern Anal. Mach. Intell. 2008, 30, 2067–2083. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wöllmer, M.; Kaiser, M.; Eyben, F.; Schuller, B.; Rigoll, G. LSTM-modeling of continuous emotions in an audiovisual affect recognition framework. Image Vis. Comput. 2013, 31, 153–163. [Google Scholar] [CrossRef] [Scilit]
- Yu, B.; Lu, P.; Wang, R.; Hu, W.; Tan, X.; Ye, W.; Zhang, S.; Qin, T.; Liu, T.Y. Museformer: Transformer with fine-and coarse-grained attention for music generation. Adv. Neural Inf. Process. Syst. 2022, 35, 1376–1388. [Google Scholar]
- Livingstone, S.R.; Muhlberger, R.; Brown, A.R.; Thompson, W.F. Changing musical emotion: A computational rule system for modifying score and performance. Comput. Music J. 2010, 34, 41–64. [Google Scholar] [CrossRef] [Scilit]
- Katahira, K.; Matsuda, Y.T.; Fujimura, T.; Ueno, K.; Asamizuya, T.; Suzuki, C.; Cheng, K.; Okanoya, K.; Okada, M. Neural basis of decision making guided by emotional outcomes. J. Neurophysiol. 2015, 113, 3056–3068. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Shakya, A.K.; Pillai, G.; Chakrabarty, S. Reinforcement learning algorithms: A brief survey. Expert Syst. Appl. 2023, 231, 120495. [Google Scholar] [CrossRef] [Scilit]
- Gil, M.; Pelechano, V.; Fons, J.; Albert, M. Designing the human in the loop of self-adaptive systems. In Proceedings of the International Conference on Ubiquitous Computing and Ambient Intelligence; Springer: New York, NY, USA, 2016; pp. 437–449. [Google Scholar]
- Vincent, E.; Virtanen, T.; Gannot, S. Audio Source Separation and Speech Enhancement; John Wiley & Sons: New York, NY, USA, 2018. [Google Scholar]
- Wang, D.; Chen, J. Supervised speech separation based on deep learning: An overview. IEEE/ACM Trans. Audio Speech Lang. Process. 2018, 26, 1702–1726. [Google Scholar] [CrossRef] [Scilit]
- Mehrabian, A. Pleasure-arousal-dominance: A general framework for describing and measuring individual differences in temperament. Curr. Psychol. 1996, 14, 261–292. [Google Scholar] [CrossRef] [Scilit]
- Ni, T.; Jiang, Y.; Lin, Z.; Ruan, J.; Wang, Y.; Liu, Y.; Han, Z. The results of short-course acoustic test could act as an effective predictor of the efficacy of customized music therapy for chronic tinnitus. Front. Neurosci. 2025, 19, 1544723. [Google Scholar] [CrossRef] [Scilit]
- Greenberg, D.M.; Bodner, E.; Shrira, A.; Fricke, K.R. Decreasing stress through a spatial audio and immersive 3D environment: A pilot study with implications for clinical and medical settings. Music Sci. 2021, 4, 2059204321993992. [Google Scholar] [CrossRef] [Scilit]
- Panteliodi, E.; Hudson, D. A sense of space in the core of the bore: Enhancing the MRI experience through use of spatial audio. Radiography 2024, 30, 1451–1454. [Google Scholar] [CrossRef] [Scilit]
- Bruschi, V.; Generosi, A.; Terenzi, A.; Mengoni, M.; Cecchi, S. A Preliminary Study on the Effect of Spatial Sound Reproduction based on Physiological Responses and Facial Expressions of the Listener. In Proceedings of the 2025 Immersive and 3D Audio: From Architecture to Automotive (I3DA); IEEE: New York, NY, USA, 2025; pp. 1–7. [Google Scholar]
- Chi, T.; Gao, L.; Zhang, Y. STASE: A spatialized text-to-audio synthesis engine for music generation. arXiv 2025, arXiv:2509.11124. [Google Scholar]
- Yang, J.; Barde, A.; Billinghurst, M. Audio augmented reality: A systematic review of technologies, applications, and future research directions. J. Audio Eng. Soc. 2022, 70, 788–809. [Google Scholar] [CrossRef] [Scilit]
- Hssayeni, M.D.; Ghoraani, B. Multi-modal physiological data fusion for affect estimation using deep learning. IEEE Access 2021, 9, 21642–21652. [Google Scholar] [CrossRef] [Scilit]
- Qu, W.; Wang, M.J.S. Structured and Factorized Multi-Modal Representation Learning for Physiological Affective State and Music Preference Inference. Symmetry 2026, 18, 488. [Google Scholar] [CrossRef] [Scilit]
- Suno, Inc. SUNO Music Generation API. 2024. Available online: https://suno.com (accessed on 18 January 2026).
- Le, D.V.T.; Bigo, L.; Herremans, D.; Keller, M. Natural language processing methods for symbolic music generation and information retrieval: A survey. ACM Comput. Surv. 2025, 57, 175. [Google Scholar] [CrossRef] [Scilit]
- Pandey, A.; Singh, J.; Kaur, M. Bridging Text and Speech for Emotion Understanding: An Explainable Multimodal Transformer Fusion Framework with Unified Audio–Text Attribution. J. Intell. 2025, 13, 159. [Google Scholar] [CrossRef] [Scilit]
- Naik, P.S.; Ranjan, H.; Uma, D. Cross-Modal Emotion-Aware Music Generation for Therapeutic Healing Using LoRA-Tuned Transformers. In Proceedings of the 2025 20th International Joint Symposium on Artificial Intelligence and Natural Language Processing (iSAI-NLP); IEEE: New York, NY, USA, 2025; pp. 1–6. [Google Scholar]
- Li, S.; Ji, S.; Wang, Z.; Wu, S.; Yu, J.; Zhang, K. A survey on music generation from single-modal, cross-modal, and multi-modal perspectives. ACM Comput. Surv. 2025, 58, 279. [Google Scholar] [CrossRef] [Scilit]
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning transferable visual models from natural language supervision. Proc. Int. Conf. Mach. Learn. 2021, 139, 8748–8763. [Google Scholar]
- Elizalde, B.; Deshmukh, S.; Al Ismail, M.; Wang, H. Clap learning audio concepts from natural language supervision. In Proceedings of the ICASSP 2023—2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); IEEE: New York, NY, USA, 2023; pp. 1–5. [Google Scholar]
- Liu, X.; Zhang, Y.; Yan, Z.; Ge, Y. Defining ‘seamlessly connected’: User perceptions of operation latency in cross-device interaction. Int. J. Hum.-Comput. Stud. 2023, 177, 103068. [Google Scholar] [CrossRef] [Scilit]
- Fang-Yi Tan, F.; Nov, O. Counting the Wait: Effects of Temporal Feedback on Downstream Task Performance and Perceived Wait-Time Experience during System-Imposed Delays. arXiv 2026, arXiv:2602.04138. [Google Scholar]
- Lam, M.W.; Tian, Q.; Li, T.; Yin, Z.; Feng, S.; Tu, M.; Ji, Y.; Xia, R.; Ma, M.; Song, X.; et al. Efficient neural music generation. Adv. Neural Inf. Process. Syst. 2023, 36, 17450–17463. [Google Scholar]
- Caspe, F.; Shier, J.; Sandler, M.; Saitis, C.; McPherson, A. Designing neural synthesizers for low-latency interaction. arXiv 2025, arXiv:2503.11562. [Google Scholar] [CrossRef] [Scilit]
- Fan, X.; Zou, W.; Moghimi, M.; Yi, W. Integrating AI and cloud-edge technologies for music creation in educational and performance domains. J. Cloud Comput. 2026, 15, 32. [Google Scholar] [CrossRef] [Scilit]
- Wei, X.; Zhang, Z.; Yue, Z.; Chen, H.T. Context-AI Tunes: Context-Aware AI-Generated Music for Stress Reduction. In Proceedings of the International Conference on Human-Computer Interaction; Springer: New York, NY, USA, 2025; pp. 330–345. [Google Scholar]
- Shen, L.; Zhang, H.; Zhu, C.; Li, R.; Qian, K.; Meng, W.; Tian, F.; Hu, B.; Schuller, B.W.; Yamamoto, Y. A first look at generative artificial intelligence based music therapy for mental disorders. IEEE Trans. Consum. Electron. 2025, 71, 7439–7453. [Google Scholar] [CrossRef] [Scilit]
- Shen, L.; Zhang, H.; Zhu, C.; Li, R.; Qian, K.; Tian, F.; Hu, B.; Schuller, B.W.; Yamamoto, Y. Enhancing emotion regulation in mental disorder treatment: An aigc-based closed-loop music intervention system. IEEE Trans. Affect. Comput. 2025, 16, 2245–2260. [Google Scholar] [CrossRef] [Scilit]
- Venkatachalam, N. MusiciAI: A Hybrid Generative Model for Music Therapy using Cross-Modal Transformer and Variational Autoencoder. In Proceedings of the 2024 2nd International Conference on Intelligent Data Communication Technologies and Internet of Things (IDCIoT); IEEE: New York, NY, USA, 2024; pp. 1176–1180. [Google Scholar]
- Nikolakakis, E.; Ching, J.; Karystinaios, E.; Sipin, G.; Widmer, G.; Marinescu, R. Language Models for Music Medicine Generation. In Proceedings of the 2024 International Society for Music Information Retrieval Conference, San Francisco, CA, USA, 10–14 November 2024. [Google Scholar]
- Agres, K.R.; Dash, A.; Chua, P. AffectMachine-Classical: A novel system for generating affective classical music. Front. Psychol. 2023, 14, 1158172. [Google Scholar] [CrossRef] [Scilit]
- Braun Janzen, T.; Koshimori, Y.; Richard, N.M.; Thaut, M.H. Rhythm and music-based interventions in motor rehabilitation: Current evidence and future perspectives. Front. Hum. Neurosci. 2022, 15, 789467. [Google Scholar] [CrossRef] [Scilit]
- Sun, J.; Yang, J.; Zhou, G.; Jin, Y.; Gong, J. Understanding human-AI collaboration in music therapy through co-design with therapists. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, Honolulu, HI, USA, 11–16 May 2024; pp. 1–21. [Google Scholar]
- Qiu, Z.; Yuan, R.; Xue, W.; Jin, Y. Generated Therapeutic Music Based on the ISO Principle. In Summit on Music Intelligence; Springer: New York, NY, USA, 2023; pp. 32–45. [Google Scholar]
- Bao, J.; Lyu, Y.; Yang, J.; Jin, Y.; Gong, J. How Generative Music Affects the ISO Principle-Based Emotion-Focused Therapy: An EEG Study. In Proceedings of the Annual Meeting of the Cognitive Science Society, San Francisco, CA, USA, 30 July–2 August 2025; Volume 47. [Google Scholar]
- Jin, Y.; Cai, W.; Chen, L.; Zhang, Y.; Doherty, G.; Jiang, T. Exploring the design of generative AI in supporting music-based reminiscence for older adults. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, Honolulu, HI, USA, 11–16 May 2024; pp. 1–17. [Google Scholar]
- Low, B.; Liu, X.; Li, R.Z.; Ren, E.; Zhang, J.X. Music therapy for autism Spectrum disorder: A comprehensive literature review on therapeutic efficacy, limitations, and AI integration. In Proceedings of the 2024 IEEE 15th Annual Ubiquitous Computing, Electronics & Mobile Communication Conference (UEMCON); IEEE: New York, NY, USA, 2024; pp. 90–99. [Google Scholar]
- Yang, W.; Huang, C.F.; Huang, H.Y.; Zhang, Z.; Li, W.; Wang, C. Research on the improvement of children’s attention through binaural beats music therapy in the context of ai music generation. In Summit on Music Intelligence; Springer: New York, NY, USA, 2023; pp. 19–31. [Google Scholar]
- Robb, S.L.; Hanson-Abromeit, D.; May, L.; Hernandez-Ruiz, E.; Allison, M.; Beloat, A.; Daugherty, S.; Kurtz, R.; Ott, A.; Oyedele, O.O.; et al. Reporting quality of music intervention research in healthcare: A systematic review. Complement. Ther. Med. 2018, 38, 24–41. [Google Scholar] [CrossRef] [Scilit]
- Lacson, C.; Myers-Coffman, K.; Kesslick, A.; Krater, C.; Bradt, J. Conducting Clinical Studies in Community Health Settings: Challenges and Opportunities for Music Therapists. Music Ther. Perspect. 2021, 39, 105–112. [Google Scholar] [CrossRef] [Scilit]
- Rodgers-Melnick, S.N.; Block, S.; Rivard, R.L.; Dusek, J.A. Optimizing patient-reported outcome collection and documentation in medical music therapy: Process-improvement study. JMIR Hum. Factors 2023, 10, e46528. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Rodgers-Melnick, S.N.; Rivard, R.L.; Block, S.; Dusek, J.A. Effectiveness of medical music therapy practice: Integrative research using the electronic health record: Rationale, design, and population characteristics. J. Integr. Complement. Med. 2024, 30, 57–65. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Baltaxe-Admony, L.B.; Hope, T.; Watanabe, K.; Teodorescu, M.; Kurniawan, S.; Nishimura, T. Exploring the creation of useful interfaces for music therapists. In Proceedings of the Audio Mostly 2018 on Sound in Immersion and Emotion; ACM: New York, NY, USA, 2018; pp. 1–7. [Google Scholar]
- Kader, F.B.; Karmaker, S. A survey on evaluation metrics for music generation. arXiv 2025, arXiv:2509.00051. [Google Scholar]
- Spiro, N.; Tsiris, G.; Cripps, C. A systematic review of outcome measures in music therapy. Music Ther. Perspect. 2018, 36, 67–78. [Google Scholar] [CrossRef] [Scilit]
- Sabbatella, P.E. Assessment and clinical evaluation in music therapy: An overview from literature and clinical practice. Music Ther. Today 2004, 5, 1–32. [Google Scholar]
- Zhao, F.; Sun, Z.; Niu, W. Effect of ward noise reduction technology combined with music therapy on negative emotions in inpatients undergoing gastric cancer radiotherapy: A retrospective study. Noise Health 2023, 25, 257–263. [Google Scholar] [CrossRef] [Scilit]
- Zhang, X.; Zheng, B.; Lu, D.; Qi, W.; Lin, L.; He, S. Retrospective Analysis of the Influence of Music Therapy Combined with Noise Reduction Technology in Dental Implant Patients. Noise Health 2025, 27, 785–793. [Google Scholar] [CrossRef] [Scilit]
- Dallı, Ö.E.; Yıldırım, Y.; Aykar, F.Ş.; Kahveci, F. The effect of music on delirium, pain, sedation and anxiety in patients receiving mechanical ventilation in the intensive care unit. Intensive Crit. Care Nurs. 2023, 75, 103348. [Google Scholar] [CrossRef] [Scilit]
- Witek, S.; Schmoor, C.; Montigel, F.; Grotejohann, B.; Ziegler, S. Sustainable reduction in sound levels on intensive care units through noise management-an implementation study. BMC Health Serv. Res. 2025, 25, 9. [Google Scholar] [CrossRef] [Scilit]
- Elgammal, Z.; Albrijawi, M.T.; Alhajj, R. Digital twins in healthcare: A review of AI-powered practical applications across health domains. J. Big Data 2025, 12, 234. [Google Scholar] [CrossRef] [Scilit]
- Khoshfekr Rudsari, H.; Tseng, B.; Zhu, H.; Song, L.; Gu, C.; Roy, A.; Irajizad, E.; Butner, J.; Long, J.; Do, K.A. Digital twins in healthcare: A comprehensive review and future directions. Front. Digit. Health 2025, 7, 1633539. [Google Scholar] [CrossRef] [Scilit]
- Yan, R.; Shen, X.; Wachi, A.; Gros, S.; Zhao, A.; Hu, X. Offline Guarded Safe Reinforcement Learning for Medical Treatment Optimization Strategies. arXiv 2025, arXiv:2505.16242. [Google Scholar] [CrossRef] [Scilit]
- Lauer-Schmaltz, M.W.; Cash, P.; Hansen, J.P.; Das, N. Human digital twins in rehabilitation: A case study on exoskeleton and serious-game-based stroke rehabilitation using the ETHICA methodology. IEEE Access 2024, 12, 180968–180991. [Google Scholar] [CrossRef] [Scilit]


| Functional Layer | Primary Role | Typical Inputs/Methods | Studies |
|---|---|---|---|
| Multimodal acquisition layer | Capture observable signals reflecting user emotional, physiological, and behavioral states | EEG, HR/HRV, EDA, respiration; motion and posture; facial expression and gaze; speech and acoustic features; self-report measures | [32,33,34,35,36,37,38,39,40] |
| Affective representation and modeling layer | Fuse heterogeneous inputs into unified, machine-interpretable affective state representations | Cross-modal feature alignment; joint latent embeddings; arousal–valence or PAD models; temporal smoothing and state estimation | [41,42,43,44,45] |
| AI-assisted music synthesis core | Generate or control music content conditioned on affective state and therapeutic intent | Symbolic generation (transformer-based); audio-domain generation (diffusion models); affect-to-music parameter mapping; safety-constrained generation | [13,14,18,23,24,27,31,46] |
| Feedback and adaptive optimization layer | Close the loop between generated music and user response to enable dynamic adaptation | Rule-based control; reward modeling from affective change; reinforcement learning; online optimization; human-in-the-loop intervention | [47,48,49,50] |
| Case | Input Modality | Affective Modeling | Generation & Adaptation | Architectural Emphasis |
|---|---|---|---|---|
| S1 [74] | Environmental video | Implicit affect and contextual cues extracted via a visual-language model; mediated through user or therapist prompts | Commercial text-to-music system (Suno API) with manual prompt refinement; no automated closed-loop adaptation | Prompt-mediated contextual control without explicit physiological affect modeling |
| S2 [75] | Real-time EEG | Latent affect embeddings from a pretrained EEG encoder used as continuous conditioning variables | Affect-conditioned music generation integrated with reinforcement learning-based adaptive control and selective user intervention | EEG-conditioned closed-loop generation emphasizing neural affect modeling and adaptation |
| S3 [76] | EEG | Discrete or continuous affect labels inferred from EEG-based emotion recognition | Pretrained audio generative models (HiFiGAN + VAE); no explicit adaptive feedback | Feasibility-driven EEG-audio integration without closed-loop optimization |
| S4 [77] | Text prompts and symbolic or latent music | Implicit affective descriptors embedded via cross-modal attention | Hybrid cross-modal transformer and VAE; feedback conceptually acknowledged but weakly specified | Cross-modal latent alignment prioritizing expressive diversity over adaptive therapy |
| S5 [78] | Iterative text prompts with partial audio | Affective intent encoded through iterative textual descriptors updated across generations | MusicGen guided by prompt engineering and generation history; adaptation across iterations | Iterative affect-aware prompting with temporal refinement of affect alignment |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the author. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Seo, J.S. A Focused Survey of Generative AI-Based Music Therapy Systems: Recent Progress and Open Challenges. Appl. Sci. 2026, 16, 4120. https://doi.org/10.3390/app16094120
Seo JS. A Focused Survey of Generative AI-Based Music Therapy Systems: Recent Progress and Open Challenges. Applied Sciences. 2026; 16(9):4120. https://doi.org/10.3390/app16094120
Chicago/Turabian StyleSeo, Jin S. 2026. "A Focused Survey of Generative AI-Based Music Therapy Systems: Recent Progress and Open Challenges" Applied Sciences 16, no. 9: 4120. https://doi.org/10.3390/app16094120
APA StyleSeo, J. S. (2026). A Focused Survey of Generative AI-Based Music Therapy Systems: Recent Progress and Open Challenges. Applied Sciences, 16(9), 4120. https://doi.org/10.3390/app16094120
