Application of Machine Learning and Natural Language Processing Techniques for the Analysis of Surveys with Open-Ended Questions: A Scoping Review
Abstract
1. Introduction
2. Materials and Methods
2.1. Research Questions
- RQ1. What types of machine learning are most used for data analysis in surveys with open-ended questions?
- RQ2. What are the most common ML and NLP tasks for analyzing surveys with open-ended questions?
- RQ3. What are the main ML algorithms applied for analyzing surveys with open-ended questions?
- RQ4. What are the most popular NLP models and architectures for analyzing surveys with open-ended questions?
- RQ5. What technologies are frequently employed for analyzing surveys with open-ended questions?
- RQ6. In which areas of research are surveys with open-ended questions predominant?
2.2. Inclusion and Exclusion Criteria
- ‘Analysis’ AND ‘Machine learning’ AND ‘Survey’ AND (‘open-ended question’ OR ‘open-ended’) AND ‘Natural language processing’. Analysis of the preliminary results of this query revealed relevant search terms related to different applications of machine learning in survey analysis. Query 2 includes these search terms to broaden the identified relationship.
- (‘Machine learning’ OR ‘Supervised descriptive rule discovery’ OR ‘SDRD’ OR ‘Natural language processing’ OR ‘NLP’) AND (open-ended question’ OR ‘open-ended’ OR ‘free text questions’ OR ‘open-ended answers’ OR ‘open-ended survey’).
2.3. Selection and Eligibility of Studies
- Relevant studies were identified through an exhaustive search of five digital databases (IEEE, ACM, Springer, ScienceDirect, Wiley) and Google Scholar, yielding 29,288 initial records.
- Following the PRISMA flowchart, the selection was conducted in three phases. First, a manual review of titles from 16,808 records; next, a review of abstracts and keywords from 5111 records; followed by an evaluation of the full text of 96 reports to ensure eligibility.
- Studies that did not use at least one search criterion were not directly related to the application of surveys, or were theses of any level were excluded, ultimately yielding a total of 79 relevant studies.
- Finally, the results were compiled, summarized, and presented in graphical and tabular formats.
2.4. Data Collection and Analysis
3. Results
Results of the State-of-the-Art Analysis
| Article | SL | UL | SDRD | OQ | NLP | OC | SC |
|---|---|---|---|---|---|---|---|
| Singer and Couper [7] | NO | NO | NO | YES | YES | YES | NS |
| Kosmajac et al. [8] | NO | YES | NO | YES | YES | YES | 50,792 |
| Baburajan et al. [9] | YES | YES | NO | YES | YES | YES | 364 |
| Bardutz and Bigazzi [10] | NO | YES | NO | YES | YES | YES | 366 |
| Moreo, Esuli, and Sebastiani [11] | YES | NO | NO | YES | YES | YES | 201–10,788 |
| Schonlau and Couper [12] | NO | NO | NO | YES | YES | YES | 1212 and 1758 |
| Espinoza et al. [13] | NO | YES | NO | YES | YES | YES | 9800–20,000 |
| Gweon and Schonlau [14] | YES | NO | NO | YES | YES | YES | 585–1212 |
| Olmos-Vallejo et al. [18] | YES | YES | CS and SD | YES | NO | YES | NS |
| Jojoa et al. [23] | YES | NO | NO | YES | YES | YES | 365 |
| Zucco et al. [24] | YES | NO | NO | YES | YES | YES | 169 |
| Jayaratne and Jayatilleke [25] | YES | YES | NO | YES | YES | YES | 46,888 |
| He and Schonlau [26] | YES | YES | NO | YES | YES | NO | 1096–1756 |
| Ríos-Méndez et al. [27] | YES | YES | EPM | YES | NO | YES | 7856 |
| Supraja et al. [28] | YES | YES | NO | YES | YES | YES | 39 |
| Ferrario and Stantcheva [29] | NO | YES | NO | YES | YES | YES | 5144 |
| Spasić et al. [30] | YES | NO | NO | YES | YES | YES | 55 |
| Roberts et al. [31] | YES | YES | NO | YES | YES | NO | 2323 |
| Duan et al. [32] | YES | YES | EPM | YES | YES | YES | 359–1091 |
| Machorro-Cano et al. [33] | YES | YES | EPM | YES | NO | YES | 7856 |
| Olmos-Vallejo et al. [34] | YES | YES | SD | YES | NO | YES | 289 and 47,093 |
| Koufakou [36] | YES | YES | NO | YES | YES | NO | 10,610 |
| Haensch et al. [37] | YES | NO | NO | YES | YES | YES | 5000 |
| Jacennik et al. [38] | YES | NO | NO | YES | YES | YES | 104 |
| van Buchem et al. [39] | YES | YES | NO | YES | YES | YES | 534 |
| Nanda et al. [40] | NO | YES | NO | YES | YES | YES | 130,500–158,000 |
| Rubio Delgado et al. [41] | YES | YES | NO | YES | NO | NO | 86 |
| Schonlau et al. [42] | YES | NO | NO | YES | YES | YES | 1006–2350 |
| Smoll et al. [43] | NO | YES | NO | YES | YES | YES | 723 |
| Wang et al. [44] | YES | YES | NO | YES | YES | YES | 402 |
| Yawson et al. [45] | NO | NO | NO | NO | YES | YES | 119 |
| Pestian et al. [46] | NO | NO | NO | YES | YES | YES | 30 |
| Kjell et al. [47] | YES | YES | NO | YES | YES | YES | 92–854 |
| Robinson et al. [48] | YES | NO | NO | YES | YES | YES | 9862 |
| Tvinnereim and Fløttum [49] | YES | YES | NO | YES | YES | YES | 4634 |
| Maramba et al. [50] | YES | NO | NO | YES | YES | NO | 3426 |
| Guetterman et al. [51] | NO | YES | NO | YES | YES | NO | 58 and 68 |
| Hahn et al. [52] | NO | YES | NO | YES | YES | YES | 3183 |
| McGillivray et al. [53] | NO | YES | NO | YES | YES | NO | 599 |
| Nawaz et al. [54] | YES | YES | NO | YES | YES | YES | 4400 |
| Popping [55] | NO | NO | NO | YES | YES | NO | NS |
| Liu et al. [56] | YES | YES | NO | YES | NO | YES | 1448 |
| Jaeger and Rasmussen [57] | YES | NO | NO | YES | YES | YES | 4341 |
| Küfner et al. [58] | YES | NO | NO | YES | NO | YES | 674,138 |
| Mourtgos and Adams [59] | NO | YES | NO | YES | YES | YES | 396 |
| Yamano et al. [60] | NO | YES | NO | YES | YES | YES | 107 |
| Lasri et al. [61] | YES | YES | NO | YES | NO | YES | 4388 |
| Linton et al. [62] | NO | NO | NO | YES | YES | YES | 5634 and 59,768 |
| González Canché [63] | YES | YES | NO | YES | YES | YES | 10,684 |
| Marengo et al. [64] | YES | YES | NO | YES | NO | YES | 5.048 |
| Buenano-Fernández et al. [65] | NO | YES | NO | YES | YES | YES | 900 |
| Pietsch and Lessmann [66] | YES | YES | NO | YES | YES | YES | 5001 |
| Baumer et al. [67] | NO | YES | NO | YES | YES | NO | 1095 |
| Shiwaku et al. [68] | NO | NO | NO | YES | YES | YES | 1.978 |
| Koufakou et al. [69] | YES | YES | NO | YES | YES | YES | 204 |
| Etz et al. [70] | NO | YES | NO | YES | YES | NO | 320,500 |
| Costales et al. [71] | NO | YES | NO | YES | YES | YES | 65 |
| Zhang et al. [72] | NO | YES | NO | YES | YES | YES | 168 |
| Suadaa et al. [73] | YES | NO | NO | YES | YES | NO | 19,944 |
| Grönberg et al. [74] | NO | YES | NO | YES | YES | YES | 742 and 6087 |
| Singha and Mohapatra [75] | YES | YES | NO | YES | NO | YES | 170 |
| Mahendher et al. [76] | YES | NO | NO | YES | NO | YES | 153 |
| Kazi et al. [77] | YES | YES | NO | YES | YES | YES | 124 |
| Alshaikh et al. [78] | YES | NO | NO | YES | YES | YES | 1.506 |
| Wang [79] | YES | NO | NO | YES | NO | YES | 1019 |
| Pestian et al. [80] | YES | NO | NO | YES | YES | YES | 371 |
| Clark et al. [81] | YES | YES | NO | YES | YES | YES | 4.068 |
| Li et al. [82] | YES | YES | NO | YES | YES | YES | 3336 |
| Torrao et al. [83] | YES | YES | NO | YES | YES | YES | 41 |
| Kobra et al. [84] | YES | NO | NO | YES | YES | YES | 400 |
| Inoue et al. [85] | YES | YES | NO | YES | YES | YES | 201 |
| Nikulchev et al. [86] | NO | YES | NO | YES | YES | YES | 20,443 |
| Cook et al. [87] | YES | NO | NO | YES | YES | YES | 1453 |
| Onan [88] | YES | YES | NO | YES | YES | YES | 154,000 |
| Wijngaards et al. [89] | YES | NO | NO | YES | YES | YES | 1122 |
| Robin et al. [90] | YES | YES | NO | YES | YES | YES | 13,705 |
| Tvinnereim and Fløttum [91] | YES | YES | NO | YES | YES | YES | 2115 |
| Gish et al. [92] | NO | YES | NO | YES | YES | YES | 510 |
| Cheese et al. [93] | YES | YES | NO | YES | YES | YES | 98 |
4. Discussion
4.1. Question 1: What Types of Machine Learning Are Most Used for Data Analysis in Surveys with Open-Ended Questions?
4.2. Question 2: What Are the Most Common ML and Natural Language Processing Tasks for Analyzing Surveys with Open-Ended Questions?
4.3. Question 3: What Are the Main ML Algorithms Applied for Analyzing Surveys with Open-Ended Questions?
4.4. Question 4: What Are the Most Popular NLP Models and Algorithms for Analyzing Surveys with Open-Ended Questions?
4.5. Question 5: What Are the Most Frequently Employed Technologies for Analyzing Surveys with Open-Ended Questions?
4.6. Question 6: In Which Areas of Research Are Surveys with Open-Ended Questions Predominant?
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Russell, S.J.; Norvig, P. Artificial Intelligence: A Modern Approach, 4th ed.; Pearson: Boston, MA, USA, 2022. [Google Scholar]
- Mahesh, B. Machine Learning Algorithms—A Review. Int. J. Sci. Res. 2020, 9, 381–386. [Google Scholar] [CrossRef] [Scilit]
- Zhu, X.; Goldberg, A.B. Introduction to Semi-Supervised Learning; Springer International Publishing: Cham, Switzerland, 2009. [Google Scholar]
- Carmona, C.J.; Elizondo, D. Supervised Descriptive Rule Discovery: A Survey of the State-of-the-Art. SIMIDAT, 2016. Available online: https://simidat.ujaen.es/sites/default/files/2024-05/TR2016.pdf (accessed on 10 May 2026).
- Garcıa, A.M.; Carmona, C.J.; Gonzalez, P. Minería de Patrones Emergentes: Una Oportunidad Para la Extracción Evolutiva de Conocimiento. SIMIDAT, 2016. Available online: https://simidat.ujaen.es/sites/default/files/2024-05/2016-Garcia-EPMReview.pdf (accessed on 10 May 2026).
- Compilación, C. Historia de las encuestas en el mundo. La Sociol. En Sus Escen. 2011, 24. Available online: https://revistas.udea.edu.co/index.php/ceo/article/view/10988 (accessed on 10 May 2026).
- Singer, E.; Couper, M.P. Some Methodological Uses of Responses to Open Questions and Other Verbatim Comments in Quantitative Surveys. Methods Data Anal. 2017, 11, 115–134. [Google Scholar]
- Kosmajac, D.; Smith, K.; Keselj, V.; Kirkland, S. Graph-based Topic Extraction Using Centroid Distance of Phrase Embeddings on Healthy Aging Open-ended Survey Questions. In Proceedings of the 2020 International Conference on Data Mining Workshops (ICDMW), Sorrento, Italy, 17–20 November 2020; pp. 621–628. [Google Scholar]
- Baburajan, V.; Silva, J.D.A.E.; Pereira, F.C. Open-Ended Versus Closed-Ended Responses: A Comparison Study Using Topic Modeling and Factor Analysis. IEEE Trans. Intell. Transp. Syst. 2021, 22, 2123–2132. [Google Scholar] [CrossRef] [Scilit]
- Bardutz, E.; Bigazzi, A. Communicating perceptions of pedestrian comfort and safety: Structural topic modeling of open response survey comments. Transp. Res. Interdiscip. Perspect. 2022, 14, 100600. [Google Scholar] [CrossRef] [Scilit]
- Moreo, A.; Esuli, A.; Sebastiani, F. Building automated survey coders via interactive machine learning. Int. J. Mark. Res. 2019, 61, 408–429. [Google Scholar] [CrossRef] [Scilit]
- Schonlau, M.; Couper, M.P. Semi-automated categorization of open-ended questions. Surv. Res. Methods 2016, 10, 143–152. [Google Scholar]
- Espinoza, F.; Hamfors, O.; Karlgren, J.; Olsson, F.; Persson, P.; Hamberg, L.; Sahlgren, M. Analysis of Open Answers to Survey Questions through Interactive Clustering and Theme Extraction. In Proceedings of the 2018 Conference on Human Information Interaction & Retrieval (CHIIR ’18), New Brunswick, NJ, USA, 11–15 March 2018; pp. 317–320. [Google Scholar]
- Gweon, H.; Schonlau, M. Automated classification for open-ended questions with BERT. J. Surv. Stat. Methodol. 2024, 12, 493–504. [Google Scholar] [CrossRef] [Scilit]
- Kastrati, Z.; Dalipi, F.; Imran, A.S.; Pireva Nuci, K.; Wani, M.A. Sentiment Analysis of Students’ Feedback with NLP and Deep Learning: A Systematic Mapping Study. Appl. Sci. 2021, 11, 3986. [Google Scholar] [CrossRef] [Scilit]
- Shaik, T.; Tao, X.; Li, Y.; Dann, C.; McDonald, J.; Redmond, P.; Galligan, L. A Review of the Trends and Challenges in Adopting Natural Language Processing Methods for Education Feedback Analysis. IEEE Access 2022, 10, 56720–56739. [Google Scholar] [CrossRef] [Scilit]
- O’bRien, K.K.; Colquhoun, H.; Levac, D.; Baxter, L.; Tricco, A.C.; Straus, S.; Wickerson, L.; Nayar, A.; Moher, D.; O’mAlley, L. Advancing scoping study methodology: A web-based survey and consultation of perceptions on terminology, definition and methodological steps. BMC Health Serv. Res. 2016, 16, 305. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Olmos-Vallejo, A.; Rodríguez-Mazahua, L.; Palet-Guzmán, J.A.; Machorro-Cano, I.; Alor-Hernández, G.; Sánchez-Cervantes, J.L. Application of Supervised Descriptive Rule Discovery Methods: Review and Architecture. In Smart Technologies, Systems and Applications. In SmartTech-IC 2022: Third International Conference on Smart Technologies, Systems and Applications, 1st ed.; Narváez, F.R., Urgilés, F., Bastos-Filho, T.F., Salgado-Guerrreo, J.P., Eds.; Editorial Universitaria Abya-Yala: Quito, Ecuador, 2023; pp. 39–58. [Google Scholar]
- García-Vico, A.M.; Carmona, C.J.; Martín, D.; García-Borroto, M.; Del Jesus, M.J. An overview of emerging pattern mining in supervised descriptive rule discovery: Taxonomy, empirical study, trends, and prospects. WIREs Data Min. Knowl. Discov. 2018, 8, e1231. [Google Scholar] [CrossRef] [Scilit]
- Le Glaz, A.; Haralambous, Y.; Kim-Dufor, D.H.; Lenca, P.; Billot, R.; Ryan, T.C.; Marsh, J.; DeVylder, J.; Walter, M.; Berrouiguet, S.; et al. Machine learning and natural language processing in mental health: Systematic review. J. Med. Internet Res. 2021, 23, e15708. [Google Scholar] [CrossRef] [Scilit]
- Arksey, H.; O’Malley, L. Scoping studies: Towards a methodological framework. Int. J. Soc. Res. Methodol. 2005, 8, 19–32. [Google Scholar] [CrossRef] [Scilit]
- Levac, D.; Colquhoun, H.; O’Brien, K.K. Scoping studies: Advancing the methodology. Implement. Sci. 2010, 5, 69. [Google Scholar] [CrossRef] [Scilit]
- Jojoa, M.; Garcia-Zapirain, B.; Gonzalez, M.J.; Perez-Villa, B.; Urizar, E.; Ponce, S.; Tobar-Blandon, M.F. Analysis of the Effects of Lockdown on Staff and Students at Universities in Spain and Colombia Using Natural Language Processing Techniques. Int. J. Environ. Res. Public Health 2022, 19, 5705. [Google Scholar] [CrossRef] [Scilit]
- Zucco, C.; Paglia, C.; Graziano, S.; Bella, S.; Cannataro, M. Sentiment Analysis and Text Mining of Questionnaires to Support Telemonitoring Programs. Information 2020, 11, 550. [Google Scholar] [CrossRef] [Scilit]
- Jayaratne, M.; Jayatilleke, B. Predicting Personality Using Answers to Open-Ended Interview Questions. IEEE Access 2020, 8, 115345–115355. [Google Scholar] [CrossRef] [Scilit]
- He, Z.; Schonlau, M. Coding Text Answers to Open-ended Questions: Human Coders and Statistical Learning Algorithms Make Similar Mistakes. Methods Data Anal. 2020, 15, 103–120. [Google Scholar]
- Rios-Mendez, I.A.; Rodriguez-Mazahua, L.; Guzman, J.A.P.; Machorro-Cano, I.; Pelaez-Camarena, S.G.; Romero-Torres, C.; Contreras, H.M. Discovering Emerging Patterns from Medical Opinions about the Decrease of Autopsies Performed in a Mexican Hospital. In Proceedings of the 2020 IEEE 16th International Conference on Automation Science and Engineering (CASE), Hong Kong, China, 20–24 August 2020; pp. 798–803. [Google Scholar]
- Supraja, S.; Lim, F.S.; Tan, S.; Ho, S.Y.; Ng, B.K.; Khong, A.W.H. Factors Impacting Students’ Creativity-related Self-efficacy in an Undergraduate Makerspace-based Course. In Proceedings of the 2022 IEEE Global Engineering Education Conference (EDUCON), Tunis, Tunisia, 28–31 March 2022; pp. 513–522. [Google Scholar]
- Ferrario, B.; Stantcheva, S. Eliciting People’s First-Order Concerns: Text Analysis of Open-Ended Survey Questions. Am. Econ. Assoc. Pap. Proc. 2022, 112, 163–169. [Google Scholar]
- Spasić, I.; Owen, D.; Smith, A.; Button, K. KLOSURE: Closing in on open–ended patient questionnaires with text mining. J. Biomed. Semant. 2019, 10, 24. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Roberts, M.E.; Stewart, B.M.; Tingley, D.; Lucas, C.; Leder-Luis, J.; Gadarian, S.K.; Albertson, B.; Rand, D.G. Structural Topic Models for Open-Ended Survey Responses. Am. J. Polit. Sci. 2014, 58, 1064–1082. [Google Scholar] [CrossRef] [Scilit]
- Duan, L.; Liu, L.; Dong, G.; Nummenmaa, J.; Wang, T.; Qin, P.; Yang, H. Mining distinguishing customer focus sets from online customer reviews. Computing 2018, 100, 335–351. [Google Scholar] [CrossRef] [Scilit]
- Machorro-Cano, I.; Ríos-Méndez, I.A.; Palet-Guzmán, J.A.; Rodríguez-Mazahua, N.; Rodríguez-Mazahua, L.; Alor-Hernández, G.; Olmedo-Aguirre, J.O. Medical Opinions Analysis about the Decrease of Autopsies Using Emerging Pattern Mining. Data 2023, 9, 2. [Google Scholar] [CrossRef] [Scilit]
- Olmos-Vallejo, A.; Rodríguez-Mazahua, L.; Palet-Guzmán, J.A.; Machorro-Cano, I.; Alor-Hernández, G.; Cervantes, J. Comparison of Medical Opinions About the Decrease in Autopsies in Mexican Hospitals Using Data Mining. Electronics 2024, 13, 4686. [Google Scholar] [CrossRef] [Scilit]
- Ventura, S.; Luna, J.M. Supervised Descriptive Pattern Mining; Springer: Berlin/Heidelberg, Germany, 2018. [Google Scholar]
- Koufakou, A. Deep learning for opinion mining and topic classification of course reviews. Educ. Inf. Technol. 2024, 29, 2973–2997. [Google Scholar] [CrossRef] [Scilit]
- Haensch, A.C.; Weiß, B.; Steins, P.; Chyrva, P.; Bitz, K. The semi-automatic classification of an open-ended question on panel survey motivation and its application in attrition analysis. Front. Big Data 2022, 5, 880554. [Google Scholar] [CrossRef] [Scilit]
- Jacennik, B.; Zawadzka-Gosk, E.; Moreira, J.P.; Glinkowski, W.M. Evaluating Patients’ Experiences with Healthcare Services: Extracting Domain and Language-Specific Information from Free-Text Narratives. Int. J. Environ. Res. Public Health 2022, 19, 10182. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Van Buchem, M.M.; Neve, O.M.; Kant, I.M.J.; Steyerberg, E.W.; Boosman, H.; Hensen, E.F. Analyzing patient experiences using natural language processing: Development and validation of the artificial intelligence patient reported experience measure (AI-PREM). BMC Med. Inform. Decis. Mak. 2022, 22, 183. [Google Scholar] [CrossRef] [Scilit]
- Nanda, G.; Douglas, K.A.; Waller, D.R.; Merzdorf, H.E.; Goldwasser, D. Analyzing Large Collections of Open-Ended Feedback from MOOC Learners Using LDA Topic Modeling and Qualitative Analysis. IEEE Trans. Learn. Technol. 2021, 14, 146–160. [Google Scholar] [CrossRef] [Scilit]
- Rubio-Delgado, E.; Rodríguez-Mazahua, L.; Palet-Guzmán, J.A.; Cervantes, J.; Sánchez-Cervantes, J.L.; Peláez-Camarena, S.G.; López-Chau, A. Analysis of Medical Opinions about the Nonrealization of Autopsies in a Mexican Hospital Using Association Rules and Bayesian Networks. Sci. Program. 2018, 2018, 1–21. [Google Scholar] [CrossRef] [Scilit]
- Schonlau, M.; Gweon, H.; Wenemark, M. Automatic Classification of Open-Ended Questions: Check-All-That-Apply Questions. Soc. Sci. Comput. Rev. 2021, 39, 562–572. [Google Scholar] [CrossRef] [Scilit]
- Smoll, N.R.; Walker, J.; Khandaker, G. The barriers and enablers to downloading the COVIDSafe app—A topic modelling analysis. Aust. N. Z. J. Public Health 2021, 45, 344–347. [Google Scholar] [CrossRef] [Scilit]
- Wang, Z.; Wang, Q.; Wang, D. Reducing Employees’ Time Theft through Leader’s Developmental Feedback: The Serial Multiple Mediating Effects of Perceived Insider Status and Work Passion. Behav. Sci. 2024, 14, 269. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yawson, A.E.; Tette, E.; Tettey, Y. Through the lens of the clinician: Autopsy services and utilization in a large teaching hospital in Ghana. BMC Res. Notes 2014, 7, 943. [Google Scholar] [CrossRef] [Scilit]
- Pestian, J.P.; Grupp-Phelan, J.; Cohen, K.B.; Meyers, G.; Richey, L.A.; Matykiewicz, P.; Sorter, M.T. A Controlled Trial Using Natural Language Processing to Examine the Language of Suicidal Adolescents in the Emergency Department. Suicide Life-Threat. Behav. 2016, 46, 154–159. [Google Scholar] [CrossRef] [Scilit]
- Kjell, O.N.E.; Kjell, K.; Garcia, D.; Sikström, S. Semantic measures: Using natural language processing to measure, differentiate, and describe psychological constructs. Psychol. Methods 2019, 24, 92–115. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Robinson, C.; Yeomans, M.; Reich, J.; Hulleman, C.; Gehlbach, H. Forecasting student achievement in MOOCs with natural language processing. In Proceedings of the Sixth International Conference on Learning Analytics & Knowledge (LAK ’16), Edinburgh, UK, 25–29 April 2016; pp. 383–387. [Google Scholar]
- Tvinnereim, E.; Fløttum, K.; Gjerstad, Ø.; Johannesson, M.P.; Nordø, Å.D. Citizens’ preferences for tackling climate change. Quantitative and qualitative analyses of their freely formulated solutions. Glob. Environ. Change 2017, 46, 34–41. [Google Scholar] [CrossRef] [Scilit]
- Maramba, I.D.; Davey, A.; Elliott, M.N.; Roberts, M.; Roland, M.; Brown, F.; Burt, J.; Boiko, O.; Campbell, J. Web-Based Textual Analysis of Free-Text Patient Experience Comments from a Survey in Primary Care. JMIR Med. Inform. 2015, 3, e20. [Google Scholar] [CrossRef] [Scilit]
- Guetterman, T.C.; Chang, T.; DeJonckheere, M.; Basu, T.; Scruggs, E.; Vydiswaran, V.V. Augmenting Qualitative Text Analysis with Natural Language Processing: Methodological Study. J. Med. Internet Res. 2018, 20, e231. [Google Scholar] [CrossRef] [Scilit]
- Hahn, S.; Kroehne, U.; Merk, S. Improving and Analyzing Open-Ended Survey Responses: A Case Study Linking Psychological Theories and Analysis Approaches for Text Data. Z. Psychol. 2024, 232, 171–180. [Google Scholar] [CrossRef] [Scilit]
- McGillivray, B.; Jenset, G.; Heil, D. Extracting Keywords from Open-Ended Business Survey Questions. J. Data Min. Digit. Humanit. 2020, 2020, 5077. [Google Scholar] [CrossRef] [Scilit]
- Nawaz, R.; Sun, Q.; Shardlow, M.; Kontonatsios, G.; Aljohani, N.R.; Visvizi, A.; Hassan, S.-U. Leveraging AI and Machine Learning for National Student Survey: Actionable Insights from Textual Feedback to Enhance Quality of Teaching and Learning in UK’s Higher Education. Appl. Sci. 2022, 12, 514. [Google Scholar] [CrossRef] [Scilit]
- Popping, R. Analyzing Open-ended Questions by Means of Text Analysis Procedures. Bull. Sociol. Methodol. 2015, 128, 23–39. [Google Scholar] [CrossRef] [Scilit]
- Liu, X.; Tvinnereim, E.; Grimsrud, K.M.; Lindhjem, H.; Velle, L.G.; Saure, H.I.; Lee, H. Explaining landscape preference heterogeneity using machine learning-based survey analysis. Landsc. Res. 2021, 46, 417–434. [Google Scholar] [CrossRef] [Scilit]
- Jaeger, S.R.; Rasmussen, M.A. Importance of data preparation when analysing written responses to open-ended questions: An empirical assessment and comparison with manual coding. Food Qual. Prefer. 2021, 93, 104270. [Google Scholar] [CrossRef] [Scilit]
- Küfner, B.; Sakshaug, J.W.; Zins, S. Analysing Establishment Survey Non-Response Using Administrative Data and Machine Learning. J. R. Stat. Soc. Ser. A 2022, 185, S310–S342. [Google Scholar] [CrossRef] [Scilit]
- Mourtgos, S.M.; Adams, I.T. The rhetoric of de-policing: Evaluating open-ended survey responses from police officers with machine learning-based structural topic modeling. J. Crim. Justice 2019, 64, 101627. [Google Scholar] [CrossRef] [Scilit]
- Yamano, H.; Park, J.J.; Choe, N.H.; Sakata, I. Understanding Students’ Perception of Sustainability: Educational NLP in the Analysis of Free Answers. Sustainability 2022, 14, 13970. [Google Scholar] [CrossRef] [Scilit]
- Lasri, I.; RiadSolh, A.; El Belkacemi, M. Toward an Effective Analysis of COVID-19 Moroccan Business Survey Data using Machine Learning Techniques. In Proceedings of the 2021 13th International Conference on Machine Learning and Computing, Shenzhen, China, 26–28 February 2021; pp. 50–58. [Google Scholar]
- Linton, A.G.; Dimitrova, V.; Downing, A.; Wagland, R.; Glaser, A. Weakly Supervised Text Classification on Free Text Comments in Patient-Reported Outcome Measures. Front. Digit. Health 2025, 7, 1345360. [Google Scholar] [CrossRef] [Scilit]
- González Canché, M.S. Machine driven classification of open-ended responses (MDCOR): An analytic framework and no-code, free software application to classify longitudinal and cross-sectional text responses in survey and social media research. Expert Syst. Appl. 2023, 215, 119265. [Google Scholar] [CrossRef] [Scilit]
- Marengo, D.; Hoeboer, C.M.; Veldkamp, B.P.; Olff, M. Text mining to improve screening for trauma-related symptoms in a global sample. Psychiatry Res. 2022, 316, 114753. [Google Scholar] [CrossRef] [Scilit]
- Buenano-Fernandez, D.; Gonzalez, M.; Gil, D.; Lujan-Mora, S. Text Mining of Open-Ended Questions in Self-Assessment of University Teachers: An LDA Topic Modeling Approach. IEEE Access 2020, 8, 35318–35330. [Google Scholar] [CrossRef] [Scilit]
- Pietsch, A.S.; Lessmann, S. Topic modeling for analyzing open-ended survey responses. J. Bus. Anal. 2018, 1, 93–116. [Google Scholar] [CrossRef] [Scilit]
- Baumer, E.P.S.; Mimno, D.; Guha, S.; Quan, E.; Gay, G.K. Comparing grounded theory and topic modeling: Extreme divergence or unlikely convergence? J. Assoc. Inf. Sci. Technol. 2017, 68, 1397–1410. [Google Scholar] [CrossRef] [Scilit]
- Shiwaku, A.; Kobayashi, N.; Kitagawa, F.; Shiina, H. Evaluation of Free Answer Comment Using Machine Learning by Word Evaluation. In Proceedings of the 2016 5th IIAI International Congress on Advanced Applied Informatics (IIAI-AAI), Kumamoto, Japan, 10–14 July 2016; IEEE: Piscataway, NJ, USA, 2016; pp. 134–137. [Google Scholar]
- Koufakou, A.; Gosselin, J.; Guo, D. Using data mining to extract knowledge from student evaluation comments in undergraduate courses. In Proceedings of the 2016 International Joint Conference on Neural Networks (IJCNN), Vancouver, BC, Canada, 24–29 July 2016; IEEE: Piscataway, NJ, USA, 2016; pp. 3138–3142. [Google Scholar]
- Etz, R.S.; Gonzalez, M.M.; Eden, A.R.; Winship, J. Rapid Sense Making: A Feasible, Efficient Approach for Analyzing Large Data Sets of Open-Ended Comments. Int. J. Qual. Methods 2018, 17, 1609406918765509. [Google Scholar] [CrossRef] [Scilit]
- Costales, J.A.; Albino, M.G.; Palaoag, T.D. Students’ Experiences in Using Learning Management System (Canvas): Application of TAM and Sentiment Analysis. In Proceedings of the TENCON 2022–2022 IEEE Region 10 Conference, Hong Kong, China, 1–4 November 2022; pp. 1–6. [Google Scholar]
- Zhang, T.; Moody, M.; Nelon, J.P.; Boyer, D.M.; Smith, D.H.; Visser, R.D. Using Natural Language Processing to Accelerate Deep Analysis of Open-Ended Survey Data. In Proceedings of the 2019 SoutheastCon, Huntsville, AL, USA, 11–14 April 2019; pp. 1–3. [Google Scholar]
- Suadaa, L.H.; Ridho, F.; Monika, A.K.; Projo, N.W.K. Automatic text categorization to standard classification of indonesian business fields (KBLI) 2020. In Proceedings of the 2023 International Conference on Electrical Engineering and Informatics (ICEEI), Gwangju, South Korea, 10–11 October 2023; IEEE: Piscataway, NJ, USA, 2023; pp. 1–6. [Google Scholar]
- Grönberg, N.; Knutas, A.; Hynninen, T.; Hujala, M. Palaute: An online text mining tool for analyzing written student course feedback. IEEE Access 2021, 9, 134518–134529. [Google Scholar] [CrossRef] [Scilit]
- Singha, C.; Mohapatra, R.L. Student Satisfaction on Online Learning During COVID-19 Using Machine Learning Techniques. In Proceedings of the 2023 International Conference on Sustainable Communication Networks and Application (ICSCNA), Theni, India, 15–17 November 2023; IEEE: Piscataway, NJ, USA, 2023; pp. 1683–1691. [Google Scholar]
- Mahendher, S.; Singhal, S.; Devaraj, J.; Markandey, P. Using machine learning to comprehend and forecast Post-COVID-19 pharmaceutical sales. In Proceedings of the 2023 Advanced Computing and Communication Technologies for High Performance Applications (ACCTHPA), Kochi, India, 19–20 January 2023; pp. 1–5. [Google Scholar]
- Kazi, N.; Kahanda, I.; Rupassara, S.I.; Kindt, J.W. Zero-Shot Information Extraction with Community-Fine-Tuned Large Language Models From Open-Ended Interview Transcripts. In Proceedings of the 2023 International Conference on Machine Learning and Applications (ICMLA), Jacksonville, FL, USA, 15–17 December 2023; pp. 932–937. [Google Scholar]
- Alshaikh, K.A.; Almatrafi, O.A.; Abushark, Y.B. BERT-Based Model for Aspect-Based Sentiment Analysis for Analyzing Arabic Open-Ended Survey Responses: A Case Study. IEEE Access 2024, 12, 2288–2302. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y. Evaluation of Tourist Satisfaction Based Light Gradient Boosting Machine Technique. In Proceedings of the 2024 International Conference on Intelligent Algorithms for Computational Intelligence Systems (IACIS), Hassan, India, 19–21 August 2024; IEEE: Piscataway, NJ, USA, 2024; pp. 1–5. [Google Scholar]
- Pestian, J.P.; Sorter, M.; Connolly, B.; Cohen, K.B.; McCullumsmith, C.; Gee, J.T.; Morency, L.; Scherer, S.; Rohlfs, L. A Machine Learning Approach to Identifying the Thought Markers of Suicidal Subjects: A Prospective Multicenter Trial. Suicide Life-Threat. Behav. 2017, 47, 112–121. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Clark, N.J.; Tozer, S.; Wood, C.; Firestone, S.M.; Stevenson, M.; Caraguel, C.; Chaber, A.; Heller, J.; Magalhães, R.J.S. Unravelling animal exposure profiles of human Q fever cases in Queensland, Australia, using natural language processing. Transbound. Emerg. Dis. 2020, 67, 2133–2145. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Li, Y.; Du, H.; Martin, I.; Hidalgo, E.; Jiang, Y.; Xing, W.; Popov, V. Exploring adolescents’ occupational possible selves: The role of gender and socioeconomic status. Career Dev. Q. 2023, 71, 189–205. [Google Scholar] [CrossRef] [Scilit]
- Torrao, G.; Htait, A.; Wong, S.H.S. Perceptions of Women’s Safety in Transient Environments and the Potential Role of AI in Enhancing Safety: An Inclusive Mobility Study in India. Sustainability 2024, 16, 8631. [Google Scholar] [CrossRef] [Scilit]
- Kobra, K.; Sammi, S.S.; Rahman, N.; Khushbu, S.A.; Islam, M. Multihead Text Mining from COVID-19 Feedback Using Machine Learning, Deep Learning, and Hybrid Deep Learning Approaches. J. Sens. 2024, 2024, 3027199. [Google Scholar] [CrossRef] [Scilit]
- Inoue, M.; Fukahori, H.; Matsubara, M.; Yoshinaga, N.; Tohira, H. Latent Dirichlet allocation topic modeling of free-text responses exploring the negative impact of the early COVID-19 pandemic on research in nursing. Jpn. J. Nurs. Sci. 2023, 20, e12520. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Nikulchev, E.; Ilin, D.; Silaeva, A.; Kolyasnikov, P.; Belov, V.; Runtov, A.; Pushkin, P.; Laptev, N.; Alexeenko, A.; Magomedov, S.; et al. Digital Psychological Platform for Mass Web-Surveys. Data 2020, 5, 95. [Google Scholar] [CrossRef] [Scilit]
- Cook, B.L.; Progovac, A.M.; Chen, P.; Mullin, B.; Hou, S.; Baca-Garcia, E. Novel Use of Natural Language Processing (NLP) to Predict Suicidal Ideation and Psychiatric Symptoms in a Text-Based Mental Health Intervention in Madrid. Comput. Math. Methods Med. 2016, 2016, 8708434. [Google Scholar] [CrossRef] [Scilit]
- Onan, A. Mining opinions from instructor evaluation reviews: A deep learning approach. Comput. Appl. Eng. Educ. 2020, 28, 117–138. [Google Scholar] [CrossRef] [Scilit]
- Wijngaards, I.; Burger, M.; Exel, J.V. The promise of open survey questions—The validation of text-based job satisfaction measures. PLoS ONE 2019, 14, e0226408. [Google Scholar] [CrossRef] [Scilit]
- Robin, C.; Mashinchi, M.I.; Zeleti, F.A.; Ojo, A.; Buitelaar, P. A Term Extraction Approach to Survey Analysis in Health Care. In Proceeding of the 12th Conference on Language Resources and Evaluation (LREC 2020), Marseille, France, 11–16 May 2020. [Google Scholar]
- Tvinnereim, E.; Fløttum, K. Explaining topic prevalence in answers to open-ended survey questions about climate change. Nat. Clim. Change 2015, 5, 744–747. [Google Scholar] [CrossRef] [Scilit]
- Gish, M.; Nowominski, A.; Dror, R. Breaking the Ceiling: Mitigating Extreme Response Bias in Surveys Using an Open-Ended Adaptive-Testing System and LLM-Based Response Analysis. AI 2026, 7, 73. [Google Scholar] [CrossRef] [Scilit]
- Cheese, E.; Bichoo, R.A.; Grover, K.; Dumitru, D.; Zenonos, A.; Groark, J.; Gibson, D.; Pope, R. Using Natural Language Processing to Explore Patient Perspectives on AI Avatars in Support Materials for Patients with Breast Cancer: Survey Study. J. Med. Internet Res. 2025, 27, e70971. [Google Scholar] [CrossRef] [Scilit]
- Padilla-Rascón, M.A.; González, P.; Carmona, C.J. SDRDPy: An application to graphically visualize the knowledge obtained with supervised descriptive rule algorithms. SoftwareX 2024, 28, 101939. [Google Scholar] [CrossRef] [Scilit]
- Carmona, C.J.; Del Jesus, M.J.; Herrera, F. A unifying analysis for the supervised descriptive rule discovery via the weighted relative accuracy. Knowl.-Based Syst. 2018, 139, 89–100. [Google Scholar] [CrossRef] [Scilit]
- Goldberg, Y.; Levy, O. word2vec Explained: Deriving Mikolov et al.’s negative-sampling word-embedding method. arXiv 2014, arXiv:1402.3722. [Google Scholar] [CrossRef] [Scilit]
- Al-Radhi, M.S.; Csapó, T.G.; Németh, G. Adaptive Refinements of Pitch Tracking and HNR Estimation within a Vocoder for Statistical Parametric Speech Synthesis. Appl. Sci. 2019, 9, 2460. [Google Scholar] [CrossRef] [Scilit]













| Area | Keywords | Related Concepts |
|---|---|---|
| Surveys with open-ended questions | Analysis | Data mining, Data analysis, Artificial intelligence |
| Machine learning | ||
| Survey | ||
| Open-ended question | ||
| Open-ended | ||
| NLP Natural language processing | ||
| SDRD | ||
| Supervised descriptive rule discovery | ||
| Free text questions | ||
| Open-ended answers | ||
| Open-ended survey |
| Types of Learning | Total | Article |
|---|---|---|
| Supervised and Unsupervised | 33 | [9,18,25,26,27,28,31,32,33,34,36,39,41,44,47,49,54,56,61,63,64,66,69,75,77,81,82,83,85,88,90,91,93] |
| Supervised | 20 | [11,14,23,24,30,37,38,42,48,50,57,58,73,76,78,79,80,84,87,89] |
| Unsupervised | 19 | [8,10,13,29,40,43,51,52,53,59,60,65,67,70,71,72,74,86,92] |
| Semi-supervised | 5 | [12,29,46,62,68] |
| Not specified | 2 | [7,55] |
| Not applicable | 1 | [45] |
| Task | AI Area | Article |
|---|---|---|
| Classification | ML | [11,14,23,24,26,30,36,37,38,39,41,42,47,48,50,54,55,56,57,58,61,62,63,66,68,69,75,76,77,78,79,80,83,84,87,88,89,90] |
| Topic modeling | NLP | [8,9,10,25,28,31,39,40,43,44,49,52,54,56,59,63,65,66,67,71,74,82,85,86,91,93] |
| Clustering | ML | [8,13,28,32,43,44,46,51,61,65,70,72,86] |
| Regression | ML | [12,25,28,31,44,47,49,58,64,75,81,82,91] |
| Correlation | ML | [25,26,28,44,47,60,64,75,83,88,90] |
| Sentiment analysis | NLP | [23,24,28,36,39,69,77,78,83,89,93] |
| SDRD | ML | [18,27,32,33,34] |
| Association | ML | [41,69,81] |
| Key phrase extraction | NLP | [28] |
| Keyword-based topic analysis | NLP | [29] |
| Topic-based classification | NLP | [36] |
| Keyword extraction | NLP | [53] |
| Semantic Text Matching | NLP | [77] |
| Exact Answer Extraction (EAE) | NLP | [77] |
| Emotion classification | NLP | [83] |
| Term extraction | NLP | [90] |
| Not specified | NA | [7] |
| Not applicable | NA | [45] |
| Algorithm | Article |
|---|---|
| SVM | [11,14,23,26,36,37,42,46,54,66,73,75,76,79,80,84,88] |
| RF | [14,25,26,30,37,54,58,64,73,75,76,79,84,88] |
| LR (Logistic regression) | [37,50,58,75,76,81,82,87,88] |
| Naive Bayes | [30,36,37,69,75,79,84,88] |
| K-Means | [28,32,44,46,61,86,93] |
| KNN (K-Nearest Neighbor) | [36,69,84,88] |
| Gradient boosting | [12,58,79] |
| Multilayer Perceptron | [23,41] |
| XGBoost | [14,58] |
| Lasso regression | [58,64] |
| Ridge regression | [58,64] |
| PA (Passive Aggressive) | [11] |
| Multinomial boosting | [12] |
| Best-first decision tree | [30] |
| SMO (Sequential Minimal Optimization) | [41] |
| Multilevel Bayesian Modeling | [52] |
| MNB (Multinomial Naïve Bayes) | [54] |
| Ordinal logistic regression | [56] |
| PLS-DA (Partial Least-Squares Discriminant Analysis) | [57] |
| BART (Bayesian additive regression trees) | [58] |
| C-Tree (Conditional inference tree) | [58] |
| CART (Classification and Regression Trees) | [58] |
| Stacking ensemble | [64] |
| SMOTE (Synthetic Minority Over-sampling Technique) | [76] |
| Models/Algorithms | Article |
|---|---|
| LDA | [8,9,25,28,31,36,40,43,44,54,64,65,66,67,71,74,81,82,85,86,93] |
| STM | [10,31,36,49,52,56,59,74,91] |
| BERT | [14,28,36,39,52,73,78] |
| RoBERTa | [36,77,83,93] |
| BERTopic | [52,62,93] |
| BTM | [8,66] |
| NMF (Non-negative Matrix Factorization) | [39,93] |
| Gibbs Sampling | [63,66] |
| WNTM (Word Network Topic Model) | [8] |
| XLNet | [36] |
| Sentence-BERT (SBERT) | [44] |
| Based on lexicon and rules (Vader) | [71] |
| IndoBERT | [73] |
| SRoBERTa | [77] |
| AraBERT | [78] |
| MARBERT | [78] |
| QARiB | [78] |
| Local Context Focus-Aspect Term Extraction and Polarity Classification (LCF-ATEPC) | [78] |
| SentiStrength | [89] |
| OpenNLP Perceptron | [90] |
| GPT-4.1 | [92] |
| GPT-4o | [92] |
| GPT-5 | [92] |
| Gemini-2.0-Flash | [92] |
| Gemini-2.5-Flash | [92] |
| Twitter-roBERTa-base for Sentiment Analysis-Updated | [93] |
| Pegasus | [93] |
| GPT-2 | [93] |
| Bidirectional and autoregressive transformers | [93] |
| T5 | [93] |
| FLAN-T5-model | [93] |
| Model | Total | Article |
|---|---|---|
| n-grams | 16 | [12,14,29,37,42,47,48,50,57,62,64,69,80,84,85,87] |
| TF-IDF | 12 | [25,36,39,53,60,62,65,69,73,85,88,93] |
| BoW (Bag-of-Words) | 10 | [9,36,43,57,59,65,66,69,82] |
| Word2Vec | 8 | [8,23,25,36,39,66,69,88] |
| BERT vectors | 7 | [14,28,36,52,62,73,78] |
| Linguistic Inquiry and Word Count (LIWC) dictionary | 4 | [25,48,60,64,89] |
| Dense vector embeddings | 2 | [62,84] |
| GloVe (Global Vectors for Word Representation) | 2 | [66,88] |
| fastText | 1 | [72,88] |
| Swivel (Submatrix-wise Vector Embedding Learner) | 1 | [23] |
| Doc2Vec | 1 | [25] |
| Word/sentence embeddings | 1 | [37] |
| SBERT-based sentence embedding vectors | 1 | [44] |
| BoW vectorization model | 1 | [86] |
| WeSTClass embedding | 1 | [62] |
| Term-presence | 1 | [88] |
| Term-Frequency | 1 | [88] |
| LDA2Vec | 1 | [88] |
| Libraries | Article |
|---|---|
| stm package | [10,31,52,59,74] |
| NLTK | [60,82,86,93] |
| Gensim | [40,85,86,93] |
| spaCy | [8,24,38,60] |
| pyLDAvis | [9,40,74,85] |
| Keras | [36,61,84,88] |
| MALLET (MAchine Learning for LanguagE Toolkit) | [40,54,64,67] |
| pandas | [24,38,75] |
| quanteda | [37,52,56] |
| tidyverse | [57,81] |
| FastText | [8,88] |
| TensorFlow | [61,88] |
| HuggingFace | [36,77] |
| Caret | [57,58] |
| LiblineaR | [37,87] |
| VADER Library | [24,93] |
| Schedule | [24] |
| scikit-learn | [79] |
| Stanford Core NLP | [30] |
| Jieba package | [44] |
| SentiWordNet | [24] |
| MultiWordNet | [24] |
| PyTorch | [36] |
| Janome | [85] |
| dplyr packages | [57] |
| stringr | [57] |
| tidyr | [57] |
| R splitTools library | [64] |
| tableone de R | [57] |
| Hunspell packages | [37] |
| PrimeFaces | [18] |
| stm package | [10,31,52,59,74] |
| Broom | [57] |
| MLeval | [57] |
| BartMachine | [58] |
| Rpart | [58] |
| partykit | [58] |
| Gam | [58] |
| gamsel | [58] |
| R glmnet packages | [58] |
| Sentix | [24] |
| corpus | [57] |
| tyditext | [57] |
| Transformers Python library | [93] |
| Programming Language | Article |
|---|---|
| Python | [9,14,24,32,36,38,39,40,44,53,66,75,76,83,84,85,86,93] |
| R | [10,42,43,49,56,57,58,59,64,66,74,76,81,82,89,91] |
| Java | [18,32,34,66] |
| Visual Basic | [70] |
| C++ | [66] |
| JavaScript | [86] |
| Bash | [66] |
| Coffeescript | [86] |
| Graphic | Article |
|---|---|
| Word cloud | [23,24,25,36,43,64,71,77,81,85,93] |
| Heatmap | [25,65] |
| Inter-topic distance plot | [74,85] |
| Violin plot | [24] |
| TagCrowd | [50] |
| Many Eyes | [50] |
| Topic-document | [74] |
| Bipartite network | [65] |
| Word map | [87] |
| Histogram charts | [23] |
| Statistical Test | Article |
|---|---|
| Chi-square tests | [18,85] |
| ANOVA | [57,88] |
| F-test statistics | [24] |
| Granger causality hypothesis test model | [24] |
| Augmented Dickey–Fuller Test | [24] |
| Likelihood-ratio test | [24] |
| The Wilcoxon test | [28] |
| Harman’s one-factor testing | [44] |
| Wilcoxon rank-sum (Mann–Whitney) test | [87] |
| Programming Language | Article |
|---|---|
| WEKA | [18,30,34,41,64,88] |
| Stata | [12,42,50,58,87] |
| IBM SPSS | [38,44,45,53,83] |
| EPM Framework | [27,33] |
| VIKAMINE | [18] |
| Rapidminer | [69] |
| Orange3 | [79] |
| Technology | Article |
|---|---|
| Text mining (Minería de texto) | [12,51,64,65,69,84,87] |
| Microsoft Excel | [45,70] |
| Canvas | [69,71] |
| Gavagai Explorer | [13] |
| LimeSurvey | [24] |
| MetaMap | [30] |
| Monte Carlo simulation | [31] |
| MPLus 8 | [44] |
| Semantic Excel | [47] |
| Voyant Tools | [50] |
| Qualitative software MAXQDA 12 | [51] |
| PsychArchives | [52] |
| SPSS Survey Analyzer 4.0.1 | [53] |
| T-Lab | [55] |
| Text Component Analysis (TCA) | [55] |
| TextQuest | [55] |
| Wordscores | [55] |
| WordStat | [55] |
| Yoshikoder | [55] |
| Syuzhet | [74] |
| R/RStudio for statistical analysis | [76] |
| Pre-trained language models (CLLMs) | [77] |
| COVAREP (Cooperative Voice Analysis Repository for Speech Technologies) | [80] |
| NVIDIA Tesla T4 GPU | [84] |
| CUDA 12.0 toolkit | [84] |
| Plataforma Digital para la Investigación Psicológica Interdisciplinaria | [86] |
| Radboud Faces Database | [47] |
| cTakes (Clinical Text Analysis Knowledge Extract System) | [87] |
| Bayesian optimization with a Gaussian process | [88] |
| Minitab statistical software | [88] |
| Text mining | [12,51,64,65,69,84,87] |
| Computer-Aided Text Analysis | [89] |
| Saffron knowledge extraction tool | [90] |
| Methodological framework (ARC framework) | [90] |
| Areas | Total | Article |
|---|---|---|
| Health | 30 | [18,24,27,30,33,34,38,39,41,43,45,46,47,50,52,61,62,64,70,76,80,81,83,84,85,86,87,90,91,93] |
| Research | 24 | [7,8,11,12,13,14,26,31,37,42,49,51,53,55,56,57,58,63,66,67,72,77,78,92] |
| Education | 17 | [23,25,28,29,36,40,48,54,60,65,68,69,71,74,75,82,88] |
| Service | 8 | [9,10,32,44,59,73,79,89] |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Olmos-Vallejo, A.; Rodríguez-Mazahua, L.; Machorro-Cano, I.; Palet-Guzmán, J.A.; Alor-Hernández, G.; Cervantes, J.; Sánchez-Cervantes, J.L. Application of Machine Learning and Natural Language Processing Techniques for the Analysis of Surveys with Open-Ended Questions: A Scoping Review. Computers 2026, 15, 342. https://doi.org/10.3390/computers15060342
Olmos-Vallejo A, Rodríguez-Mazahua L, Machorro-Cano I, Palet-Guzmán JA, Alor-Hernández G, Cervantes J, Sánchez-Cervantes JL. Application of Machine Learning and Natural Language Processing Techniques for the Analysis of Surveys with Open-Ended Questions: A Scoping Review. Computers. 2026; 15(6):342. https://doi.org/10.3390/computers15060342
Chicago/Turabian StyleOlmos-Vallejo, Araceli, Lisbeth Rodríguez-Mazahua, Isaac Machorro-Cano, José Antonio Palet-Guzmán, Giner Alor-Hernández, Jair Cervantes, and José Luis Sánchez-Cervantes. 2026. "Application of Machine Learning and Natural Language Processing Techniques for the Analysis of Surveys with Open-Ended Questions: A Scoping Review" Computers 15, no. 6: 342. https://doi.org/10.3390/computers15060342
APA StyleOlmos-Vallejo, A., Rodríguez-Mazahua, L., Machorro-Cano, I., Palet-Guzmán, J. A., Alor-Hernández, G., Cervantes, J., & Sánchez-Cervantes, J. L. (2026). Application of Machine Learning and Natural Language Processing Techniques for the Analysis of Surveys with Open-Ended Questions: A Scoping Review. Computers, 15(6), 342. https://doi.org/10.3390/computers15060342

