Evidence-Grounded LLM Summarization for Actionable Student Feedback Analysis
Abstract
1. Introduction
2. Literature Review
3. Dataset Collection and Annotation
3.1. Primary Dataset Collection
3.2. Student Responses and Annotation
3.3. Preprocessing and Data Cleaning
3.4. Data Augmentation and Expansion
3.5. External Validation Datasets
4. Methodology
4.1. Framework Overview
- Layer 1: Data Collection and Preparation constructs standardized, analysis-ready corpora from institutional and external datasets through cleaning, standardization, and harmonized preprocessing.
- Layer 2: Parallel Analysis Pipelines executes three complementary analytical modules: (A) supervised classification for category-aware feedback organization, (B) unsupervised clustering for latent theme discovery, and (C) retrieval-augmented generation for evidence-grounded summarization.
- Layer 3: Integration and Decision Support synthesizes outputs from all pipelines, cross-validates analytical signals, and produces interpretable summaries and diagnostic reports to support institutional decision-making.
4.2. Supervised Classification Module
4.2.1. Supervised Classification Models
- TF–IDF + Logistic Regression (LR) [6].
- TF–IDF + Support Vector Machine (SVM) [5].
- Hybrid TF–IDF + MPNet [34].
- Deep Text-Level Processing (DTLP) [35].
- Multi-Head Attention Fusion (MHAF) [36].
- SetFit (contrastive few-shot learning) [14].
- DeBERTa-v3 with Low-Rank Adaptation (LoRA) [12].
- Weighted ensemble (TF–IDF + SVM, SetFit, DeBERTa-LoRA) (ours).
4.2.2. Ensemble Fusion Strategy
4.2.3. Training and Evaluation
4.3. Unsupervised Clustering Module
4.3.1. Sentence Embedding Models
4.3.2. Embedding Fusion
4.3.3. Clustering and Pseudo-Label Induction
4.3.4. Unsupervised Clustering Evaluation Metrics
4.4. Retrieval-Augmented Generation (RAG) Module
4.5. Evidence-Grounded Generation
5. Results
5.1. Supervised Learning
5.2. Unsupervised Clustering
5.3. RAG-Based Summarization
6. Discussion and Limitations
6.1. Value of Integrated Analysis
6.2. From Analysis to Actionability
6.3. Robustness and Interpretability
6.4. Generalizability and Scope
6.5. Implications for Educational Practice
6.6. Limitations
6.7. Future Research Directions
7. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Kastrati, Z.; Imran, A.; Kurti, A. Weakly Supervised Framework for Aspect-Based Sentiment Analysis on Students’ Reviews of MOOCs. IEEE Access 2020, 8, 106799–106810. [Google Scholar] [CrossRef]
- Ren, P.; Yang, L.; Luo, F. Automatic Scoring of Student Feedback for Teaching Evaluation Based on Aspect-Level Sentiment Analysis. Educ. Inf. Technol. 2023, 28, 797–814. [Google Scholar] [CrossRef]
- Kastrati, Z.; Dalipi, F.; Imran, A.; Pireva Nuci, K.; Wani, M. Sentiment Analysis of Students’ Feedback with NLP and Deep Learning: A Systematic Mapping Study. Appl. Sci. 2021, 11, 3986. [Google Scholar] [CrossRef]
- Sindhu, I.; Daudpota, S.; Badar, K.; Bakhtyar, M.; Baber, J.; Nurunnabi, M. Aspect-Based Opinion Mining on Student’s Feedback for Faculty Teaching Performance Evaluation. IEEE Access 2019, 7, 108729–108741. [Google Scholar] [CrossRef]
- Joachims, T. Text Categorization with Support Vector Machines: Learning with Many Relevant Features. In Proceedings of the European Conference on Machine Learning (ECML), Chemnitz, Germany, 21–23 April 1998; Springer: Berlin/Heidelberg, Germany, 1998; pp. 137–142. [Google Scholar] [CrossRef]
- Sebastiani, F. Machine Learning in Automated Text Categorization. ACM Comput. Surv. 2002, 34, 1–47. [Google Scholar] [CrossRef]
- Dalipi, F.; Imran, A.; Kastrati, Z. MOOC Dropout Prediction Using Machine Learning Techniques: Review and Research Challenges. In Proceedings of the IEEE Global Engineering Education Conference (EDUCON), Tenerife, Spain, 17–20 April 2018; IEEE: New York, NY, USA, 2018; pp. 1007–1014. [Google Scholar] [CrossRef]
- Devlin, J.; Chang, M.W.; Lee, K.; Toutanova, K. BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics (NAACL-HLT), Minneapolis, MN, USA, 2–7 June 2019; Association for Computational Linguistics: Stroudsburg, PA, USA, 2019; pp. 4171–4186. [Google Scholar] [CrossRef]
- Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; Stoyanov, V. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv 2019, arXiv:1907.11692. [Google Scholar]
- Yang, Z.; Dai, Z.; Yang, Y.; Carbonell, J.; Salakhutdinov, R.; Le, Q. XLNet: Generalized Autoregressive Pretraining for Language Understanding. In Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada, 8–14 December 2019; Curran Associates Inc.: Red Hook, NY, USA, 2019; pp. 5753–5763. [Google Scholar]
- Lan, Z.; Chen, M.; Goodman, S.; Gimpel, K.; Sharma, P.; Soricut, R. ALBERT: A Lite BERT for Self-Supervised Learning of Language Representations. In Proceedings of the International Conference on Learning Representations (ICLR), Addis Ababa, Ethiopia, 26–30 April 2020. [Google Scholar]
- Hu, E.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W. LoRA: Low-Rank Adaptation of Large Language Models. In Proceedings of the International Conference on Learning Representations (ICLR), Virtual, 25–29 April 2022. [Google Scholar]
- Ding, N.; Qin, Y.; Yang, G.; Wei, F.; Yang, Z.; Su, Y.; Hu, S.; Shi, Y.; Liu, Z.; Sun, M.; et al. Parameter-Efficient Fine-Tuning of Large-Scale Pre-Trained Language Models. Nat. Mach. Intell. 2023, 5, 220–235. [Google Scholar] [CrossRef]
- Tunstall, L.; Reimers, N.; Jo, U.; Bates, L.; Korat, D.; Wasserblat, M.; Pereg, O. Efficient Few-Shot Learning Without Prompts. arXiv 2022, arXiv:2209.11055. [Google Scholar] [CrossRef]
- Ganaie, M.; Hu, M.; Malik, A.; Tanveer, M.; Suganthan, P. Ensemble Deep Learning: A Review. Eng. Appl. Artif. Intell. 2022, 115, 105151. [Google Scholar] [CrossRef]
- Wolpert, D. Stacked Generalization. Neural Netw. 1992, 5, 241–259. [Google Scholar] [CrossRef]
- Blei, D.; Ng, A.; Jordan, M. Latent Dirichlet Allocation. J. Mach. Learn. Res. 2003, 3, 993–1022. [Google Scholar]
- Grootendorst, M. BERTopic: Neural Topic Modeling with a Class-Based TF-IDF Procedure. arXiv 2022, arXiv:2203.05794. [Google Scholar]
- Staubitz, T.; Petrick, D.; Bauer, M.; Renz, J.; Meinel, C. Improving the Peer Assessment Experience on MOOC Platforms. In Proceedings of the Third ACM Conference on Learning at Scale (L@S), Edinburgh, UK, 25–26 April 2016; Association for Computing Machinery: New York, NY, USA, 2016; pp. 389–398. [Google Scholar] [CrossRef]
- Reimers, N.; Gurevych, I. Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), Hong Kong, China, 3–7 November 2019; Association for Computational Linguistics: Stroudsburg, PA, USA, 2019; pp. 3982–3992. [Google Scholar] [CrossRef]
- Reimers, N.; Gurevych, I. Making Monolingual Sentence Embeddings Multilingual Using Knowledge Distillation. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), Online, 16–20 November 2020; Association for Computational Linguistics: Stroudsburg, PA, USA, 2020; pp. 4512–4525. [Google Scholar] [CrossRef]
- Jolliffe, I.; Cadima, J. Principal Component Analysis: A Review and Recent Developments. Philos. Trans. R. Soc. A 2016, 374, 20150202. [Google Scholar] [CrossRef]
- McInnes, L.; Healy, J.; Melville, J. UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction. arXiv 2018, arXiv:1802.03426. [Google Scholar]
- Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; Küttler, H.; Lewis, M.; Yih, W.t.; Rocktäschel, T.; et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Proceedings of the Advances in Neural Information Processing Systems, Virtual, 6–12 December 2020; Curran Associates Inc.: Red Hook, NY, USA, 2020; pp. 9459–9474. [Google Scholar]
- Asai, A.; Wu, Z.; Wang, Y.; Sil, A.; Hajishirzi, H. Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection. In Proceedings of the 12th International Conference on Learning Representations (ICLR), Vienna, Austria, 7–11 May 2024. [Google Scholar]
- Guu, K.; Lee, K.; Tung, Z.; Pasupat, P.; Chang, M.W. REALM: Retrieval-Augmented Language Model Pre-Training. In Proceedings of the International Conference on Machine Learning (ICML), Virtual, 13–18 July 2020; JMLR.org: Norfolk, MA, USA, 2020; pp. 3929–3938. [Google Scholar]
- Jeong, S.; Baek, J.; Cho, S.; Hwang, S.; Park, J. Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), Mexico City, Mexico, 16–21 June 2024; Association for Computational Linguistics: Stroudsburg, PA, USA, 2024; pp. 7036–7050. [Google Scholar] [CrossRef]
- Mialon, G.; Dessì, R.; Lomeli, M.; Nalmpantis, C.; Pasunuru, R.; Raj, K.; Riedel, S.; Khot, T.; Roller, S.; Weston, J.; et al. Augmented Language Models: A Survey. Trans. Mach. Learn. Res. 2023. [Google Scholar] [CrossRef]
- Yu, H.; Gan, A.; Zhang, K.; Tong, S.; Liu, Q.; Liu, Z. Evaluation of Retrieval-Augmented Generation: A Survey. arXiv 2024, arXiv:2405.07437. [Google Scholar] [CrossRef]
- Zhang, Y.; Jin, R.; Zhou, Z.H. Understanding Bag-of-Words Model: A Statistical Framework. Int. J. Mach. Learn. Cybern. 2010, 1, 43–52. [Google Scholar] [CrossRef]
- Hua, Y.; Denny, P.; Wicker, J.S.; Taskova, K. The EduRABSA Dataset and the ASQE-DPT Annotation Tool. Zenodo. 2025. Available online: https://zenodo.org/records/16935018 (accessed on 20 March 2026).
- Mungunshagai, M. Coursera Course Reviews Dataset. Hugging Face Datasets. 2024. Available online: https://huggingface.co/datasets/MungunshagaiT/coursera-reviews (accessed on 20 March 2026).
- McInnes, L.; Healy, J.; Astels, S. HDBSCAN: Hierarchical Density Based Clustering. J. Open Source Softw. 2017, 2, 205. [Google Scholar] [CrossRef]
- Song, K.; Tan, X.; Qin, T.; Lu, J.; Liu, T.Y. MPNet: Masked and Permuted Pre-Training for Language Understanding. In Proceedings of the Advances in Neural Information Processing Systems, Virtual, 6–12 December 2020; Curran Associates Inc.: Red Hook, NY, USA, 2020; pp. 16857–16867. [Google Scholar]
- Abdi, A.; Sedrakyan, G.; Veldkamp, B.; van Hillegersberg, J.; van den Berg, S. Students Feedback Analysis Model Using Deep Learning-Based Method and Linguistic Knowledge for Intelligent Educational Systems. Soft Comput. 2023, 27, 14073–14094. [Google Scholar] [CrossRef]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.; Kaiser, Ł.; Polosukhin, I. Attention Is All You Need. In Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA, 4–9 December 2017; Curran Associates Inc.: Red Hook, NY, USA, 2017; pp. 5998–6008. [Google Scholar]
- Li, Z.; Zhang, X.; Li, Y.; Long, D.; Xie, P.; Zhang, M. Towards General Text Embeddings with Multi-Stage Contrastive Learning. arXiv 2023, arXiv:2308.03281. [Google Scholar] [CrossRef]
- Su, H.; Shi, W.; Kasai, J.; Wang, Y.; Hu, Y.; Ostendorf, M.; Yih, W.t.; Smith, N.; Zettlemoyer, L.; Yu, T. One Embedder, Any Task: Instruction-Finetuned Text Embeddings. arXiv 2022, arXiv:2212.09741. [Google Scholar]
- Es, S.; James, J.; Espinosa-Anke, L.; Schockaert, S. RAGAS: Automated Evaluation of Retrieval Augmented Generation. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, St. Julians, Malta, 17–22 March 2024; Association for Computational Linguistics: Stroudsburg, PA, USA, 2024; pp. 150–158. [Google Scholar] [CrossRef]
- Conneau, A.; Kruszewski, G.; Lample, G.; Barrault, L.; Baroni, M. What you can cram into a single vector and what you can’t: Probing sentence embeddings for linguistic properties. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL), Melbourne, Australia, 15–20 July 2018; Association for Computational Linguistics: Stroudsburg, PA, USA, 2018; pp. 2126–2136. [Google Scholar] [CrossRef]




| Category | Example | Count | Percentage |
|---|---|---|---|
| Teaching Quality | Instructor explained topics clearly and answered questions patiently. | 105 | 10.9 |
| General Comments & Satisfaction | Good course, but pacing could be adjusted for beginners. | 210 | 21.9 |
| Improvement Suggestions | More coding examples would make the lectures more engaging. | 95 | 9.9 |
| Facilities/Inclusivity | Computers in the lab were slow, and Wi-Fi often disconnected. | 172 | 17.9 |
| Career Preparation | The assignments helped me understand industry-relevant projects. | 96 | 10.0 |
| Support Services | Technical support responded quickly when I had setup issues. | 89 | 9.3 |
| Social & Extracurricular | Hackathons organized by the department motivated me to participate. | 192 | 20.0 |
| Total | 959 | 100.0 |
| Characteristic | Primary (Ours) | EduRABSA [31] | Coursera [32] |
|---|---|---|---|
| Sample size | 959 | 6500 | 8200+ |
| Average length (tokens) | 87.3 | 42.3 | 67.8 |
| Learning context | In-person | Mixed | Online |
| Geographic scope | Central Asia | US/EU | Global |
| Label structure | 7 categories | Aspect-level | Course-level |
| Model | Primary | EduRABSA | Coursera | |||
|---|---|---|---|---|---|---|
| Acc | F1 | Acc | F1 | Acc | F1 | |
| Classical Models | ||||||
| TF-IDF + LR/SVM | 80.0 (±2.7) | 0.798 | 72.9 (±0.9) | 0.700 | 48.3 (±1.6) | 0.492 |
| Hybrid (TF-IDF + Emb) | 82.2 (±2.2) | 0.821 | 74.4 (±1.1) | 0.715 | 49.4 (±0.9) | 0.519 |
| Neural Models | ||||||
| DTLP | 74.5 (±3.4) | 0.746 | 59.4 (±0.7) | 0.546 | 47.5 (±1.0) | 0.402 |
| MHAF | 71.6 (±1.9) | 0.714 | 61.5 (±0.7) | 0.547 | 48.8 (±0.8) | 0.384 |
| Transformer Models | ||||||
| SetFit | 84.5 (±2.5) | 0.846 | 80.2 (±1.1) | 0.774 | 49.3 (±0.7) | 0.469 |
| DeBERTa-LoRA | 83.3 (±1.2) | 0.820 | 73.2 (±1.1) | 0.699 | 47.1 (±1.2) | 0.420 |
| MPNet FT | 79.2 (±2.5) | 0.799 | 72.9 (±0.5) | 0.688 | 48.4 (±1.9) | 0.438 |
| Ensemble (Ours) | 83.0 (±2.0) | 0.829 | 81.1 (±1.4) | 0.778 | 49.8 (±0.7) | 0.534 |
| Fusion Strategy | Silhouette ↑ | Davies–Bouldin ↓ | Calinski–Harabasz ↑ | Clusters | Noise |
|---|---|---|---|---|---|
| Average Ensemble | 0.238 | 1.369 | 24.12 | 8 | 612 (63.7%) |
| Weighted Ensemble | 0.259 | 1.287 | 25.67 | 9 | 564 (58.7%) |
| Concatenation | 0.219 | 1.521 | 22.34 | 11 | 721 (75.1%) |
| Concat + PCA | 0.271 | 1.241 | 26.94 | 10 | 412 (42.9%) |
| Model | Faithfulness ↑ | Cluster Coverage ↑ |
|---|---|---|
| Baseline RAG | 0.71 | 0.59 |
| Supervised RAG | 0.89 | 0.60 |
| Supervised + Unsupervised RAG (Ours) | 0.93 | 0.64 |
| Component | Example 1: High Agreement | Example 2: Split Clusters |
|---|---|---|
| Category | Teaching Quality (supervised ensemble) | General Comments & Satisfaction (supervised ensemble) |
| Samples analyzed | 105 total → 10 via RAG (top-k) | 210 total → 10 via RAG (top-k) |
| Cluster mapping | 84% → C0 (96.9% purity) | 41% → C1 (90.8% purity); 32% → C3; 27% noise |
| Agreement level | High (C0 predominantly Teaching Quality) | Moderate (category spans multiple clusters) |
| Top keywords | lectures, clarity, examples, pacing, difficult, fast | overall, satisfied, workload, balance, experience |
| Generated summary | Lecture clarity and pacing issues dominate, with 68% negative sentiment. Difficulty following explanations (84%), rapid pacing (76%), and insufficient worked examples (62%) were frequent. Actions include 3–4 worked examples per lecture, pause intervals, scaffolding materials, and reduced content density. | Feedback is heterogeneous: 58% overall satisfaction, with helpfulness (45%) and challenging-but-rewarding perceptions (38%). Concerns include workload/resources (32%) and work–life balance (27%). Outputs include both actionable items and diagnostic flags for manual review. |
| Statistics provided | 4 indicators (68%, 84%, 76%, 62%) | 6 indicators (58%, 45%, 38%, 32%, 42%, 35%) |
| Actionability | 4 quantified recommendations | Mixed: actionable + diagnostic |
| Traceability | Category: Teaching Quality; Cluster: C0; IDs: FB_0023, FB_0087, FB_0134, FB_0156, FB_0189, FB_0203, FB_0267, FB_0289, FB_0312, FB_0378 | Category: General Comments; Clusters: C1, C3, noise; IDs: FB_0156, FB_0189, FB_0234, FB_0278, FB_0301, FB_0345, FB_0389, FB_0412, FB_0456, FB_0489 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Baimukanova, Z.; Saparbekov, Y.; Ha, H.; Lee, M. Evidence-Grounded LLM Summarization for Actionable Student Feedback Analysis. Information 2026, 17, 351. https://doi.org/10.3390/info17040351
Baimukanova Z, Saparbekov Y, Ha H, Lee M. Evidence-Grounded LLM Summarization for Actionable Student Feedback Analysis. Information. 2026; 17(4):351. https://doi.org/10.3390/info17040351
Chicago/Turabian StyleBaimukanova, Zhanerke, Yerassyl Saparbekov, Hyesong Ha, and Minho Lee. 2026. "Evidence-Grounded LLM Summarization for Actionable Student Feedback Analysis" Information 17, no. 4: 351. https://doi.org/10.3390/info17040351
APA StyleBaimukanova, Z., Saparbekov, Y., Ha, H., & Lee, M. (2026). Evidence-Grounded LLM Summarization for Actionable Student Feedback Analysis. Information, 17(4), 351. https://doi.org/10.3390/info17040351

