When Do Undergraduate Students Prefer AI? Insights into AI Scoring and Feedback
Abstract
1. Introduction
- What are students’ preferred characteristics of AI feedback, and do these preferences differ from chance?
- How do students’ preferences for AI versus human scoring and feedback vary across (a) assessment scenarios and (b) types, and are these differences statistically significant?
- What are students’ perceptions of AI scoring and feedback?
- What are students’ preferences regarding (a) the role of AI and (b) AI-based scoring relative to human graders, and what factors predict these preferences?
- What preferences do students express regarding the AI scoring and feedback after evaluating AI-generated assessment outputs in terms of (a) emerging themes and (b) sentiment and topic patterns?
2. Literature Review
2.1. The Integration of AI in Higher Education Assessment
2.2. Conceptualizing Student Preferences
2.3. Perceptions and Preferences for AI-Based Scoring and Feedback
2.4. The Impact of Interaction with AI on Student Perceptions and Preferences
3. Methods
3.1. Participants
3.2. Measure
3.3. Data Analyses
4. Results
4.1. RQ1. Preferences for AI Feedback Characteristics
4.2. RQ2. Preferences for AI Versus Human Scoring and Feedback
4.2.1. Preferences Across Scenarios
4.2.2. Preferences Across Assignment Types
4.3. RQ3. Perceptions of AI Scoring and Feedback
4.4. RQ4. Preferences for AI and AI Relative to Human Graders
4.4.1. Preferences for AI Scoring and Feedback
4.4.2. Preferences for AI Relative to Human Grading
4.5. RQ5. Post-Evaluation Preferences for AI Scoring and Feedback
4.5.1. Results from Thematic Analysis
4.5.2. Results from Sentiment Analysis and Topic Modelling
5. Discussion
5.1. RQ1. Preferences for AI Feedback Characteristics
5.2. RQ2. Preferences for AI Versus Human Scoring and Feedback
5.3. RQ3. Perceptions of AI Scoring and Feedback
5.4. RQ4. Preferences for AI and AI Relative to Human Graders
5.5. RQ5. Post-Evaluation Preferences for AI Scoring and Feedback
6. Conclusions
6.1. Implications
6.2. Limitations and Future Directions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AI | Artificial Intelligence |
| LDA | Latent Dirichlet Allocation |
Appendix A. Background & AI Familiarity
- What is your age? ______
- What is your gender?Male | Female | Non-binary | Prefer not to say
- What is your current academic level?First-year | Second-year | Third-year | Fourth-year or higher
- What is your major or area of study? ______
- What is your current GPA (Grade Point Average)?4.0 | 3.99–3.5 | 3.49–3.0 | 2.99–2.5 | 2.45–2.0 | 1.99 or below | Prefer not to disclose
- How comfortable are you with technology in general?Very uncomfortable | Somewhat uncomfortable | Somewhat comfortable | Very comfortable
- How familiar are you with artificial intelligence (AI)?Not familiar at all | Not very familiar | Somewhat familiar | Very familiar
- How familiar are you with AI scoring in writing tasks?Not familiar at all | Not very familiar | Somewhat familiar | Very familiar
- How familiar are you with AI-generated feedback in writing tasks?Not familiar at all | Not very familiar | Somewhat familiar | Very familiar
- How often do you receive feedback on your assignments from human graders?Daily | Weekly | Monthly | Occasionally | Rarely/Never
- How frequently do you use AI tools (e.g., ChatGPT, Grammarly)?Daily | Weekly | Monthly | Occasionally | Rarely/Never
- For which of the following purposes do you use AI tools (e.g., ChatGPT, Grammarly)? (Select all that apply)
- ⬤
- Enhancing learning and understanding
- ⬤
- Summarizing course content or materials
- ⬤
- Proofreading and correcting grammar, spelling, and punctuation
- ⬤
- Generating ideas or brainstorming
- ⬤
- Writing drafts or essays
- ⬤
- Language translation
- ⬤
- Other (please specify)
Appendix B. Task 1
Appendix C. AI Scoring and Feedback Preferences
- ⬤
- Score and feedback by AI
- ⬤
- Score by AI, feedback by a human
- ⬤
- Score by a human, feedback by AI
- ⬤
- Score and feedback by a human
- Which scoring and feedback method do you prefer in general for academic writing assignments?
- Which scoring and feedback method do you prefer for low-stakes assignments (e.g., practice essays, writing exercises)?
- Which scoring and feedback method do you prefer for high-stakes assignments (e.g., final papers, term projects)?
- Which scoring and feedback method do you prefer if the assignment involves creative or subjective writing (e.g., personal reflections)?
- Which scoring and feedback method do you prefer if you have the option to appeal or request a second review afterward?
- What tone do you prefer in AI-generated feedback?Encouraging | Neutral | Critical | A mix depending on the task
- What length of AI feedback do you prefer?Concise—short, to-the-point suggestionsDetailed—thorough explanations and examplesA combination of both—key points plus some detailed guidanceDepends on the assignment—concise for low-stakes, detailed for high-stakes
- What format do you prefer for AI feedback?Bulleted points | Paragraph-style | Inline comments on the text | A mix of formats
- When would you prefer to receive AI feedback in your writing process?After the first draft | After revisions | Before submission | At multiple stages
- Do you prefer AI feedback to focus most on:Content and ideas | Grammar and mechanics | Organization and structure | Style and tone
Appendix D. Scenarios
- ⬤
- Score and feedback by AI
- ⬤
- Score by AI, feedback by a human
- ⬤
- Score by a human, feedback by AI
- ⬤
- Score and feedback by a human
- Scenario 1: You are submitting a final research paper that counts for 40% of your overall course grade.
- Scenario 2: You submit a draft of an essay to receive feedback before the final version is due. The goal is to improve your work through revision.
- Scenario 3: You are asked to write a personal reflection that includes creative elements.
- Scenario 4: You are submitting a paper and know you will have the option to request a second review or appeal the score if needed.
- Scenario 5: You are submitting an essay on a highly technical topic where the accuracy of the content is critical.
- Scenario 6: Your instructor assigns a short, ungraded writing task to help you practice structuring arguments. You are told it will not impact your final grade.
- Scenario 7: You are submitting a final research paper that counts for 10% of your course grade.
- Scenario 8: You are nearing a submission deadline in one week. Your instructor offers an option to have your draft scored and receive feedback from an AI system within 24 h, allowing you to revise before the final submission. Alternatively, you can wait for human scoring and feedback, which will take 6 days.
- Scenario 9: You are asked to write a personal narrative about a difficult life experience for a psychology course.
- Scenario 10: You are a non-native English speaker submitting an argumentative essay for a writing-intensive course. You are concerned about misinterpretation of your phrasing or being penalized for grammar issues.
Appendix E. AI Perceptions
Appendix E.1. AI Scoring
- AI scoring is objective.
- AI scoring is consistent.
- AI scoring is transparent and explainable.
- AI can fairly grade diverse writing styles and voices.
- AI can accurately grade grammar and mechanics in writing.
- AI can accurately grade the organization and clarity of ideas in writing.
- AI can accurately grade complex or creative writing.
- AI assigns similar scores to similar quality work.
- AI grading might overlook important aspects of my writing. (reverse-coded)
- AI grading might struggle to interpret the tone or emotion in my writing. (reverse-coded)
Appendix E.2. AI Feedback
- AI feedback is easy to understand.
- AI feedback gives me clear next steps for improvement.
- AI feedback clearly identifies what I did well in my writing.
- AI feedback clearly identifies what I need to improve.
- AI feedback avoids vague or generic comments.
- AI feedback is relevant to the content of my writing.
- AI feedback helps me become more independent as a writer.
- AI feedback is too general to be useful. (reverse-coded)
- AI feedback is repetitive or redundant. (reverse-coded)
- AI feedback sometimes contradicts itself. (reverse-coded)
Appendix F. AI Preferences
Appendix F.1. Scoring Process Preferences
- I prefer to see a breakdown of how the AI grades each part of my writing.
- I prefer to receive a confidence level or accuracy rating along with the AI scoring and feedback.
- I prefer to have the choice to opt in or opt out of AI scoring and feedback.
- I prefer AI if it includes a justification based on a clear rubric.
- I prefer AI if its decision-making process is fully transparent.
- I prefer AI to only offer feedback, not assign scores.
- I prefer AI feedback to highlight both strengths and weaknesses in my writing.
- I prefer AI feedback to reference specific examples from my writing.
- I prefer AI feedback to include links or resources for further learning.
- I prefer AI to allow me to ask follow-up questions about its score and feedback.
Appendix F.2. AI vs. Human Comparison
- I prefer AI over a human grader for low-stakes assignments (e.g., practice essays).
- I prefer AI over a human grader for high-stakes assignments (e.g., final papers).
- I prefer AI scoring and feedback to always be reviewed by a human.
- I prefer to receive AI-generated feedback but human-assigned scores.
- I prefer that AI be used only to assist human graders, not replace them.
- I prefer my score to be equally based on AI and human evaluation.
- I prefer AI for fast self-assessment before submission, instead of waiting for a human grader.
- I prefer AI involvement in final grading, rather than relying entirely on a human grader.
- I prefer to see a side-by-side comparison of AI and human scoring and feedback.
- I prefer AI scoring and feedback first, then human, to maximize both speed and insight.
Appendix G. Task 2
- ⬤
- Copy the question, scoring rubric, and short student response below.
- ⬤
- Paste them into ChatGPT and ask it to: “Provide a score and feedback for this response based on the rubric”.
- ⬤
- Review the AI’s score and feedback. Once done, return to the survey and answer the question below:
- ⬤
- Please copy and paste the score and feedback generated by ChatGPT. Then, reflect on your impressions: To what extent did you find the score accurate? Was the feedback clear and constructive? In your view, how useful (or not) might AI-generated scoring and feedback be for students?
- ⬤
- 4 points—Clear, well-developed argument with specific examples; logical structure; minimal grammar errors.
- ⬤
- 3 points—Good argument but may lack depth or specific examples; mostly clear; few grammar errors.
- ⬤
- 2 points—Argument is underdeveloped or unclear in places; limited examples; some grammar errors.
- ⬤
- 1 point—Weak or unclear argument; few or no examples; multiple grammar issues.
- ⬤
- 0 points—Off-topic, incomprehensible, or no answer provided.
References
- Al Harrasi, N. H., & El Din, M. S. (Eds.). (2024). Utilizing AI for assessment, grading, and feedback in higher education. IGI Global. [Google Scholar]
- Bouchet-Valat, M. (2023). SnowballC: Snowball stemmers based on the C ‘libstemmer’ (R package version 0.7.1). CRAN. Available online: https://CRAN.R-project.org/package=SnowballC (accessed on 25 January 2026).
- Braun, V., & Clarke, V. (2006). Using thematic analysis in psychology. Qualitative Research in Psychology, 3(2), 77–101. [Google Scholar] [CrossRef] [Scilit]
- Braun, V., & Clarke, V. (2019). Reflecting on reflexive thematic analysis. Qualitative Research in Sport, Exercise and Health, 11(4), 589–597. [Google Scholar] [CrossRef] [Scilit]
- Bryer, J., & Speerschneider, K. (2016). likert: Analysis and visualization of Likert items (R package version 1.3.5). CRAN. Available online: https://CRAN.R-project.org/package=likert (accessed on 25 January 2026).
- Buzick, H., Oliveri, M. E., Attali, Y., & Flor, M. (2016). Comparing human and automated essay scoring for prospective graduate students with learning disabilities and/or ADHD. Applied Measurement in Education, 29(3), 161–172. [Google Scholar] [CrossRef] [Scilit]
- Carless, D., & Boud, D. (2018). The development of student feedback literacy: Enabling uptake of feedback. Assessment & Evaluation in Higher Education, 43(8), 1315–1325. [Google Scholar] [CrossRef] [Scilit]
- Chai, F., Ma, J., Wang, Y., Zhu, J., & Han, T. (2024). Grading by AI makes me feel fairer? How different evaluators affect college students’ perception of fairness. Frontiers in Psychology, 15, 1221177. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chan, C. K. Y., & Hu, W. (2023). Students’ voices on generative AI: Perceptions, benefits, and challenges in higher education. International Journal of Educational Technology in Higher Education, 20, 43. [Google Scholar] [CrossRef] [Scilit]
- Davis, F. D. (1989). Perceived usefulness, perceived ease of use, and user acceptance of information technology. MIS Quarterly, 13(3), 319–340. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Dietvorst, B. J., Simmons, J. P., & Massey, C. (2015). Algorithm aversion: People erroneously avoid algorithms after seeing them err. Journal of Experimental Psychology: General, 144(1), 114–126. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Er, E., Akcapinar, G., Bayazit, A., Noroozi, O., & Banihashem, S. K. (2025). Assessing student perceptions and use of instructor versus AI-generated feedback. British Journal of Educational Technology, 56(3), 1074–1091. [Google Scholar] [CrossRef] [Scilit]
- Feinerer, I., & Hornik, K. (2025). tm: Text mining package (R Package Version 0.7-17). CRAN. Available online: https://CRAN.R-project.org/package=tm (accessed on 25 January 2026).
- Flodén, J. (2025). Grading exams using large language models: A comparison between human and AI grading of exams in higher education using ChatGPT. British Educational Research Journal, 51(1), 201–224. [Google Scholar] [CrossRef] [Scilit]
- Fu, Q. K., Zou, D., Xie, H., & Cheng, G. (2022). A review of AWE feedback: Types, learning outcomes, and implications. Computer Assisted Language Learning, 37(1–2), 179–221. [Google Scholar] [CrossRef] [Scilit]
- Gaube, S., Suresh, H., Raue, M., Merritt, A., Berkowitz, S. J., Lermer, E., Coughlin, J. F., Guttag, J. V., & Ghassemi, M. (2021). Do as AI say: Susceptibility in deployment of clinical decision-aids. npj Digital Medicine, 4(1), 31. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Grün, B., & Hornik, K. (2011). topicmodels: An R package for fitting topic models. Journal of Statistical Software, 40(13), 1–30. [Google Scholar] [CrossRef] [Scilit]
- Hahn, M. G., Navarro, S. M. B., De La Fuente Valentín, L., & Burgos, D. (2021). A systematic review of the effects of automatic scoring and automatic feedback in educational settings. IEEE Access, 9, 108190–108198. [Google Scholar] [CrossRef] [Scilit]
- Hockly, N. (2019). Automated writing evaluation. ELT Journal, 73(1), 82–88. [Google Scholar] [CrossRef] [Scilit]
- Jin, F. J.-Y., Maheshi, B., Lai, W., Li, Y., Gasevic, D., Chen, G., Charwat, N., Chan, P. W. K., Martinez-Maldonado, R., Gašević, D., & Tsai, Y.-S. (2025). Students’ perceptions of generative AI–powered learning analytics in the feedback process: A feedback literacy perspective. Journal of Learning Analytics, 12(1), 152–168. [Google Scholar] [CrossRef] [Scilit]
- Kasneci, E., Sessler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., Günnemann, S., Hüllermeier, E., Krusche, S., Kutyniok, G., Michaeli, T., Nerdel, C., Pfeffer, J., Poquet, O., Sailer, M., Schmidt, A., Seidel, T., … Kasneci, G. (2023). ChatGPT for good? On opportunities and challenges of large language models for education. Learning and Individual Differences, 103, 102274. [Google Scholar] [CrossRef] [Scilit]
- Kelly, A., Sullivan, M., & Strampel, K. (2023). Generative artificial intelligence: University student awareness, experience, and confidence in use across disciplines. Journal of University Teaching & Learning Practice, 20(5), 12. [Google Scholar] [CrossRef] [Scilit]
- Kotlyar, I., & Krasman, J. (2025). Student reactions to AI versus human feedback in teamwork skills assessment. International Journal of Educational Technology in Higher Education, 22, 57. [Google Scholar] [CrossRef] [Scilit]
- Kumar, V. S., & Boulanger, D. (2021). Automated essay scoring and the deep learning black box: How are rubric scores determined? International Journal of Artificial Intelligence in Education, 31(3), 538–584. [Google Scholar] [CrossRef] [Scilit]
- Le, H., Shen, Y., Li, Z., Xia, M., Tang, L., Li, X., Jia, J., Wang, Q., Gašević, D., & Fan, Y. (2025). Breaking human dominance: Investigating learners’ preferences for learning feedback from generative AI and human tutors. British Journal of Educational Technology, 56(5), 1758–1783. [Google Scholar] [CrossRef] [Scilit]
- Lee, A. V. Y., Luco, A. C., & Tan, S. C. (2023). A human-centric automated essay scoring and feedback system for the development of ethical reasoning. Educational Technology & Society, 26(1), 147–159. Available online: https://www.jstor.org/stable/48707973 (accessed on 25 January 2026).
- Lipnevich, A. A., Berg, D. A. G., & Smith, J. K. (2016). Toward a model of student response to feedback. In G. T. L. Brown, & L. R. Harris (Eds.), Handbook of human and social conditions in assessment (pp. 169–185). Routledge. [Google Scholar] [CrossRef] [Scilit]
- Long, D., & Magerko, B. (2020). What is AI literacy? Competencies and design considerations. In Proceedings of the 2020 CHI conference on human factors in computing systems (pp. 1–16). Association for Computing Machinery. [Google Scholar] [CrossRef] [Scilit]
- Molloy, E., Boud, D., & Henderson, M. (2020). Developing a learning-centred framework for feedback literacy. Assessment & Evaluation in Higher Education, 45(4), 527–540. [Google Scholar] [CrossRef] [Scilit]
- Nazaretsky, T., Mejia-Domenzain, P., Swamy, V., Frej, J., & Käser, T. (2025). The critical role of trust in adopting AI-powered educational technology for learning: An instrument for measuring student perceptions. Computers & Education: Artificial Intelligence, 8, 100368. [Google Scholar] [CrossRef] [Scilit]
- Ngo, T. T. A. (2023). The perception by university students of the use of ChatGPT in education. International Journal of Emerging Technologies in Learning, 18(4), 4–19. [Google Scholar] [CrossRef] [Scilit]
- R Core Team. (2020). R: A language and environment for statistical computing. R Foundation for Statistical Computing. Available online: https://www.R-project.org/ (accessed on 25 January 2026).
- Shoufan, A. (2023). Exploring students’ perceptions of ChatGPT: Thematic analysis and follow-up survey. IEEE Access, 11, 38805–38818. [Google Scholar] [CrossRef] [Scilit]
- Siregar, R., Subagiharti, H., Handayani, D. S., Sutarno, S., Hasibuan, A. L., & Barus, E. (2024). Student preferences on using artificial intelligence (AI) platform in language learning. International Journal of Educational Research Excellence, 3(2), 746–754. [Google Scholar] [CrossRef] [Scilit]
- Slepankova, M., Kilianova, K., Kockova, P., Kostolanyova, K., Kotyrba, M., & Habiballa, H. (2025). Student perceptions and preferences in personalized AI-driven learning. Acta Informatica Pragensia, 14(2), 261–271. [Google Scholar] [CrossRef] [Scilit]
- Stowell, J. R., & Zhu, J. (2025). Evaluating ChatGPT for automated creation and grading of essay questions in higher education. Teaching of Psychology. Advance online publication. [Google Scholar] [CrossRef] [Scilit]
- Thomas, M. L., Yildirim-Erbasli, S. N., & Hariharan, S. (2026). Exploring undergraduate students’ perceptions of AI vs. human scoring and feedback. The Internet and Higher Education, 68, 101052. [Google Scholar] [CrossRef] [Scilit]
- Tossell, C. C., Tenhundfeld, N. L., Momen, A., Cooley, K., & de Visser, E. J. (2024). Student perceptions of ChatGPT use in a college essay assignment: Implications for learning, grading, and trust in artificial intelligence. IEEE Transactions on Learning Technologies, 17, 1069–1081. [Google Scholar] [CrossRef] [Scilit]
- Venkatesh, V., Morris, M. G., Davis, G. B., & Davis, F. D. (2003). User acceptance of information technology: Toward a unified view. MIS Quarterly, 27(3), 425–478. [Google Scholar] [CrossRef] [Scilit]
- Von Garrel, J., & Mayer, J. (2023). Artificial intelligence in studies: Use of ChatGPT and AI-based tools among students in Germany. Humanities and Social Sciences Communications, 10, 799. [Google Scholar] [CrossRef] [Scilit]
- Wickham, H. (2016). ggplot2: Elegant graphics for data analysis. Springer. [Google Scholar]
- Wickham, H., François, R., Henry, L., Müller, K., & Vaughan, D. (2023). dplyr: A grammar of data manipulation. Available online: https://CRAN.R-project.org/package=dplyr (accessed on 25 January 2026).
- Wickham, H., & Henry, L. (2020). tidyr: Tidy messy data. Available online: https://CRAN.R-project.org/package=tidyr (accessed on 25 January 2026).
- Wilson, J., Huang, Y., Palermo, C., Beard, G., & MacArthur, C. A. (2021). Automated feedback and automated scoring in the elementary grades: Usage, attitudes, and associations with writing outcomes in a districtwide implementation of MI write. International Journal of Artificial Intelligence in Education, 31(2), 234–276. [Google Scholar] [CrossRef] [Scilit]
- Yan, X., Rupp, A. A., & Foltz, P. W. (2020). Handbook of automated scoring: Applications and advances. Routledge. [Google Scholar]
- Zawacki-Richter, O., Marín, V. I., Bond, M., & Gouverneur, F. (2019). Systematic review of research on artificial intelligence applications in higher education—Where are the educators? International Journal of Educational Technology in Higher Education, 16(1), 39. [Google Scholar] [CrossRef] [Scilit]
- Zhan, Y., Boud, D., Dawson, P., & Yan, Z. (2025). Generative artificial intelligence as an enabler of student feedback engagement: A framework. Higher Education Research & Development, 44(5), 1289–1304. [Google Scholar] [CrossRef] [Scilit]



| Variable | Response Category | % (n) |
|---|---|---|
| Comfort with technology | Very uncomfortable | 10.8 (10) |
| Somewhat uncomfortable | 14.0 (13) | |
| Somewhat comfortable | 51.6 (48) | |
| Very comfortable | 23.7 (22) | |
| Familiarity with AI | Not familiar at all | 3.2 (3) |
| Not very familiar | 14.0 (13) | |
| Somewhat familiar | 62.4 (58) | |
| Very familiar | 20.4 (19) | |
| Familiarity with AI scoring | Not familiar at all | 19.4 (18) |
| Not very familiar | 41.9 (39) | |
| Somewhat familiar | 33.3 (31) | |
| Very familiar | 5.4 (5) | |
| Familiarity with AI feedback | Not familiar at all | 15.1 (14) |
| Not very familiar | 38.7 (36) | |
| Somewhat familiar | 37.6(35) | |
| Very familiar | 8.6 (8) | |
| Frequency of receiving human feedback | Daily | 2.2 (2) |
| Weekly | 35.5 (33) | |
| Monthly | 34.4 (32) | |
| Occasionally | 24.7 (23) | |
| Rarely/Never | 3.2 (3) | |
| Frequency of AI use | Daily | 9.7 (9) |
| Weekly | 28.0 (26) | |
| Monthly | 10.8 (10) | |
| Occasionally | 31.2 (29) | |
| Rarely/Never | 20.4 (19) | |
| Purpose of AI use | Enhancing learning and understanding | 66.7 (62) |
| Summarizing course content or materials | 53.8 (50) | |
| Proofreading, correcting grammar, spelling, & punctuation | 47.3 (44) | |
| Generating ideas or brainstorming | 49.5 (46) | |
| Writing drafts or essays | 9.7 (9) | |
| Language translation | 28 (26) |
| Feedback Dimension | Response Option | % (n) |
|---|---|---|
| Tone | Encouraging | 14.0 (13) |
| Neutral | 22.6 (21) | |
| Critical | 21.5 (20) | |
| A mix depending on the task | 41.9 (39) | |
| Length | Concise (short, to-the-point suggestions) | 14.0 (13) |
| Detailed (thorough explanations and examples) | 30.1 (28) | |
| A combination of both (key points & detailed guidance) | 36.6 (34) | |
| Depends on the assignment (concise for low-stakes, detailed for high-stakes) | 19.4 (18) | |
| Format | Bulleted points | 43.0 (40) |
| Paragraph-style | 7.5 (7) | |
| Inline comments on the text | 15.1 (14) | |
| A mix of formats | 34.4 (32) | |
| Timing | After the first draft | 26.9 (25) |
| After revisions | 16.1 (15) | |
| Before submission | 22.6 (21) | |
| At multiple stages | 34.4 (32) | |
| Focus | Content and ideas | 24.7 (23) |
| Organization and structure | 30.1 (28) | |
| Grammar and mechanics | 37.6 (35) | |
| Style and tone | 7.5 (7) |
| Scenario | Human Score & Feedback | AI Score & Feedback | Human Score & AI Feedback | AI Score & Human Feedback |
|---|---|---|---|---|
| Final paper (40%) | 76 (81.7%) | 4 (4.3%) | 8 (8.6%) | 5 (5.4%) |
| Draft for revision | 41 (44.1%) | 13 (14.0%) | 21 (22.6%) | 18 (19.4%) |
| Creative reflection | 66 (71.0%) | 3 (3.2%) | 13 (14.0%) | 11 (11.8%) |
| Paper with an appeal option | 50 (53.8%) | 7 (7.5%) | 21 (22.6%) | 15 (16.1%) |
| Technical essay | 64 (68.8%) | 6 (6.5%) | 11 (11.8%) | 12 (12.9%) |
| Short ungraded task | 42 (42.5%) | 22 (23.7%) | 16 (17.2%) | 13 (14.0%) |
| Final paper (10%) | 60 (64.5%) | 7 (7.5%) | 14 (15.1%) | 12 (12.9%) |
| Near-deadline AI option | 26 (28.0%) | 35 (37.6%) | 18 (19.4%) | 14 (15.1%) |
| Personal narrative | 67 (72.0%) | 8 (8.6%) | 9 (9.7%) | 9 (9.7%) |
| Non-native English essay | 50 (53.8%) | 15 (16.1%) | 18 (19.4%) | 10 (10.8%) |
| Assignment Type | Human Score & Feedback | AI Score & Feedback | Human Score & AI Feedback | AI Score & Human Feedback |
|---|---|---|---|---|
| General academic writing | 74 (79.6%) | 2 (2.2%) | 9 (9.7%) | 8 (8.6%) |
| Low-stakes assignments | 50 (53.8%) | 7 (7.5%) | 17 (18.3%) | 19 (20.4%) |
| High-stakes assignments | 79 (84.9%) | 1 (1.1%) | 7 (7.5%) | 6 (6.5%) |
| Creative/subjective writing | 66 (71.0%) | 4 (4.3%) | 17 (18.3%) | 6 (6.5%) |
| Assignments with an appeal option | 63 (67.7%) | 4 (4.3%) | 19 (20.4%) | 7 (7.5%) |
| Predictor | B | SE | t | p | OR |
|---|---|---|---|---|---|
| Comfort with technology | |||||
| Very uncomfortable | 1.28 | 1.17 | 1.09 | 0.276 | 3.59 |
| Somewhat uncomfortable | −0.59 | 0.74 | −0.81 | 0.419 | 0.55 |
| Very comfortable | −1.20 | 0.63 | −1.90 | 0.058 | 0.30 |
| Familiarity with AI (general) | |||||
| Not very familiar | 0.60 | 1.52 | 0.39 | 0.695 | 1.82 |
| Somewhat familiar | −0.17 | 1.50 | −0.11 | 0.911 | 0.85 |
| Very familiar | 0.58 | 1.63 | 0.36 | 0.720 | 1.79 |
| Familiarity with AI scoring | |||||
| Not very familiar | 0.41 | 1.14 | 0.36 | 0.717 | 1.51 |
| Somewhat familiar | −0.15 | 1.23 | −0.12 | 0.902 | 0.86 |
| Very familiar | 0.44 | 1.81 | 0.24 | 0.807 | 1.56 |
| Familiarity with AI feedback | |||||
| Not very familiar | 0.31 | 1.12 | 0.27 | 0.785 | 1.36 |
| Somewhat familiar | 1.23 | 1.22 | 1.00 | 0.316 | 3.41 |
| Very familiar | 0.05 | 1.60 | 0.03 | 0.974 | 1.05 |
| Frequency of receiving human feedback | |||||
| Frequent | 0.35 | 0.62 | 0.55 | 0.580 | 1.41 |
| Infrequent | 0.52 | 0.69 | 0.75 | 0.455 | 1.68 |
| Frequency of AI use | |||||
| Weekly | 1.28 | 0.96 | 1.33 | 0.184 | 3.59 |
| Monthly | 1.87 | 1.15 | 1.63 | 0.103 | 6.50 |
| Occasionally | 1.24 | 0.92 | 1.35 | 0.179 | 3.46 |
| Rarely/Never | 2.11 | 1.03 | 2.05 | 0.040 * | 8.24 |
| Predictor | B | SE | t | p | OR |
|---|---|---|---|---|---|
| Comfort with technology | |||||
| Very uncomfortable | −0.77 | 0.81 | −0.95 | 0.343 | 0.47 |
| Somewhat uncomfortable | −1.37 | 0.68 | −2.01 | 0.045 * | 0.26 |
| Very comfortable | −0.85 | 0.61 | −1.40 | 0.161 | 0.43 |
| Familiarity with AI (general) | |||||
| Not very familiar | −0.01 | 1.36 | −0.01 | 0.994 | 0.99 |
| Somewhat familiar | 0.54 | 1.35 | 0.40 | 0.688 | 1.72 |
| Very familiar | −0.67 | 1.45 | −0.46 | 0.644 | 0.51 |
| Familiarity with AI scoring | |||||
| Not very familiar | −0.58 | 1.02 | −0.57 | 0.569 | 0.56 |
| Somewhat familiar | −0.05 | 1.09 | −0.04 | 0.964 | 0.95 |
| Very familiar | 1.60 | 1.81 | 0.88 | 0.376 | 4.96 |
| Familiarity with AI feedback | |||||
| Not very familiar | −0.21 | 0.99 | −0.21 | 0.836 | 0.81 |
| Somewhat familiar | 0.50 | 1.09 | 0.46 | 0.646 | 1.65 |
| Very familiar | −0.29 | 1.51 | −0.19 | 0.847 | 0.75 |
| Frequency of receiving human feedback | |||||
| Frequent | 1.24 | 0.58 | 2.13 | 0.033 * | 3.45 |
| Infrequent | −0.38 | 0.62 | −0.61 | 0.545 | 0.69 |
| Frequency of AI use | |||||
| Weekly | −0.03 | 0.91 | −0.04 | 0.970 | 0.97 |
| Monthly | 0.53 | 1.09 | 0.49 | 0.625 | 1.70 |
| Occasionally | −0.08 | 0.86 | −0.09 | 0.930 | 0.93 |
| Rarely/Never | 0.34 | 0.95 | 0.36 | 0.719 | 1.41 |
| Theme | Description | n |
|---|---|---|
| AI as a formative learning aid | Preference for AI as a tool for learning, feedback, or improvement | 32 |
| Human authority | Preference for human judgment, nuance, emotional understanding, or final control | 17 |
| Context-dependent preference | Acceptance depends on the assignment type, stakes, creativity, or discipline | 9 |
| Skepticism toward educational impact | Concerns about harm to learning, fairness, bias, inconsistency, or critical thinking | 18 |
| Topic | Positive Words | Negative Words | Sentiment Score |
|---|---|---|---|
| Detailed feedback and examples | 38 | 8 | 30 |
| Usefulness and clarity of feedback | 75 | 4 | 71 |
| Accuracy of scoring and feedback | 73 | 3 | 70 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Yildirim-Erbasli, S.N.; Ilgun Dibek, M.; Thomas, M.L.; Lesoway, N. When Do Undergraduate Students Prefer AI? Insights into AI Scoring and Feedback. Behav. Sci. 2026, 16, 1196. https://doi.org/10.3390/bs16071196
Yildirim-Erbasli SN, Ilgun Dibek M, Thomas ML, Lesoway N. When Do Undergraduate Students Prefer AI? Insights into AI Scoring and Feedback. Behavioral Sciences. 2026; 16(7):1196. https://doi.org/10.3390/bs16071196
Chicago/Turabian StyleYildirim-Erbasli, Seyma N., Munevver Ilgun Dibek, Mackenzie L. Thomas, and Nicolya Lesoway. 2026. "When Do Undergraduate Students Prefer AI? Insights into AI Scoring and Feedback" Behavioral Sciences 16, no. 7: 1196. https://doi.org/10.3390/bs16071196
APA StyleYildirim-Erbasli, S. N., Ilgun Dibek, M., Thomas, M. L., & Lesoway, N. (2026). When Do Undergraduate Students Prefer AI? Insights into AI Scoring and Feedback. Behavioral Sciences, 16(7), 1196. https://doi.org/10.3390/bs16071196

