Fair Marking in the Generative AI Era: Introducing the Master’s Dissertation Marking Framework
Abstract
1. Introduction
1.1. Background
1.1.1. Marking Reliability and Fairness
1.1.2. Rubrics, Exemplars and Judgement
1.1.3. Training on How to Use the Rubric
1.1.4. Moderation and Standards
1.1.5. Marking Practices and Interpretive Bias
1.1.6. Generative AI and Feedback/Assessment
1.2. Study Rationale and Research Questions
- RQ1: What challenges affect fairness, consistency, and judgment in master’s dissertation marking?
- RQ2: How do these challenges evolve in the context of generative AI?
- RQ3: How can a structured, empirically grounded framework support fairer and more defensible dissertation assessment practices?
2. Research Design and Methodology
2.1. Research Design: A Design-Based Research Approach
Overview of Iterative Research Phases
2.2. Participants
2.3. Data Collection Methods
- Phase 2 (2022). A survey of 31 academic staff explored challenges in dissertation marking, including fairness, bias, and marking practices.
- Phase 4 (2023). Two exploratory surveys (n = 19; n = 8) captured early responses to the emergence of generative AI prior to the establishment of institutional policies.
- Phase 6 (2025). A mixed-methods survey collected both quantitative data (Likert-scale items on AI effectiveness, concerns, and support) and qualitative data (open-ended responses on marking challenges and AI impact).
- Phase 3 (2022): Seven academic staff participated in interviews exploring marking culture, fairness, bias, and the use of marking schemes.
- Phase 5 (2025): Three follow-up interviews were conducted with experienced markers after at least one full dissertation marking cycle in the post-generative-AI context. These interviews focused on changes in marking practices, perceptions of AI, and emerging fairness concerns.
2.4. Data Analysis
2.5. Framework Development and Iterative Refinement
- MDMF v1: Derived from literature (conceptual foundation)
- MDMF v2–v3: Refined using staff perceptions (survey and interviews)
- MDMF v4: Updated to incorporate AI-related challenges emerging from early post-ChatGPT 3.5 and 4 contexts
- MDMF v5: Final refinement based on mixed-methods evidence from the post-generative-AI environment
3. Master’s Dissertation Marking Framework Design Cycles in 2022
3.1. MDMF Version 1—Initial Draft Based on the Literature Review
3.2. MDMF Version 2—Framework Updates Based on Fact-Finding Study 1 (Survey Conducted in 2022)
3.3. MDMF Version 3—Framework Updates Based on Fact-Finding Study 2 (Interviews Conducted in 2022)
4. Transitional Phase Evidence: Early Staff Responses to Generative AI in Dissertation Marking in 2023
5. Post-Marking Practice Evidence in 2025: Understanding Shifts in Dissertation Marking in the AI-Dominated Era
5.1. Results from the Early 2025 Interview Data
5.1.1. Continuities in Marking Practices and Persistent Challenges
5.1.2. Post-ChatGPT Shifts: AI Awareness, Ethical Ambiguities, and Reframing Fairness
5.1.3. AI’s Disruptive Impact on the Originality and Academic Integrity of the Dissertation
5.1.4. Resistance to Using AI Tools for Marking Dissertations
5.1.5. Fairness Reframed: From Procedural Consistency to AI-Mediated Equity
5.1.6. Language Bias Reinterpreted in an AI Context
5.2. MDMF Version 4—Framework Updates Based on Early 2025 Study
5.2.1. Core Component 1: Ethical Boundaries and Policy Clarity in the Age of AI
| Sub-Components | Elements |
|---|---|
| Policy gaps | Co-develop clear, student-facing AI use policies Distinguish acceptable vs. unacceptable uses (e.g., grammar vs. argument generation) Align policies with existing integrity frameworks. |
| Ethical ambiguity | Introduce AI use declarations as a standard for dissertations Provide marker guidance on how to handle undeclared or suspicious AI use scenarios. |
| Marker tension | Develop AI-sensitivity checklists for markers. Foster reflective discussions on the boundaries of responsible AI use in research writing. |
| Disciplinary variation | Encourage departments to develop discipline-specific AI guidelines that align with broader policy. Use these to inform marking schemes and moderation approaches. |
| Shared accountability | Promote a shared responsibility model: Institutions provide guidance, markers exercise judgment, and students disclose their AI use transparently. |
| Equity and Inclusion | Include equity lens in policy-making (e.g., ensure access to AI tools is considered) Support multilingual and neurodiverse students in understanding AI guidelines. |
| Documentation and Dispute Resolution | Develop transparent documentation protocols for suspected AI use. Include processes for student appeals, staff consultation, and third-party review if needed. |
5.2.2. Core Component 2: Identifying Issues Related to Fairness and Equity in Master’s Dissertation Marking
| Sub-Components | Elements |
|---|---|
| Project Dissertation | Disparity in complexity Disparity in requirements Different types (development, research, hybrid (dev + res), unclassified) Not knowing the student contribution AI-assisted work further obscures student contribution * Difficulty in judging independent thinking and authenticity without explicit declaration * |
| Marker Characteristics | Supervisor bias Reader bias Intercultural differences Marking based on role Lack of subject knowledge Language bias Fluency bias amplified by AI tools * Greater cognitive dissonance due to AI-mediated language quality * Emotional reaction to AI-sounding text * |
| Marking Process | High marking load Lack of anonymity Reduced marking time Marking time clashes with other markings Speed up marking due to time constraints Comparing dissertations Fatigue from reviewing similar-sounding AI-generated work * Suspicion of AI use without evidence * Emotional disengagement from similar language patterns * |
| Marking Schemes | Too long Vagueness Irrelevant Not adapted to assess AI-mediated work * Lack of guidance for handling AI involvement * No differentiation between AI-generated and human voice * |
| Marking Technology | Too many systems Needs Improvement No integrated AI-detection tools * Risk of data misuse (e.g., student work used to train AI) * No tools to support transparency or traceability of AI use * |
| Marker Well-being | Stressed, worried, disappointed, resentful, tired, confused Increased emotional fatigue from AI-authored text * Moral discomfort marking suspected AI work * Frustration at policy voids and shifting norms * |
| Institutional Policy & Ethics (cross-cutting issue linked to Component 1) * | (New in V4—For full detail, see Component 6—Ethical Boundaries and Policy Clarity) No consensus on acceptable AI use * Tension between innovation and integrity * Need for clearer policy and student-facing guidelines * |
5.2.3. Core Component 3: Pre-Marking Tasks
| Sub-Components | Elements |
|---|---|
| Marking Schemes | Detailed, clear and calibrated: Break down the project into similar parts as marking. Clarify which dissertation chapter corresponds to each section of marking. Consider the subjectivity of descriptors that explain the grade using terms such as ‘excellent,’ ‘good,’ etc. Provide examples of corresponding dissertation sections and associated marks (to ensure quality and standards). Separate marking schemes for development, research-based, hybrid (dev + res) and others Consider the length of marking schemes (interactive documents could help). Explicit AI-related criteria or guidance (e.g., evaluating AI-assisted work) * Examples of how AI use affects structure or tone * Clarify how originality and critical thinking should be assessed in the AI era * |
| Exemplars | Organised by research fields, example titles can be included in marking schemes. Exemplars should include grades and feedback (seen only by markers) Exemplars could include AI-assisted and non-AI-assisted versions (for comparison)* Tag exemplars with critical thinking quality, not just polish or fluency * |
| Training for Marking and Moderation | Consider the length of training. Possible training content: What is the expected standard? Practical examples of marking using the marking schemes; how to deal with “difficult characters” during negotiation, and how to use the marking system (s). Mentoring: for new markers. Training to detect possible signs of AI-generated content (e.g., voice consistency, fabricated references) * Discussions of bias and fairness in AI-mediated language * Handling undeclared AI use in line with ethical policy * |
| Calibration Practices | Discussion round for each project within each research group Include AI-awareness calibration, e.g., how AI usage affects language, criticality, and structure * Share edge cases (e.g., suspicious fluency vs. weak reasoning) * |
| Good Practice Guidelines | Collaborative authorship by all marking staff Updated regularly to reflect institutional AI policy * Include decision-making flowcharts for suspected or declared AI involvement * |
5.2.4. Core Component 4: Assigning Marking Based on Marker Characteristics
| Sub-Components | Elements |
|---|---|
| Marker Roles & Marking Models | Supervisors summarise the project, including project field (research group or area), type (i.e., research), project difficulty level and student performance. Readers select a subject area or research group (all projects within that field appear), and dissertation type (research/development). Readers bid for projects and declare their level of expertise in the project topic. Consider the marking load. Add AI expertise as part of expertise declaration (e.g., ability to detect or handle AI-generated work) * Factor in emotional load when assigning AI-heavy disciplines or tasks *—Emotional implications of sustained AI exposure are further addressed under Emotional Engagement in Table 5: Marking Process. |
| Marker cognitive profile | Pairing readers and supervisors based on subjects’ knowledge. Pairing third markers to dissertations based on subject knowledge. If pairing the supervisor with a non-subject expert or non-experienced reader, the reader should receive additional information on example dissertations in that area (or extra guidance). If not paired with a subject expert, provide extended support to include examples of AI-influenced dissertations (to help identify patterns). * Address language-bias risk by balancing native and non-native speakers across pairings (where necessary). * |
| Equity and Transparency (New) * | Prevent hidden bias against AI-generated fluency (e.g., polished but superficial content). * Include reflective calibration sessions to align perceptions of AI use and authenticity. * See also Marking Culture in Table 5: MDMF V4—Marking Process for practical procedures to support alignment during assessment. |
5.2.5. Core Component 5: Marking Process
| Sub-Components | Elements |
|---|---|
| Marking Consistency | Blind marking, double marking. Consider using a marking process, such as conference and journal review processes Use marking schemes. Readers to consider student supplement material for marking (i.e., demonstration video) Consider AI detection uncertainty: introduce flagging protocols for suspected AI use. * Ensure schemes clearly define originality, criticality, and engagement in AI-mediated texts. * |
| Moderation | Blind negotiation (i.e., in a chat instead of email) Documentation of the negotiation process Blind third marking: If necessary, the third marker should receive AI-context training on how to assess suspected generated content fairly. * Moderation discussions include AI use declarations or absences. * |
| Marking Culture | Discussion and debates (i.e., round-the-table debates instead of emails) Create space for AI-focused calibration discussions (e.g., signs of AI use, thresholds for concern). For example, use anonymised samples and structured discussions to align marker judgments on AI use, especially regarding fluency vs. criticality. *—see Table 4: Equity and Transparency for assessor alignment practices during marker assignment. Address marker fatigue or frustration with repeated AI-generated language. * Reflective calibration sessions to address and align perceptions of AI use and authorship. * Prevent hidden bias against polished but superficial AI-generated fluency. * |
| Emotional Engagement & Well-being (new) * | Recognise fatigue from marking similar-sounding or impersonal AI-generated text. * Offer institutional debriefs or informal spaces for assessors to reflect on emotionally or ethically challenging marking situations. * (See also Table 4: Marker Roles and Cognitive profile for how emotional well-being is considered during the allocation of AI-heavy projects.) |
5.2.6. Core Component 6: Technology
| Sub-Components | Elements |
|---|---|
| Technology as Enabler | Use platforms to capture project/assessor characteristics Aid mapping of markers to dissertations Improve anonymity Integrate AI-awareness features, e.g., space for AI use declarations * Add dashboards to flag patterns in submission styles or language that may suggest AI involvement * |
| Systems Integration | Too many systems Push for unified platforms that allow end-to-end marking, moderation, and communication. * Consider carefully governed tools that support human review of possible AI involvement, while recognising the limitations of automated detection. * |
| Transparency & Traceability | Store and log AI-use declarations, any moderation discussions around AI suspicion, and actions taken. * Allow audit trails for disputed cases (e.g., if misconduct is suspected but not confirmed) * |
| Ethical Safeguards | Avoid training AI models on student data without consent * Clarify data retention, privacy, and consent processes * Avoid building in biases that privilege fluency or stylistic uniformity * |
| Equity Enhancements | Use tech to monitor whether markers are unintentionally favouring AI-assisted fluency. * Highlight discrepancies between content quality and linguistic style as a potential indicator of equity. * |
| Marker Support | Use platforms to track emotional load or marking volume per assessor. * Enable markers to flag dissertations that require support (e.g., suspected AI overuse, unclear authorship, confusing topic complexity) * |
6. Academic Staff Perceptions Regarding the Potential Role of AI Tools in the MSc Dissertation Marking Process at the End of 2025
6.1. Results
6.1.1. Persistent Structural Challenges in Dissertation Marking
6.1.2. Fairness, Bias, and AI-Mediated Distortions of Judgment
“I feel like projects all tend towards a B by old standards due to ubiquitous LLM use… all are filtered through generic competence. I want to give students who visibly did not use LLMs more credit, even if their work is less well argued, because I know it was their thoughts.”
6.1.3. Ethical Resistance and Boundary-Setting Around AI Use
6.1.4. Conditional and Peripheral Roles for AI
6.1.5. Infrastructure, Systems, and Marking Efficiency
6.2. MDMF Version 5—Framework Updates Based on Mixed-Methods Evidence from the Post-Generative-AI Context (Late 2025)
6.2.1. Core Component 1 (Strengthened): Ethical Boundaries and Policy Clarity in the Age of AI
| Focus Area | MDMF v5 Update |
|---|---|
| Role of AI | Explicitly defines dissertation grading as a non-delegable academic judgment |
| Ethical stance | Shifts from regulating AI use to defining what AI must not do |
| Marker autonomy | Reinforces markers’ professional responsibility and intellectual ownership of assessment decisions |
| Policy emphasis | Moves from permissive guidance to boundary-setting and protection of human judgment |
| Empirical grounding | Informed by strong staff resistance to AI-based grading and high ethical concern scores |
6.2.2. Core Component 2 (Refined): Identifying Fairness and Equity Issues in Master’s Dissertation Marking
| Area | MDMF v5 Refinement—New Issue: Fairness and Equity Risks |
|---|---|
| AI-mediated writing | Introduces grade compression and fluency masking as fairness risks |
| Student contribution | Reframes obscured authorship as a systemic equity issue, not an individual misconduct problem |
| Language bias | Extends from linguistic disadvantage to AI-amplified fluency bias |
| Fairness logic | Moves away from AI as bias-mitigation toward human calibration and contextual judgment |
6.2.3. Core Component 3 (Operationalised): Pre-Marking Tasks and Calibration in the AI Era
| Aspect | MDMF v5 Update |
|---|---|
| Calibration | Repositioned as the primary human mechanism for managing AI-related uncertainty, including misalignment caused by AI-polished but conceptually weak work |
| Marking schemes | Emphasis on reducing ambiguity and interpretive drift rather than outsourcing interpretation to AI tools |
| Exemplars | Prioritise critical engagement and evidential depth over linguistic polish, including exemplars of AI-fluent but weak cases |
| AI awareness | Treated as a calibration and discussion topic for markers, not a detection or automation task |
6.2.4. Core Component 4 (Clarified): Assigning Marking Based on Marker Characteristics
| Dimension | MDMF v5 Refinement |
|---|---|
| Supervisor–reader asymmetry | Explicitly recognised as a fairness risk intensified by AI-polished writing |
| Reader insight | Greater emphasis on structured supervisor summaries and project-context sharing |
| Expertise matching | Reaffirmed as essential in managing topic diversity and AI-mediated surface quality |
| AI literacy | Defined as contextual understanding, not evaluative authority |
6.2.5. Core Component 5 (Expanded): Marking Process
| Element | MDMF v5 Enhancement |
|---|---|
| Marker well-being | Integrated as a structural condition, not an ancillary concern |
| Cognitive load | Recognised impact of repetitive AI-generated language on engagement and fatigue |
| Moderation | Emphasises discussion-based resolution over algorithmic or suspicion-driven approaches |
| Workload realism | Explicit acknowledgement that AI does not resolve time and volume pressures |
6.2.6. Core Component 6 (Concretised): Technology, Infrastructure, and Systems Design
| Area | MDMF v5 Update |
|---|---|
| System design | Focus on reducing friction rather than adding AI functionality |
| Submission infrastructure | Identified as a primary determinant of marking efficiency and fairness |
| AI tooling | Positioned as optional, peripheral support rather than a core marking technology |
| Data ethics | Reaffirms prohibition of training AI on student work without consent |
| Transparency | Infrastructure must support traceability of decisions, not automated judgment |
7. Discussion
7.1. Framework Effectiveness: How the MDMF Supports Fairer Dissertation Assessment
7.1.1. Enhancing Consistency Through Structured Processes and Calibration
7.1.2. Supporting Fairness Through Contextual and Relational Judgment
7.1.3. Improving Transparency and Accountability in Decision-Making
7.1.4. Addressing AI-Mediated Fairness Risks
7.1.5. Supporting Marker Well-Being and Sustainable Assessment Practices
7.1.6. Providing a Coherent, Adaptable Structure for Complex Assessment Contexts
7.2. Framework Contribution: Positioning the Master’s Dissertation Marking Framework (MDMF) Within Assessment Theory
7.2.1. Reasserting Professional Judgment in High-Stakes Assessment
7.2.2. From Procedural Fairness to Contextual Equity
7.2.3. Assessment as a Sociotechnical System
7.2.4. Emotional Labour and Marker Well-Being as Conditions of Fair Assessment
7.2.5. Boundary-Setting in the Post-Generative-AI Era
7.3. Implications for Practice
7.4. Limitations and Future Research
8. Conclusions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Afreen, J., Mohaghegh, M., & Doborjeh, M. (2025). Systematic literature review on bias mitigation in generative AI. AI Ethics, 5, 4789–4841. [Google Scholar] [CrossRef] [Scilit]
- An, Y., Yu, J. H., & James, S. (2025). Investigating the higher education institutions’ guidelines and policies regarding the use of generative AI in teaching, learning, research, and administration. International Journal of Educational Technology in Higher Education, 22(1), 1–23. [Google Scholar] [CrossRef] [Scilit]
- Armstrong, M., Dopp, C., & Welsh, J. (2018). Design-based research. In R. Kimmons (Ed.), The students’ guide to learning design and research. EdTech Books. Available online: https://edtechbooks.org/studentguide/design-based_research (accessed on 5 January 2024).
- Ashworth, M., Bloxham, S., & Pearce, L. (2010). Examining the tension between academic standards and inclusion for disabled students: The impact on marking of individual academics’ frameworks for assessment. Studies in Higher Education, 35(2), 209–223. [Google Scholar] [CrossRef] [Scilit]
- Bearman, M., & Ajjawi, R. (2018). From “seeing through” to “seeing with”: Assessment criteria and the myths of transparency. Frontiers in Education, 3, 96. [Google Scholar] [CrossRef] [Scilit]
- Bearman, M., Tai, J., Dawson, P., Boud, D., & Ajjawi, R. (2024). Developing evaluative judgement for a time of generative artificial intelligence. Assessment & Evaluation in Higher Education, 49(6), 893–905. [Google Scholar] [CrossRef] [Scilit]
- Bikanga Ada, M. (2014, July 7–9). A case study at a Scottish university. 6th International Conference on Education and New Learning Technologies, Barcelona, Spain. [Google Scholar]
- Bikanga Ada, M. (2018). Using design-based research to develop a Mobile Learning Framework for Assessment Feedback. Research and Practice in Technology Enhanced Learning, 13(3), 1–22. [Google Scholar] [CrossRef] [Scilit]
- Bikanga Ada, M. (2023, December 4–7). I betrayed my ethical principles: Investigating master’s dissertation marking practices in CS. 2022 IEEE International Conference on Teaching, Assessment and Learning for Engineering (TALE), Hung Hom, Hong Kong. [Google Scholar]
- Bloxham, S. (2009). Marking and moderation in the UK: False assumptions and wasted resources. Assessment & Evaluation in Higher Education, 34(2), 209–220. [Google Scholar] [CrossRef] [Scilit]
- Bloxham, S., & Boyd, P. (2012). Accountability in grading student work: Securing academic standards in a twenty-first century quality assurance context. British Educational Research Journal, 38(4), 615–634. [Google Scholar] [CrossRef] [Scilit]
- Bloxham, S., Boyd, P., & Orr, S. (2011). Mark my words: The role of assessment criteria in UK higher education grading practices. Studies in Higher Education, 36(6), 655–670. [Google Scholar] [CrossRef] [Scilit]
- Bloxham, S., Hughes, C., & Adie, L. (2016). What’s the point of moderation? A discussion of the purposes achieved through contemporary moderation practices. Assessment & Evaluation in Higher Education, 41(4), 638–653. [Google Scholar] [CrossRef] [Scilit]
- Boud, D., & Falchikov, N. (Eds.). (2007). Rethinking assessment in higher education: Learning for the longer term. Routledge/Taylor & Francis Group. [Google Scholar]
- Bourke, S., & Holbrook, A. P. (2013). Examining PhD and research masters theses. Assessment & Evaluation in Higher Education, 38(4), 407–416. [Google Scholar] [CrossRef] [Scilit]
- Braun, V., & Clarke, V. (2006). Using thematic analysis in psychology. Qualitative Research in Psychology, 3(2), 77–101. [Google Scholar] [CrossRef] [Scilit]
- Burger, R. (2017). Student perceptions of the fairness of grading procedures: A multilevel investigation of the role of the academic environment. Higher Education, 74, 301–320. [Google Scholar] [CrossRef] [Scilit]
- Carless, D., & Boud, D. (2018). The development of student feedback literacy: Enabling uptake of feedback. Assessment & Evaluation in Higher Education, 43(8), 1315–1325. [Google Scholar] [CrossRef] [Scilit]
- Chai, F., Ma, J., Wang, Y., Zhu, J., & Han, T. (2024). Grading by AI makes me feel fairer? How different evaluators affect college students’ perception of fairness. Frontiers in Psychology, 15, 1221177. [Google Scholar] [CrossRef] [Scilit]
- Chakraborty, S., Dann, C., Mandal, A., Dann, B., Paul, M., & Haeez-Baig, A. (2021). Effects of rubric quality on marker variation in higher education. Studies in Educational Evaluation, 70, 100997. [Google Scholar] [CrossRef] [Scilit]
- Corbin, T., Tai, J., & Flenady, G. (2025). Understanding the place and value of GenAI feedback: A recognition-based framework. Assessment & Evaluation in Higher Education, 50(5), 718–731. [Google Scholar] [CrossRef] [Scilit]
- Crisp, V. (2013). Criteria, comparison and past experiences: How do teachers make judgements when marking coursework? Assessment in Education: Principles, Policy & Practice, 20(1), 127–144. [Google Scholar] [CrossRef] [Scilit]
- Çağlar-Özhan, Ş., Tekeli, P., & Arkün-Kocadere, S. (2025). Comparison of AI-generated and instructor feedback: No significant difference in perceived feedback quality and neither on performance. Journal of Computer Assisted Learning, 41(5), e70134. [Google Scholar] [CrossRef] [Scilit]
- Davis, L. (2016). The influence of training and experience on rater performance in scoring spoken language. Language Testing, 33(1), 117–135. [Google Scholar] [CrossRef] [Scilit]
- Elliot, E., Pearce, K., & King, S. (2011, July 4–7). Moderation of assessment tasks: Developing solutions to common problems [Paper presentation]. HERDSA Conference 2011, Gold Coast, Australia. [Google Scholar]
- Ernstzen, D., Leibbrandt, D., & Louw, Q. (2026). The use and influence of supervisor marking in graduate research project assessment: A scoping review. Assessment & Evaluation in Higher Education, 1–20. [Google Scholar] [CrossRef] [Scilit]
- Festinger, L. (1957). A theory of cognitive dissonance. Stanford University Press. [Google Scholar]
- Fischer, J., Bearman, M., Boud, D., & Tai, J. (2024). How does assessment drive learning? A focus on students’ development of evaluative judgement. Assessment & Evaluation in Higher Education, 49(2), 233–245. [Google Scholar] [CrossRef] [Scilit]
- Forsyth, R., Cullen, R., Ringan, N., & Stubbs, M. (2015). Supporting the development of assessment literacy of staff through institutional process change. London Review of Education, 13(3), 34–41. Available online: https://files.eric.ed.gov/fulltext/EJ1160159.pdf (accessed on 25 November 2025). [CrossRef] [Scilit]
- Gipps, C. (1994). Beyond testing: Towards a theory of educational assessment (1st ed.). Routledge. [Google Scholar] [CrossRef] [Scilit]
- Gipps, C., & Stobart, G. (2009). Fairness in assessment. In C. Wyatt-Smith, & J. J. Cumming (Eds.), Educational assessment in the 21st century. Springer. [Google Scholar] [CrossRef] [Scilit]
- Gonsalves, C. (2024). Addressing student non-compliance in AI use declarations: Implications for academic integrity and assessment in higher education. Assessment & Evaluation in Higher Education, 50(4), 592–606. [Google Scholar] [CrossRef] [Scilit]
- Gonsalves, C., & Lin, Z. (2025). Clear in advance to whom? Exploring ‘transparency’ of assessment practices in UK higher education institution assessment policy. Studies in Higher Education, 50(7), 1454–1470. [Google Scholar] [CrossRef] [Scilit]
- Gordon, M. E., & Fay, C. H. (2010). The effects of grading and teaching practices on students’ perceptions of grading fairness. College Teaching, 58(3), 93–98. [Google Scholar] [CrossRef] [Scilit]
- Harper, G. (2016). Creative writing you’re cognitive dissonance. New Writing, 13(1), 1–2. [Google Scholar] [CrossRef] [Scilit]
- Hawe, E., Dixon, H., Murray, J., & Chandler, S. (2021). Using rubrics and exemplars to develop students’ evaluative and productive knowledge and skill. Journal of Further and Higher Education, 45(8), 1033–1047. [Google Scholar] [CrossRef] [Scilit]
- Hudson, J., Bloxham, S., den Outer, B., & Price, M. (2017). Conceptual acrobatics: Talking about assessment standards in the transparency era. Studies in Higher Education, 42(7), 1309–1323. [Google Scholar] [CrossRef] [Scilit]
- Jackson, D. (2018). Challenges and strategies for assessing student workplace performance during work-integrated learning. Assessment & Evaluation in Higher Education, 43(4), 555–570. [Google Scholar] [CrossRef] [Scilit]
- Jakesch, M., Hancock, J. T., & Naaman, M. (2023). Human heuristics for AI-generated language are flawed. Proceedings of the National Academy of Sciences of the United States of America, 120(11), e2208839120. [Google Scholar] [CrossRef] [Scilit]
- Jönsson, A., & Prions, F. (2019). Transparency in assessment—Exploring the influence of explicit assessment criteria. Frontiers in Education, 3, 119. [Google Scholar] [CrossRef] [Scilit]
- Kinman, G., & Wray, S. (2018). Presenteeism in academic employees-occupational and individual factors. Occupational Medicine, 68(1), 46–50. [Google Scholar] [CrossRef] [Scilit]
- Man, D., Xu, Y., Chau, M. H., O’Toole, J. M., & Shunmugam, K. (2020). Assessment feedback in examiner reports on master’s dissertations in translation studies. Studies in Educational Evaluation, 64, 100823. [Google Scholar] [CrossRef] [Scilit]
- Marr, L., & Forsyth, R. (2010). Identity crisis: Working in higher education in the 21st century. Trentham Books. [Google Scholar]
- McKenney, S. E., & Reeves, T. C. (2012). Conducting educational design research. Routledge. [Google Scholar]
- McQuade, R., Kometa, S., Brown, J., Bevitt, D., & Hall, J. (2020). Research project assessments and supervisor marking: Maintaining academic rigour through robust reconciliation processes. Assessment & Evaluation in Higher Education, 45(8), 1181–1191. [Google Scholar] [CrossRef] [Scilit]
- Mei, P., Brewis, D. N., Nwaiwu, F., Sumanathilaka, D., Alva-Manchego, F., & Demaree-Cotton, J. (2025). If ChatGPT can do it, where is my creativity? Generative AI boosts performance but diminishes experience in creative writing. Computers in Human Behavior: Artificial Humans, 4, 100140. [Google Scholar] [CrossRef] [Scilit]
- Mittelstadt, B. (2019). Principles alone cannot guarantee ethical AI. Nature Machine Intelligence, 1(11), 501–507. [Google Scholar] [CrossRef] [Scilit]
- Nazaretsky, T., Mejia-Domenzain, P., Swamy, V., Frej, J., & Käser, T. (2026). Who gives feedback matters: Student biases towards human and AI-generated formative feedback. Journal of Computer Assisted Learning, 42(1), e70153. [Google Scholar] [CrossRef] [Scilit]
- O’donovan, B., Price, M., & Rust, C. (2004). Know what I mean? Enhancing student understanding of assessment standards and criteria. Teaching in Higher Education, 9(3), 325–335. [Google Scholar] [CrossRef] [Scilit]
- Peeters, M. J., Schmude, K. A., & Steinmiller, C. L. (2014). Inter-rater reliability and false confidence in precision: Using standard error of measurement within PharmD admissions essay rubric development. Currents in Pharmacy Teaching and Learning, 6(2), 298–303. [Google Scholar] [CrossRef] [Scilit]
- Perkins, M., & Roe, J. (2024). Decoding academic integrity policies: A corpus linguistics investigation of AI and other technological threats. Higher Education Policy, 37, 633–653. [Google Scholar] [CrossRef] [Scilit]
- Petricini, T., Zipf, S., & Wu, C. (2025). RESEARCH-AI: Communicating academic honesty: Teacher messages and student perceptions about generative AI. Frontiers in Communication, 10, 1544430. [Google Scholar] [CrossRef] [Scilit]
- Pérez-Ros, P., Chust-Hernández, P., Ibáñez-Gascó, J., & Martínez-Arnau, F. M. (2021). An undergraduate thesis training course for faculty reduces variability in student evaluations. Nurse Education Today, 96, 104619. [Google Scholar] [CrossRef] [Scilit]
- Pilcher, N. (2011). The UK postgraduate masters dissertation: An ‘elusive chameleon’? Teaching in Higher Education, 16(1), 29–40. [Google Scholar] [CrossRef] [Scilit]
- Pitt, E., & Winstone, N. (2018). The impact of anonymous marking on students’ perceptions of fairness, feedback and relationships with lecturers. Assessment & Evaluation in Higher Education, 43(7), 1183–1193. [Google Scholar] [CrossRef] [Scilit]
- Postmes, L., Bouwmeester, R., de Kleijn, R., & van der Schaaf, M. (2022). Supervisors’ untrained postgraduate rubric use for formative and summative purposes. Assessment & Evaluation in Higher Education, 48(1), 41–55. [Google Scholar] [CrossRef] [Scilit]
- Pufpaff, L. A., Clarke, L., & Jones, R. E. (2015). The effects of rater training on inter-rater agreement. Mid-Western Educational Researcher, 27(2), 117–141. [Google Scholar]
- Quality Assurance Agency. (2006). Code of practice for the assurance of academic quality and standards in higher education, Section 6: Assessment of students. QAA. Available online: https://dera.ioe.ac.uk/9713/2/COP_AOS.pdf (accessed on 23 May 2025).
- Quality Assurance Agency. (2018). UK quality code for higher education: Advice and guidance—Assessment. Available online: https://www.qaa.ac.uk/the-quality-code/2018/advice-and-guidance-18/assessment (accessed on 23 January 2026).
- Sadler, D. (2014). The futility of attempting to codify academic achievement standards. Higher Education, 67(3), 273–288. [Google Scholar] [CrossRef] [Scilit]
- Sadler, D. R. (1989). Formative assessment and the design of instructional systems. Instructional Science, 18, 119–144. [Google Scholar] [CrossRef] [Scilit]
- Sadler, D. R. (2009). Indeterminacy in the use of preset criteria for assessment and grading. Assessment & Evaluation in Higher Education, 34(2), 159–179. [Google Scholar] [CrossRef] [Scilit]
- Satchell, S., & Pratt, J. (2010). The dubiety of double marking. Higher Education Review, 42(2), 59–62. Available online: https://eric.ed.gov/?id=EJ877250 (accessed on 10 January 2025).
- Tai, J., Ajjawi, R., Boud, D., Dawson, P., & Panadero, E. (2018). Developing evaluative judgement: Enabling students to make decisions about the quality of work. Higher Education, 76, 467–481. [Google Scholar] [CrossRef] [Scilit]
- Tisi, J., Whitehouse, G., Maughan, S., & Burdett, N. (2013). A review of literature on marking reliability research (report for ofqual). NFER. Available online: https://assets.publishing.service.gov.uk/media/5a81cc0e40f0b62302699360/0613_JoTisi_et_al-nfer-a-review-of-literature-on-marking-reliability.pdf (accessed on 15 September 2025).
- Vinton, L., & Wilke, D. (2011). Leniency bias in evaluating clinical social work student interns. Clinical Social Work Journal, 39(3), 288–295. [Google Scholar] [CrossRef] [Scilit]
- Williams, L., & Kemp, S. (2019). Independent markers of master’s theses show low levels of agreement. Assessment & Evaluation in Higher Education, 44(5), 764–771. [Google Scholar] [CrossRef] [Scilit]
- Williamson, B., & Eynon, R. (2020). Historical threads, missing links, and future directions in AI in education. Learning, Media and Technology, 45(3), 223–235. [Google Scholar] [CrossRef] [Scilit]
- Winstone, N. E., & Boud, D. (2022). The need to disentangle assessment and feedback in higher education. Studies in Higher Education, 47(3), 656–667. [Google Scholar] [CrossRef] [Scilit]
- Wolf, K. (2015). Leniency and halo bias in industry-based assessments of student competencies: A critical, sector-based analysis. Higher Education Research & Development, 34(5), 1045–1059. [Google Scholar] [CrossRef] [Scilit]
- Zhan, Y., Boud, D., Dawson, P., & Yan, Z. (2025). Generative artificial intelligence as an enabler of student feedback engagement: A framework. Higher Education Research & Development, 44(5), 1289–1304. [Google Scholar] [CrossRef] [Scilit]





| Phase | Year | Purpose | Method | Participants | Analysis | Framework Output |
|---|---|---|---|---|---|---|
| Phase 1 | 2022 | Identify key issues in dissertation marking | Literature review | — | Narrative synthesis | MDMF v1 |
| Phase 2 | 2022 | Explore staff experiences and challenges | Survey | n = 31 academic staff | Descriptive + thematic analysis | MDMF v2 |
| Phase 3 | 2022 | Deepen understanding of marking practices | Semi-structured interviews | n = 7 academic staff | Thematic analysis | MDMF v3 |
| Phase 4 | 2023 | Capture early responses to generative AI | Two exploratory surveys | n = 19; n = 8 academic staff | Descriptive analysis | Exploratory insights informing AI-related framework adaptation |
| Phase 5 | Early 2025 | Examine post-AI marking practices | Semi-structured interviews | n = 3 academic staff | Reflexive thematic analysis | MDMF v4 |
| Phase 6 | Late 2025 | Evaluate AI-related challenges and perceptions | Mixed-methods survey | n = 13 (quantitative); n = 14 (qualitative) | Descriptive statistics + qualitative analysis | MDMF v5 |
| Item | n | Median | Mean | SD | |
|---|---|---|---|---|---|
| Project complexity | 13 | 1.00 | 1.92 | 1.605 | |
| Bias1 | Supervisor’s bias | 13 | 1.00 | 2.00 | 1.581 |
| Bias2 | Reader’s bias | 12 | 1.00 | 1.83 | 1.528 |
| Bias3 | Markers’ intercultural differences | 13 | 1.00 | 2.31 | 1.932 |
| Bias4 | Language bias | 13 | 1.00 | 2.00 | 1.581 |
| Wkld1 | Comparing dissertations | 13 | 1.00 | 2.46 | 1.984 |
| Wkld2 | Disparity in project requirements | 13 | 1.00 | 1.85 | 1.281 |
| Wkld3 | Lack of subject knowledge | 13 | 4.00 | 3.23 | 2.006 |
| Wkld4 | High marking load | 13 | 3.00 | 3.31 | 2.463 |
| Wkld5 | Time spent on marking | 13 | 3.00 | 3.38 | 2.293 |
| Wkld6 | Emotional and mental issues (stress, resentment, tiredness, confusion, disappointment) | 13 | 1.00 | 1.85 | 1.463 |
| Wkld7 | Vagueness or irrelevance of marking schemes | 13 | 1.00 | 1.62 | 1.193 |
| Wkld8 | Long marking schemes | 13 | 1.00 | 1.38 | 0.870 |
| Wkld9 | Diversity of project topics/areas | 13 | 1.00 | 2.15 | 1.519 |
| Wkld10 | Lack of anonymity | 13 | 1.00 | 1.46 | 1.391 |
| Wkld11 | Not knowing the student contribution | 13 | 1.00 | 1.31 | 0.751 |
| Wkld12 | Dissertation marking time clashes with other marking | 13 | 1.00 | 2.38 | 1.710 |
| conc1 | I am concerned about AI tools’ limited ability to evaluate originality in a student’s dissertation. | 11 | 7.00 | 6.45 | 0.688 |
| Conc1 | AI models might have difficulties aligning their grading with established human rubrics. | 11 | 7.00 | 5.09 | 2.386 |
| Conc3 | I am concerned about potential biases in AI when grading dissertations. | 11 | 7.00 | 6.09 | 1.514 |
| Conc4 | There is a risk of becoming overly dependent on AI for tasks like dissertation marking. | 11 | 7.00 | 6.64 | 0.674 |
| Conc5 | I believe the lack of personal touch in feedback from AI tools like ChatGPT is a limitation. | 11 | 7.00 | 5.09 | 2.508 |
| Conc6 | I am concerned about potential ethical issues, such as student data privacy, when using ChatGPT in the dissertation marking process. | 11 | 7.00 | 5.64 | 1.859 |
| Conc7 | The lack of context understanding in AI is a limitation for using ChatGPT in the dissertation marking process. | 11 | 7.00 | 5.36 | 2.378 |
| Conc8 | ChatGPT’s lack of understanding of context hampers its effectiveness in marking dissertations. | 11 | 5.00 | 4.45 | 2.734 |
| I have encountered other concerns/limitations, or I foresee other concerns/limitations. (Individual item) | 10 | 3.50 | 4.00 | 2.749 | |
| ChatGPT and similar technologies should replace current practices in the dissertation marking process going forward. (Individual item) | 11 | 1.00 | 1.27 | 0.647 | |
| Sup1 | I could be satisfied with the quality of support and dissertation feedback provided by ChatGPT. | 10 | 1.50 | 2.50 | 1.900 |
| Sup2 | ChatGPT and similar technologies should complement human markers in the dissertation marking process. | 11 | 3.00 | 2.73 | 1.849 |
| Sup3 | Using ChatGPT can make the dissertation marking process more fair and objective. | 11 | 1.00 | 2.09 | 1.700 |
| n | Mean | SD | Cronbach’s Alpha | |
|---|---|---|---|---|
| Effectiveness of AI tools in addressing bias issues identified in the dissertation marking process | 13 | 2.08 | 1.47 | 0.885 (items = 4) |
| Effectiveness of AI tools in addressing the workload/process issues identified in the dissertation marking process | 13 | 2.20 | 0.94 | 0.810 (items = 12) |
| Concerns or limitations encountered or foreseen when using ChatGPT in the dissertation marking process | 11 | 5.60 | 1.41 | 0.858 (items = 8) |
| Support encountered or foreseen when using ChatGPT in the dissertation marking process | 10 | 2.48 | 1.41 | 0.671 (items = 3) |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the author. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Bikanga Ada, M. Fair Marking in the Generative AI Era: Introducing the Master’s Dissertation Marking Framework. AI Educ. 2026, 2, 23. https://doi.org/10.3390/aieduc2030023
Bikanga Ada M. Fair Marking in the Generative AI Era: Introducing the Master’s Dissertation Marking Framework. AI in Education. 2026; 2(3):23. https://doi.org/10.3390/aieduc2030023
Chicago/Turabian StyleBikanga Ada, Mireilla. 2026. "Fair Marking in the Generative AI Era: Introducing the Master’s Dissertation Marking Framework" AI in Education 2, no. 3: 23. https://doi.org/10.3390/aieduc2030023
APA StyleBikanga Ada, M. (2026). Fair Marking in the Generative AI Era: Introducing the Master’s Dissertation Marking Framework. AI in Education, 2(3), 23. https://doi.org/10.3390/aieduc2030023
