Next Article in Journal
Digital Entrepreneurial Ecosystem Maturity: Developing an Integrated Assessment Framework for Cyprus and Small Economies
Previous Article in Journal
Characteristics of the Rural Cultural Landscape in Gökçeada (Imbros)
Previous Article in Special Issue
Doctoral Student Wellbeing: Conceptualization, Challenges and Pathways Forward
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

Generative Artificial Intelligence in Doctoral Supervision: Ethical Boundaries, Feedback Practices, and Supervisory Relationships

by
Carla Giuliana Guanilo Pareja
1,
Lidia Ysabel Pareja Pera
2,
Andrés Arias Lizares
3,
Lupe Marilu Huanca Rojas
4,
Henri Emmanuel López Gómez
5,*,
Zoila Rosa Díaz Tavera
6,
Clara Patricia Almonte Andrade
6,
Fernando Martin Ramirez Wong
7,
Wanda Marina Román-Santana
8,
Carmen Mata-de-Salcedo
8 and
Judith Martínez-Alonzo
9
1
Facultad de Gestión Empresarial, Escuela Profesional de Administración de Negocios Internacionales, Universidad Femenina del Sagrado Corazón, Lima 15023, Peru
2
Facultad de Gestión Empresarial, Escuela Profesional de Contabilidad y Finanzas, Universidad Femenina del Sagrado Corazón, Lima 15023, Peru
3
Facultad de Ciencias de la Educación, Escuela Profesional de Educación Secundaria, Universidad Nacional del Altiplano, Puno 21001, Peru
4
Facultad de Educación, Escuela Profesional de Educación Intercultural Bilingüe: Nivel Inicial y Nivel Primaria, Universidad Nacional Intercultural de la Selva Central Juan Santos Atahualpa, Mazamari 12301, Peru
5
Campus Huancayo, Universidad Tecnológica del Perú, Huancayo 12002, Peru
6
Escuela Profesional de Enfermería, Facultad de Ciencias de la Salud, Universidad Nacional del Callao, Callao 07011, Peru
7
Departamento Académico de Ciencias Dinámicas, Facultad de Medicina, Universidad Nacional Mayor de San Marcos, Lima 15001, Peru
8
Instituto Superior de Formación Docente Salomé Ureña, Recinto Luis Napoleón Núñez Molina, Carretera Duarte, km 10 ½, Municipio de Licey al Medio 51000, Dominican Republic
9
Recinto San Francisco de Macorís, Universidad Autónoma de Santo Domingo, Avenida Manuel Aurelio Tavárez Justo, Salida Nagua, San Francisco de Macorís 31000, Dominican Republic
*
Author to whom correspondence should be addressed.
Encyclopedia 2026, 6(10), 209; https://doi.org/10.3390/encyclopedia6100209
Submission received: 8 August 2026 / Revised: 31 August 2026 / Accepted: 15 September 2026 / Published: 24 September 2026
(This article belongs to the Collection Doctoral Supervision)

Abstract

Generative artificial intelligence is increasingly being adopted in doctoral education to support writing, literature synthesis, translation, feedback, research planning, and administrative work. While emerging scholarship has examined generative AI in higher education, academic integrity, and student writing, its specific implications for doctoral supervision remain comparatively underexplored. This review argues that generative AI should not be understood merely as a technical aid, but as a socio-technical force that may reshape ethical boundaries, feedback practices, doctoral student agency, and supervisory relationships. Drawing on literature from doctoral education, supervision studies, academic writing, feedback literacy, and AI-mediated learning, the review examines three interrelated thematic areas of AI-augmented doctoral supervision. First, it considers ethical issues related to authorship, originality, academic integrity, transparency, privacy, confidentiality, bias, data protection, and overdependence. Second, it analyses how generative AI may influence feedback preparation, interpretation, timing, and dialogic use in doctoral writing and research development. Third, it explores implications for supervisory relationships, particularly trust, power, responsibility, emotional support, and autonomy. The review also discusses institutional implications, including AI literacy, supervisor development, doctoral policy, and responsible use guidelines, and positions generative AI as a human-centred support mechanism rather than a substitute for supervisory judgement, ethical responsibility, or relational mentoring.

1. Introduction

Doctoral supervision is a relational and developmental academic practice through which researchers learn to produce original knowledge, exercise methodological judgement, communicate within a discipline, and form a scholarly identity. It extends beyond monitoring progress or correcting drafts: supervisors provide intellectual direction, ethical guidance, critical feedback, emotional support, and professional socialisation [1,2,3,4,5,6]. Candidates, in turn, are expected to move from supported participation towards increasing independence, develop an academic voice, and defend an original contribution [7,8]. The central supervisory challenge is therefore to combine guidance with autonomy. Effective support should make the candidate more capable of reading, reasoning, deciding, and writing independently rather than becoming a continuing substitute for those activities.
Generative AI introduces new possibilities into this already complex relationship. Large language model-based systems can produce, revise, translate, summarise, classify, and reorganise text in response to prompts. In doctoral work, they may support literature searches, conceptual clarification, research planning, academic writing, coding, meeting preparation, translation, editing, administrative communication, and preliminary feedback [9,10,11]. For candidates, these systems can provide immediate assistance between supervision meetings and may reduce linguistic or organisational barriers. For supervisors, they may help prepare routine materials, formulate examples, or structure developmental activities. These uses make GenAI potentially valuable, particularly where supervisory time, language support, or access to informal academic assistance is limited.
Their significance, however, extends beyond efficiency. Reading difficult literature, developing an argument, making methodological choices, interpreting criticism, and revising a scholarly text are not merely tasks to be completed; they are formative processes through which doctoral researchers acquire expertise and an independent voice [12,13]. GenAI can intervene directly in each of these processes. Used reflectively, it may help candidates compare alternatives, identify uncertainty, or articulate a question. Used substitutively, it may bypass the intellectual struggle through which judgement is developed. The same capabilities therefore generate questions about academic integrity, authorship, originality, transparency, privacy, confidentiality, bias, epistemic authority, and the boundary between legitimate assistance and outsourced intellectual labour [14,15].
Much of the higher education debate has concentrated on assessment, misconduct, AI-generated student text, and institutional policy [16,17]. Doctoral education presents a distinct context. It is organised around original research, prolonged engagement with uncertainty, discipline-specific standards, and a sustained relationship in which the supervisor both supports development and contributes to judgements about progress and quality. GenAI may therefore alter more than the production of text. It can affect how knowledge claims are formulated, how feedback is interpreted, how candidate competence is inferred, and how authority and responsibility are distributed between candidates, supervisors, and institutions. These relational consequences require analysis that cannot be reduced to plagiarism detection or productivity gains.
AI use may also remain hidden or ambiguously disclosed. Candidates may use GenAI to interpret supervisory comments, revise arguments, generate literature summaries, rehearse explanations, or seek reassurance without discussing that assistance. In this sense, the system can operate as an unofficial third actor in the supervisory process [12,18]. Such use can be supportive, but opacity creates practical and ethical problems. Supervisors may be unable to determine whether a polished passage reflects the candidate’s current understanding, whether an argument has been substantially machine-shaped, or whether confidential data have entered an external platform. Candidates may likewise be uncertain about what is permitted, what should be disclosed, and where support becomes inappropriate substitution.
Feedback makes these tensions particularly visible. Doctoral feedback is dialogic and developmental: it helps candidates interpret disciplinary standards, refine arguments, evaluate evidence, and become increasingly self-regulating scholars [19,20]. GenAI can provide rapid comments on clarity, structure, coherence, or style and can help candidates formulate questions before a meeting. Yet generated feedback may lack disciplinary nuance, methodological context, ethical sensitivity, and knowledge of the candidate’s developmental trajectory. It can also appear authoritative despite being generic or wrong. The relevant issue is therefore not whether machine feedback is available, but how it interacts with supervisor feedback, feedback literacy, and the gradual development of independent evaluative judgement.
The supervisory relationship likewise depends on negotiated expectations, intellectual challenge, emotional attunement, and shared responsibility [12,18]. If GenAI is used to avoid difficult conversations, generate impersonal comments, or replace mentoring, it may weaken trust and relational accountability. Conversely, transparent use may improve preparation, enable candidates to articulate uncertainty, and make meetings more focused. AI can support a conversation, but it cannot assume responsibility for the quality of guidance, the wellbeing of the candidate, or the ethical consequences of a research decision. The central question is therefore how GenAI can be integrated without displacing human judgement or the formative role of the supervisor–candidate relationship.
Recent syntheses address important parts of this landscape. Reviews of AI and academic integrity examine broad higher education risks and policy responses [14,15]; a systematic review of generative-chat feedback focuses on academic writing outcomes [21]; and a systematic review of AI-powered tools maps applications in doctoral supervision [22]. These studies provide essential foundations, but their primary units of analysis are usually policy problems, feedback interventions, or technological tools. They do not fully integrate ethical accountability, feedback literacy, doctoral agency, supervisory trust and power, and institutional governance within one relational system. This leaves a need for a synthesis that treats GenAI as a mediating influence across the whole supervisory ecology rather than as an isolated application.
Table 1 positions the present review in relation to these adjacent syntheses. It shows that the distinctive contribution is not a catalogue of tools, but an integrated account of how AI-mediated practices affect the responsibilities and relationships through which doctoral researchers are formed.
This review therefore shifts the unit of analysis from individual tools or isolated tasks to the socio-technical supervisory system comprising doctoral researchers, supervisors, institutions, and GenAI systems. It develops a human-centred framework connecting five dimensions: ethical accountability; feedback and learning; agency and autonomy; trust and supervisory relationships; and institutional governance. As a conceptual and integrative framework, it provides an organising structure for comparing future empirical findings and for translating broad principles into operational supervisory practices. Figure 1 represents these relationships.
Human judgement and scholarly responsibility form its governing core. GenAI may mediate writing, planning, feedback, and communication, but it cannot hold academic authorship, moral accountability, or supervisory responsibility. The four human and technological actors are therefore connected without being treated as equivalent: responsibility remains with candidates, supervisors, and institutions.
The framework is dynamic rather than a fixed configuration of actors and responsibilities. As doctoral work progresses, the relative intensity and form of interaction among the doctoral researcher, supervisor or supervisory team, institution, and GenAI system may change as researcher autonomy develops, project requirements evolve, and new methodological, data, publication, or technological conditions emerge. Responsible AI-augmented supervision therefore requires expectations and practices to be revisited at relevant milestones rather than established once at candidature entry. Human judgement and scholarly responsibility remain the governing core throughout this progression, while ethical accountability, feedback and learning, agency and autonomy, trust and supervisory relationships, and institutional governance operate as interdependent dimensions through which AI-related practices are evaluated and adjusted.
The five dimensions are analytically distinct but practically interdependent. Ethical accountability influences what information may be shared and who is responsible for a claim; feedback and learning concern how generated material is evaluated and used; agency and autonomy address whether candidates retain intellectual control; trust and supervisory relationships determine whether use is discussable and development can be judged; and institutional governance supplies the policy, infrastructure, and training that make responsible practice possible. A weakness in one dimension can undermine the others. For example, useful feedback generated from confidential data is not ethically defensible, and transparent use that nevertheless replaces candidate judgement does not meet the developmental purpose of doctoral education.
These relationships organise the review around four questions: (RQ1) How does GenAI reshape ethical accountability, authorship, transparency, and responsibility in doctoral supervision? (RQ2) How may GenAI influence feedback practices, doctoral learning, agency, autonomy, and independent scholarly judgement? (RQ3) How may GenAI reshape trust, power, and responsibility within doctoral supervisory relationships? (RQ4) What institutional governance, AI-literacy, data-protection, and support conditions are needed for responsible and human-centred GenAI use in doctoral supervision?
The article proceeds from the methodological approach and conceptual foundations to current uses and opportunities. It then examines ethical boundaries and research integrity, AI-mediated feedback and doctoral learning, trust and power in supervisory relationships, and institutional governance. Operational recommendations convert the synthesis into observable responsibilities for the principal stakeholders. The final sections identify research gaps, acknowledge the limitations of the review, and state the conditions under which GenAI may support rather than weaken doctoral formation.

2. Methodological Approach

This article uses a critical and integrative narrative review to connect heterogeneous empirical, conceptual, review, methodological, and policy literature. The term is used here as a methodological description drawing on established integrative review and critical review traditions rather than as a claim to a single standardised review protocol. Integrative reviews permit the synthesis of methodologically and conceptually diverse forms of evidence to develop a broader understanding of complex phenomena [23,24], whereas critical reviews emphasise analysis, interpretation, conceptual development, and critical engagement with the literature rather than description alone [25]. In the present review, the integrative component refers to the combination of heterogeneous evidence types; the critical component refers to the comparative interpretation of their assumptions, convergence, tensions, and evidential limits; and the narrative component refers to the structured qualitative synthesis through which these relationships are developed. Integrative review designs have also been applied in doctoral education research where heterogeneous evidence must be brought together to analyse complex developmental phenomena [26].
This approach is appropriate because GenAI in doctoral supervision is not a single intervention associated with one stable or directly comparable outcome. It spans doctoral pedagogy, academic writing, feedback, research ethics, data governance, candidate agency, supervisory relationships, and institutional policy. The review therefore seeks conceptual integration rather than a pooled estimate of effect. It does not report new empirical data; instead, it develops a conceptually structured synthesis of convergent findings, disagreements, mechanisms, responsibilities, and evidence gaps within a defined thematic scope.
The methodological purpose of this review differs from that of a systematic review designed to answer a predefined question through explicit procedures for identifying, selecting, appraising, and synthesising eligible evidence. Although the present review also uses structured searches, eligibility criteria, evidence classification, and a documented synthesis process, these procedures support the critical integration of heterogeneous empirical, conceptual, methodological, review, and normative sources rather than the synthesis of a more narrowly defined body of evidence addressing a common review question. PRISMA is therefore used here as a point of methodological comparison rather than as the reporting framework for the present review. PRISMA is a reporting guideline for systematic reviews and is not itself a review methodology [27]. Table 2 summarises the methodological distinction between the present critical and integrative narrative review and a systematic review reported using PRISMA.
Literature was identified primarily through Scopus and Web of Science searches conducted between 1 and 11 July 2026. The database searches prioritised publications from January 2019 to 11 July 2026 as a pragmatic recency boundary for capturing the contemporary literature relevant to GenAI, AI-mediated academic work, and emerging higher-education applications. The 2019 lower boundary was used as a priority filter rather than as an absolute historical cut-off or a claim about the origin of generative AI. Earlier publications were retained when they were necessary to establish foundational concepts in doctoral supervision, feedback, academic writing, researcher identity, agency, or research ethics. This two-level temporal strategy allowed the review to privilege the rapidly developing contemporary evidence base while preserving older scholarship required to interpret the relational, pedagogical, and ethical dimensions of doctoral supervision. Targeted checks of publisher pages and recognised institutional or international guidance were additionally used to verify references and include authoritative normative sources.
Twelve database-specific search sets were organised around four thematic areas: GenAI in doctoral education and supervision; GenAI in doctoral and academic writing; AI-mediated feedback, ethics, academic integrity, and governance; and supervisory relationships involving trust, power, agency, autonomy, identity, wellbeing, and emotional support. The six Scopus search sets generated 372 query-level hits, and the six Web of Science search sets generated 2215 query-level hits. The combined figure of 2587 represents non-unique search-set hits because thematic searches and databases overlapped substantially.
Eligibility was assessed against predefined topical, publication, temporal, educational, AI-related, ethical, and conceptual criteria. These criteria were designed to preserve direct relevance to doctoral supervision while permitting inclusion of adjacent scholarship when it contributed necessary empirical evidence, conceptual foundations, methodological insight, or authoritative normative guidance. Sources were therefore not included solely because they mentioned artificial intelligence or higher education; they had to contribute substantively to at least one analytical dimension of the review. The complete inclusion and exclusion criteria, together with the rationale for each criterion, are reported in Supplementary Table S2.
Source identification, organisation, relevance assessment, and synthesis involved contributions from the author team. Authors responsible for investigation and data curation supported the organisation and checking of potentially relevant sources, while methodology and formal-analysis contributors participated in reviewing source relevance, thematic interpretation, and the development of the conceptual synthesis. Final inclusion and interpretive decisions were discussed within the author team. Potentially relevant material was evaluated for direct relevance to doctoral supervision or for a clearly specified conceptual contribution from an adjacent field. Peer-reviewed empirical studies, conceptual articles, reviews, methodological papers, and authoritative guidance were eligible; unsupported commentary, unrelated technical developments, and studies without educational, supervisory, ethical, or research relevance were excluded. The final analytical corpus contains 61 sources. Five additional methodological and review design references [23,24,25,26,27] were used to define, exemplify, and position the review design and were treated separately from the analytical corpus. A retrospective audit checked the 61-source analytical corpus for duplicate records by DOI, title, and author–year and classified every source by its evidential role in the review.
Across the overlapping search sets, source selection was guided by relevance to the thematic and analytical scope of the review. Sources were retained when they addressed doctoral supervision directly or contributed empirical evidence, conceptual or methodological insight, or authoritative normative guidance to one or more of the five analytical dimensions: ethical accountability; feedback and learning; agency and autonomy; trust and supervisory relationships; and institutional governance. Priority was given to literature on GenAI in doctoral or postgraduate research, AI-mediated feedback and academic writing, academic integrity, authorship, disclosure, privacy, data protection, and responsible AI use. Sources focused only on generic educational applications of AI, technical system development without educational or supervisory relevance, or topics that did not contribute to the conceptual, ethical, pedagogical, or institutional argument were not retained.
The 2587 retrievals therefore represent overlapping query-level search results rather than a unique set of screened records. The final analytical corpus comprised 61 sources. The original review workflow did not retain stage-specific unique counts following cross-search deduplication, title and abstract assessment, full-text assessment, or exclusions by reason; these quantities are therefore not reported retrospectively. To make the documented process visually traceable without reconstructing unavailable screening counts, Supplementary Figure S1 presents the sequence from scope definition and database searching through eligibility assessment, structured extraction, evidence classification, coding, and final synthesis. The documented search strategies, eligibility criteria, source classifications, and thematic mapping provide the traceable basis for the narrative synthesis.
A structured extraction framework was applied to the included sources. Extracted fields comprised author and year; country, institutional, or disciplinary context when reported; publication and evidence type; study design; participants or documentary corpus; the GenAI or supervision practice examined; principal finding or proposition; analytical domain; and the source’s contribution to the synthesis.
The synthesis combined deductive and inductive coding. Deductive coding was anchored in three broad thematic domains established in the review scope—ethical boundaries, feedback practices, and supervisory relationships—while inductive coding captured recurring themes including transparency, confidentiality, trust, power, equity, bias, overdependence, scholarly identity, human accountability, and institutional governance. Through iterative comparison across candidate, supervisor, and institutional levels, these deductive and emergent themes were organised into the five-dimensional human-centred framework. The four review questions articulate this integrated analytical structure across the ethical, developmental, relational, and governance issues examined in the review.
Evidence was not treated as interchangeable. Instead, each source type was assigned a distinct evidential role, and the wording of claims was calibrated accordingly (Table 3).
These distinctions were used throughout the manuscript to preserve the provenance and appropriate strength of each claim. No pooled effect estimate or universal causal claim is made. Practice-oriented statements developed through the synthesis are therefore presented as interpretive recommendations rather than as empirically demonstrated effects.
Given the heterogeneous nature of the corpus, no single formal risk-of-bias instrument was applicable across all evidence types. Interpretive weight was therefore informed by direct relevance to doctoral supervision, transparency of design or argument, contextual specificity, methodological coherence, recency, and consistency with related evidence. These criteria were used to calibrate the strength of the review’s interpretations rather than to generate a numerical quality score.
The review was developed with reference to the Scale for the Assessment of Narrative Review Articles (SANRA) [28], a six-item instrument designed to support the quality assessment of narrative review articles. SANRA addresses six review-level domains: justification of the article’s importance, statement of concrete aims or review questions, description of the literature search, referencing, scientific reasoning, and appropriate presentation of evidence. Each domain is rated from 0 to 2, yielding a maximum total score of 12.
SANRA was selected because the present study is a critical and integrative narrative review rather than a systematic review or meta-analysis, and its domains correspond directly to the transparency, reasoning, and reporting requirements relevant to this review design. It was used as a review-level quality-alignment and reporting aid, not as a risk-of-bias instrument for the individual sources in the corpus. This distinction is important because the review integrates empirical studies, conceptual and design-science scholarship, reviews, methodological sources, and normative guidance that cannot appropriately be reduced to a single study-level quality metric. Complete search strategies, retrieval counts, eligibility criteria, thematic mappings, and the SANRA-informed alignment checklist are provided in Supplementary Material S1.

3. Conceptual Foundations of Doctoral Supervision

Doctoral supervision is best understood as a developmental relationship rather than a sequence of administrative transactions. Through repeated discussion, critique, revision, and participation in scholarly practice, candidates are inducted into disciplinary communities, ethical norms, research methods, and forms of academic communication [2,5]. Supervisors help refine questions, interrogate assumptions, consider methodological alternatives, interpret literature, and respond to criticism. The quality of this process depends not only on subject expertise but also on the clarity of expectations, the timing and form of feedback, and the ways in which authority and responsibility are negotiated over time. This relational conception is essential for evaluating GenAI because the technology enters practices that already have pedagogical, ethical, and interpersonal meanings.
A defining tension is the balance between support and autonomy. Supervisors must guide without dominating, challenge without discouraging, and advise without replacing the candidate’s intellectual labour [1,3]. Candidates need substantial support at some stages, but the direction of development should be towards stronger independent judgement. The doctorate certifies not simply the production of a polished document, but the capacity to formulate and defend an original contribution. Assistance becomes problematic when it obscures whose reasoning, interpretation, or methodological judgement produced the work. This distinction cannot be resolved by the amount of text generated alone; it depends on the function of the assistance and on whether the candidate remains able to explain and defend the relevant decisions.
GenAI complicates this balance because it can supply immediate and plausible responses to tasks that candidates are expected to master through sustained practice. It can summarise literature, propose research questions, revise prose, translate text, suggest analytical categories, interpret feedback, or simulate a critical reader [11,29]. These functions may scaffold reflection and give candidates material to evaluate. They may also short-circuit the processes through which expertise develops. GenAI can therefore be treated as a consequential quasi-participant: not because it has academic authority or moral agency, but because its outputs can shape how candidates formulate ideas, understand criticism, and present their contribution. Its influence must be visible enough for supervisors to evaluate development and for candidates to remain accountable.
This influence changes the conditions of trust. Supervisors need credible evidence that submitted work reflects the candidate’s understanding, while candidates need fair and informed guidance about acceptable assistance [10,30]. Ambiguity can lead candidates to conceal legitimate use for fear of accusation or to use AI uncritically because no boundary has been established. Detection-led responses are a poor substitute for dialogue because current detectors are technically unreliable and do not reveal the quality of the candidate’s reasoning [14,15]. Clear expectations, proportional disclosure, and discussion of the decision process are more educationally useful. Trust also depends on supervisors being transparent when AI materially shapes comments or other guidance offered to candidates.
Equity and institutional context complete the conceptual picture. GenAI may support multilingual writers, distance candidates, part-time researchers, or students with limited access to frequent feedback [31,32]. However, access to high-quality tools, institutional licences, digital confidence, and supportive supervision is uneven [33]. A human-centred model must therefore locate responsibility at three levels. Candidates remain responsible for claims, sources, and final decisions; supervisors retain evaluative, pedagogical, and relational responsibility; and institutions must provide policy, training, approved systems, and data governance. GenAI may mediate these relationships, but its use should be assessed by whether it strengthens learning, transparency, trust, and independent research capacity rather than by whether it produces fluent output [34].
These foundations lead to three analytical propositions. First, the acceptability of an AI-supported activity depends on its function and context rather than on the tool name alone. Second, transparency is necessary but not sufficient: disclosed use can still be educationally inappropriate if it replaces the reasoning that the candidate must demonstrate. Third, responsibility cannot be delegated to the system. Candidates, supervisors, and institutions retain different but connected obligations even when a generated output appears accurate. These propositions provide the criteria used in the following sections to evaluate opportunities, ethical boundaries, feedback practices, and governance.

4. Current Uses and Opportunities

GenAI can support doctoral supervision across writing, literature engagement, research planning, feedback preparation, accessibility, and administrative coordination. These opportunities are attractive because doctoral work is cognitively demanding, iterative, and often conducted with long intervals between formal meetings. The value of AI assistance, however, does not lie in producing finished answers. It lies in helping candidates prepare, question, compare, and revise, and in allowing supervisors to redirect time from routine preparation towards substantive intellectual dialogue [9,10]. Each potential use therefore needs to be evaluated together with its learning purpose, intellectual risk, data sensitivity, and the safeguards required to preserve candidate ownership and supervisory judgement.
In doctoral writing, GenAI may help reorganise a section, test the coherence of an argument, clarify a sentence, identify unexplained transitions, or support translation and language editing. Such functions can reduce barriers for multilingual writers and allow attention to shift towards higher-order conceptual and methodological issues [5,31,35,36]. Yet writing is also a mode of thinking and identity formation. A grammatically improved passage may conceal weak understanding or flatten a distinctive scholarly voice. Responsible use therefore requires candidates to remain active editors: they should compare suggestions with their intended meaning, verify any factual additions, and be able to explain which changes were accepted or rejected. Supervisors should evaluate the reasoning behind the text, not only its surface fluency.
For literature engagement, AI can generate search terms, propose conceptual relationships, explain unfamiliar terminology, summarise difficult texts, or compare theoretical positions [37,38]. For research planning, it may help candidates refine preliminary questions, anticipate methodological tensions, prepare supervision agendas, construct timelines, or formulate issues for discussion [9,30]. These uses are particularly helpful at an exploratory stage, when a candidate needs alternatives or a structure for reflection. They are unreliable as substitutes for systematic searching, close reading, or methodological expertise. Generated summaries may omit debate, invent references, or reproduce dominant perspectives. Suggestions should therefore lead back to original sources and to discussion with the supervisor rather than being treated as evidence of comprehensive understanding.
Supervisors may use GenAI for low-risk preparatory work, including draft agendas, examples of feedback language, milestone checklists, formative exercises, or non-confidential administrative communication [17,39]. The technology can also be used pedagogically rather than merely operationally. Candidates can critique an AI-generated summary, compare machine feedback with supervisor feedback, identify errors in methodological advice, or analyse how a prompt shaped an answer [40,41]. Such activities make the system an object of critical inquiry and teach verification, epistemic caution, and prompt awareness. The supervisor must still review any generated material, adapt it to the candidate’s project and disciplinary context, and avoid delegating evaluative judgement or relational communication.
Accessibility benefits may arise for international candidates, multilingual writers, students with disabilities, part-time researchers, and distance candidates who have fewer informal opportunities for academic discussion [42,43]. GenAI may help organise complex material, explain conventions, or sustain momentum between meetings. It may also reduce anxiety associated with asking preliminary or seemingly basic questions. These benefits should not be assumed to be universal. Paid features, reliable connectivity, institutional licences, data-protection guarantees, and AI literacy vary across students and institutions. A technology presented as inclusive can become another source of inequality if access, training, and supervisory acceptance are uneven or if the system reproduces linguistic, cultural, and epistemic bias.
Across these uses, GenAI is most defensible as a reflective scaffold. Candidates should verify claims, consult original sources, record substantive assistance where appropriate, and remain able to reconstruct the reasoning behind final decisions. The supervisor’s role is to ask how the tool contributed to understanding rather than merely whether it was used. Efficiency is a secondary criterion: doctoral learning often depends on slow engagement with uncertainty, repeated revision, methodological deliberation, and response to critique [38,44]. Excessive automation may produce shallow synthesis, generic argumentation, or premature closure. A useful application should therefore strengthen the candidate’s capacity to work independently after the support is removed.
Coding and analytical support illustrate the importance of context. GenAI may explain syntax, suggest debugging strategies, propose possible categories, or help document an analytical workflow. In some disciplines, these uses resemble technical assistance; in others, the selection of a model, category, or interpretation is itself a central scholarly contribution. The candidate must therefore understand and validate any generated code or analytical suggestion, retain reproducible records where appropriate, and avoid entering protected data into unapproved systems. Administrative uses such as agendas, reminders, and routine emails are generally lower risk, but they still require attention to confidentiality and to the possibility that automated communication may become impersonal or misleading.
Table 4 summarises the main areas of use together with their potential benefits, principal risks, and supervisory safeguards. The table is intended as a decision aid: the same activity may be appropriate or inappropriate depending on whether the candidate remains intellectually engaged, whether the information is sensitive, and whether the output is verified and discussed.

5. Ethical Boundaries and Research Integrity

Ethical analysis of GenAI in doctoral supervision extends beyond plagiarism or technical compliance. It concerns how intellectual contribution is represented, how research material is governed, how generated claims acquire authority, and whether AI use supports or displaces the independent judgement that the doctorate is intended to demonstrate [45,46]. The central boundary is functional rather than tool-based. Grammar correction, brainstorming, and question generation may be legitimate when they help a candidate express or test original thinking. The same system becomes problematic when it performs unacknowledged analysis, synthesis, interpretation, or decision-making that should evidence the candidate’s competence. Ethical judgement therefore requires attention to the task, the stage of the doctorate, the discipline, the data, and the extent of human evaluation.
Academic integrity, authorship, and originality are closely connected. Doctoral work is never produced in complete isolation: supervisors, peers, editors, translators, and reviewers contribute in recognised ways. GenAI introduces a contribution that is less visible and difficult to classify. It may merely improve expression, or it may substantially shape an argument, literature synthesis, theoretical interpretation, or research design [46,47,48]. When that influence is substantial and undisclosed, the candidate’s original contribution becomes harder to evaluate. AI systems cannot qualify as authors or assume responsibility for errors and decisions. Human authors must therefore remain accountable for accuracy, originality, source use, and the coherence of the final work, irrespective of the technical assistance involved.
Transparency is consequently central. Candidates should be able to explain when, why, and how GenAI contributed to writing, literature work, planning, translation, coding assistance, or feedback interpretation [15,49]. Appropriate disclosure need not reproduce every minor interaction; it should be proportional to the intellectual significance of the use and to journal, thesis, or institutional requirements. Supervisors have parallel obligations. If AI materially shapes formative comments, examples, or expectations, students should not be left to assume that every element represents unaided expert judgement [48,50]. Disclosure is most useful when it supports an educational conversation about what was generated, what was checked, what was rejected, and how the final decision remained human.
Privacy, confidentiality, and data protection require stricter limits. Doctoral supervision routinely involves unpublished chapters, ethics applications, interview transcripts, field notes, participant data, clinical information, commercial material, institutional records, or politically sensitive evidence. Entering such material into inadequately governed systems may create risks involving data retention, unauthorised processing, intellectual property, informed consent, and research ethics [34,49]. A public literature review and a confidential transcript are not equivalent use cases. Institutions should identify approved platforms, explain their data conditions, and define material that may not be uploaded. Where sensitive data are involved, research ethics approval and data management plans should explicitly consider AI processing rather than treating the system as a neutral writing space.
Bias, hallucination, and epistemic authority create a further boundary. Generated outputs may be fluent and confidently expressed while being inaccurate, incomplete, fabricated, culturally narrow, or disconnected from the relevant discipline [38,51]. In doctoral research, these defects can distort literature mapping, theoretical framing, methods, or interpretation. They can also narrow inquiry by presenting dominant or highly represented perspectives as comprehensive. Candidates and supervisors must verify claims and references, compare generated material with primary sources, and ask which voices or forms of knowledge are absent. The apparent authority of polished language should never replace source evaluation, methodological scrutiny, or disciplinary contextualisation.
Overdependence is both an ethical and pedagogical risk. Doctoral education is intended to develop researchers who can read critically, make methodological decisions, respond to critique, and write persuasively. Repeated delegation of literature synthesis, revision, feedback interpretation, or analytical choice may weaken confidence and produce deskilling [7,52]. The relevant issue is not simply frequency of use, but whether the candidate’s capacity grows. Reflective questions—what was generated, what was independently checked, what was rejected, and what was learned—can convert AI use into evidence of judgement [8,53]. Supervisors may also vary the level of permitted support across stages so that scaffolding does not become a permanent authority.
Institutions must make these boundaries actionable through tiered and discipline-sensitive guidance. Low-risk uses, such as editing non-confidential prose, may require verification only. Conditional uses may require disclosure, supervisor approval, an authorised system, or an explanation of the candidate’s contribution. Prohibited uses should include transferring confidential or identifiable material to unapproved services and representing generated analysis as independent work. Baseline requirements should cover authorship, originality, source verification, disclosure, data protection, research ethics, and human accountability [17,54]. Policies should also provide routes for advice because acceptable use can vary across methods, disciplines, and stages of doctoral development.
Ethical responsibility also applies to supervisors’ own use. A supervisor who asks a system to summarise a confidential chapter, formulate evaluative comments, or advise outside the supervisor’s expertise may expose the candidate to risks that the candidate did not choose. Supervisory efficiency cannot justify weak verification or undisclosed delegation of judgement. Supervisors should use only material they are authorised to process, review generated content against the project and relevant scholarship, and remain prepared to explain the basis of every recommendation. Institutions should ensure that workload pressures do not create incentives to automate the relational and evaluative work that candidates reasonably expect from supervision.
Table 5 links the principal ethical issues to examples, risks, and recommended practices. Its purpose is to distinguish support from substitution and to show that a defensible practice normally combines transparency, verification, human responsibility, and appropriate data governance.

6. AI-Mediated Feedback and Doctoral Learning

Feedback in doctoral education is a long-term, dialogic process through which candidates learn disciplinary standards, evaluate their own work, and develop independent scholarly judgement [19,53,55,56,57]. It is embedded in knowledge of the research question, the evidence, the candidate’s history, and the wider supervisory relationship. A comment on a chapter may simultaneously address argument, disciplinary positioning, confidence, progress, and readiness for a milestone. GenAI may alter how feedback is prepared, interpreted, and acted upon, but it cannot be evaluated simply by the number or speed of comments produced. Its educational value depends on whether it deepens engagement with the research problem and improves the quality of dialogue between candidate and supervisor.
Supervisors may use AI for low-risk preparatory tasks such as generating alternative phrasings, questions, examples, or structural prompts. It may help organise notes, identify recurrent surface problems, or create contrasting examples for discussion. Such uses can release time for substantive engagement, especially where supervisory workloads are high. Nevertheless, the supervisor must verify and contextualise every material contribution, avoid uploading confidential work to unauthorised systems, and retain responsibility for the evaluative message. Automated comments should not be passed to a candidate as though they were expert judgement without review. Feedback that ignores the project’s methods, disciplinary debate, or developmental history may be fluent but educationally unhelpful.
Candidates may use GenAI for preliminary self-review, asking about clarity, coherence, terminology, logical flow, or possible revision options [20,21]. This can help identify surface-level problems, formulate more precise questions, and maintain momentum between meetings. Multilingual, distance, and part-time candidates may particularly value immediate assistance with language or organisation. AI may also help a student translate an unclear supervisor comment into questions for the next meeting. These uses are productive when they prepare a more focused supervisory exchange. They become problematic when the candidate pastes comments into a system and accepts a rewritten passage without understanding the underlying academic problem.
AI-generated feedback remains provisional. It may sound constructive while lacking disciplinary nuance, methodological sensitivity, reliable sources, or awareness of previous drafts and emotional context [19,58,59,60]. It may recommend standard structures that weaken an original argument, focus on easily detectable surface issues, or generate confident but irrelevant criticism. It cannot know whether a particular comment has already been discussed, whether a candidate is ready for a specific level of challenge, or how feedback will affect motivation. Machine feedback should therefore be evaluated against the research aims, the relevant literature, disciplinary standards, and supervisor guidance rather than treated as an independent assessment.
Feedback literacy provides the appropriate pedagogical lens: candidates must learn to appreciate, interpret, evaluate, and act on feedback rather than simply receive it [8,53]. GenAI can support this capacity by explaining terminology, generating questions for clarification, comparing revision strategies, or revealing that different readers may prioritise different issues. It weakens feedback literacy when it supplies a response before the candidate has attempted to diagnose the problem. The same ambivalence applies to autonomy. Preliminary self-review can make candidates more active in dialogue [35,61], but habitual reliance on generated criticism may prevent them from internalising disciplinary standards. The test is whether AI use expands the candidate’s evaluative repertoire or makes judgement increasingly dependent on the tool.
Supervisors can humanise AI-mediated feedback by asking candidates to bring generated comments to meetings, identify which suggestions were useful, justify accepted and rejected changes, and compare machine advice with disciplinary expectations [20,53]. Human feedback remains essential for setting priorities, recognising effort, addressing uncertainty or confidence, and connecting a textual revision to the candidate’s development as a researcher. It also enables disagreement and negotiation rather than presenting feedback as a fixed answer. Supervisory agreements and doctoral training should therefore make substantive AI use in feedback visible, establish expectations for confidentiality and verification, and treat critical comparison of human and machine feedback as a legitimate learning activity.
Timing and sequence are also important. Immediate AI feedback can be useful before a meeting, but constant availability may encourage endless revision, fragment attention, or reduce tolerance for unresolved questions. Doctoral feedback frequently develops across cycles: a candidate submits work, receives comments, interprets priorities, attempts revision, and discusses the result. GenAI should support rather than collapse this cycle. Supervisors can ask candidates to attempt their own diagnosis before consulting a system, use machine comments to generate discussion rather than final edits, and revisit whether reliance is decreasing as competence grows. This staged approach treats scaffolding as temporary and developmental rather than as an unlimited parallel supervisor.

7. Trust, Power, Agency, and Supervisory Relationships

GenAI affects not only supervisory tasks, but also the relationship through which candidates learn to think, decide, and act as researchers. Doctoral supervision is sustained over time and involves intellectual dependence, evaluative authority, identity formation, and often emotional vulnerability. The relational consequences of AI should therefore be assessed through trust, power, agency, scholarly identity, and emotional support rather than through productivity alone [3,12]. A system that improves a paragraph may still weaken supervision if it makes candidate understanding opaque or allows participants to avoid necessary dialogue. Conversely, a modest use can be valuable when it makes uncertainty visible and supports more substantive human interaction.
Trust depends on credible representations of understanding and on fair, context-sensitive guidance [1,4]. Undisclosed AI-generated arguments or revisions can make candidate competence difficult to judge. Unreviewed AI-generated supervisor comments can make supervisory engagement appear impersonal or inauthentic. The problem is intensified when policy is ambiguous: candidates may hide legitimate assistance because they fear allegations, while supervisors may interpret polished writing as suspicious without evidence [18,48]. Trust in AI-mediated supervision must therefore be actively constructed through explicit expectations, proportionate disclosure, and discussion of the reasoning process. The goal is not continuous surveillance, but sufficient transparency for developmental and evaluative judgements to remain credible.
Power already enters supervision through authority, expertise, access to disciplinary networks, evaluation, and institutional status [2,4]. GenAI adds new asymmetries. Supervisors may control what is permitted while candidates possess greater practical AI literacy; alternatively, candidates may feel unable to question AI-assisted feedback presented by an authority figure. Access also matters. Some candidates have premium tools, institutional licences, training, and supportive supervisors, whereas others depend on limited systems or unclear guidance [33,36]. Responsible practice should avoid unilateral prohibition and uncritical acceptance. Supervisors should invite dialogue about the strengths and limitations of the technology, and institutions should prevent access and policy differences from creating hidden educational advantages.
Doctoral agency is strengthened when candidates use AI to generate alternatives, rehearse explanations, identify uncertainty, or prepare questions while retaining final judgement [7,11,29]. It is weakened when generated suggestions determine what literature matters, how an argument should be framed, or how criticism should be answered. Agency is not equivalent to producing a finished text; it involves owning the decisions through which the text was created. Scholarly identity develops through repeated participation in disciplinary reasoning, critique, and communication [61,62]. Candidates should therefore be able to explain how AI affected their process and why their final position reflects their own evaluation. Supervisors can support this by focusing discussion on choices, evidence, and rejected alternatives rather than attempting to infer authorship from style alone.
The emotional dimension requires similar caution. Doctoral work can involve isolation, anxiety, uncertainty, and fear of failure. GenAI may offer immediately available encouragement, help candidates articulate a concern, or provide a low-stakes space for rehearsing a difficult question. It does not, however, recognise sustained distress, understand the history of the relationship, negotiate expectations, or assume a duty of care. It becomes harmful when it substitutes for conversations about weak arguments, delayed progress, methodological difficulties, conflict, or wellbeing [13,18]. Supervisors and institutions must preserve accessible human support and should not frame availability of an AI system as a replacement for mentoring or pastoral responsibility.
Transparent and reflective use can nevertheless strengthen relationships. Candidates may arrive with clearer questions, supervisors may use examples to stimulate discussion, and both parties may compare interpretations of a generated output. Relational clarity requires expectations to be established early and revisited as the project, methods, data, and available systems change. Supervisors should model critical engagement, candidates should explain the basis of final decisions, and supervisory teams should document material agreements. Institutions should provide routes for advice when ethical, disciplinary, or data-governance risks exceed the expertise of the participants. These arrangements make AI use part of doctoral learning rather than an invisible influence operating outside the relationship.
Relational effects may also differ in supervisory teams and cross-cultural settings. Multiple supervisors can hold different expectations about acceptable assistance, disclosure, or writing support, leaving candidates to navigate inconsistent rules. Cultural norms concerning authority, face, authorship, and disagreement can influence whether candidates feel able to disclose use or challenge an AI-assisted comment. Supervisory teams should therefore agree on a common position, communicate it jointly, and provide a process for resolving differences. Periodic review is preferable to a fixed statement because the candidate’s level of independence, the research methods, and the available systems change over the life of a doctorate.

8. Institutional Governance and Quality Assurance for Responsible AI Use

Responsible GenAI use cannot be left to individual supervisors and candidates. Doctoral research involves original intellectual work, confidential information, long-term assessment, publication, and discipline-specific practices. Without institutional support, expectations may vary according to a supervisor’s personal confidence or attitude, creating uncertainty, inequity, and avoidable risk [50,54]. Institutions must therefore provide a governance framework that enables informed judgement rather than relying on detection or blanket prohibition. Such a framework should recognise that the same tool may be low risk when editing public prose, higher risk when shaping analysis, and unacceptable when processing identifiable data in an unauthorised environment.
Policy should address the principal doctoral use cases—writing, literature synthesis, planning, coding or analytical support, feedback interpretation, translation, administration, and supervision meetings—and classify them by intellectual and data risk [34,63]. It should state when use is permitted, when disclosure, supervisor agreement, or an authorised platform is required, and when a task or dataset may not be transferred to an external system. Doctoral-specific guidance is necessary because the central question is not only misconduct, but whether the thesis or research output demonstrates the candidate’s original contribution. Authorship, accuracy, originality, and final responsibility must remain human [14,45,47,48].
Governance must be accompanied by AI literacy. Supervisors need enough understanding to discuss capabilities, limitations, bias, hallucination, confidentiality, and overdependence, and to evaluate AI-supported work without becoming machine learning specialists [32,40]. Candidates require training in search and source verification, prompt and output evaluation, disclosure, data protection, and ethical decision-making [34,64]. Training should use discipline-relevant cases rather than generic warnings and should include situations in which use is appropriate, conditional, or prohibited. Shared development reduces fear, supports consistent expectations, and enables supervisors and candidates to discuss AI as part of research practice rather than only after a suspected problem.
Doctoral schools, graduate colleges, ethics committees, libraries, information security teams, and research offices have complementary roles. They can provide approved tools, model clauses for supervisory agreements, disclosure templates, discipline-sensitive examples, ethics guidance, and routes for escalation. Governance should be connected to ethics applications, data management planning, research integrity, and milestone review rather than isolated in a general AI policy. It should also address equity by providing accessible training, institutional licences where justified, disability support, and alternatives for candidates who cannot or do not wish to use particular systems. Flexibility across disciplines is necessary, but it should operate within a common institutional baseline.
That baseline should require candidates to remain accountable for all submitted work, disclose substantive AI contributions, verify claims and references, protect confidential material, and retain evidence of their reasoning where the intellectual or data risk warrants it. Supervisors should discuss acceptable use at the beginning of candidature, review expectations when methods or tools change, and verify any AI-supported material that enters formal feedback. Institutions should identify authorised systems, publish clear routes for advice, and update guidance as technologies and contractual data conditions evolve. These minimum expectations preserve local judgement while making responsibility and quality assurance visible.
Quality assurance should extend beyond policy publication and operate as a documented and recurring institutional process. A doctoral AI quality-assurance system should connect formally authorised policy with operational guidance, approved technological environments, AI literacy development, proportionate documentation of substantive AI use, periodic evaluation, and improvement actions. Formal authorization by the appropriate senior academic or research authority establishes institutional responsibility for the policy and provides a common baseline across doctoral programmes, while disciplinary units retain flexibility to interpret that baseline in relation to methods, data sensitivity, authorship conventions, and disciplinary practice.
Implementation should generate proportionate evidence rather than rely solely on stated compliance. Institutions should maintain version-controlled policies and guidance, approved-system registers, discipline-sensitive examples of permitted, conditional, and prohibited practices, disclosure templates, and mechanisms for recording substantive AI use where intellectual or data risk warrants it. At relevant doctoral milestones, candidates and supervisors should each contribute a brief written evaluation of how GenAI has been used, what material was independently verified, whether the candidate remains able to explain and defend relevant intellectual decisions, what uncertainties or weaknesses have emerged, and whether current arrangements continue to support learning, integrity, trust, confidentiality, autonomy, and equitable participation. The purpose of this evaluation is developmental and preventive rather than surveillance-oriented.
Evaluation should lead to documented improvement where weaknesses, risks, or recurrent uncertainties are identified. Depending on the issue, improvement actions may include clarification of supervisory expectations, additional source-verification or data-protection guidance, targeted AI literacy development, refresher training, case-based workshops, revision of supervisory agreements, or referral to ethics, data-governance, information-security, or disciplinary expertise. Subsequent milestone review should determine whether the agreed action has addressed the identified issue. At the institutional level, aggregated and appropriately anonymised findings can inform revisions to guidance, training provision, approved-system requirements, and support resources. Quality assurance is therefore conceived here as a continuous cycle of expectation setting, documentation, evaluation, improvement, and re-evaluation rather than as a one-time compliance exercise.
Institutional procurement and technology approval are also part of responsible governance. Universities should assess whether a system retains prompts or uploaded files, uses them for model training, transfers data across jurisdictions, supports institutional access controls, and permits audit or deletion. Approval should not be based only on functionality or price. Data protection, information security, accessibility, intellectual property, and research integrity expertise should be involved, and candidates should receive plain-language guidance on what institutional approval does and does not guarantee. Even an approved system may be unsuitable for a particular dataset or intellectual task.
To translate these governance and quality-assurance principles into observable responsibilities, Table 6 presents a minimum accountability and continuous-improvement framework for supervisors, doctoral researchers, supervisory teams, and institutions. The framework links preventive expectations and documentation with periodic evaluation, improvement actions, and institutional learning. The recommendations are not intended as a universal script; rather, they can be adapted to disciplinary methods, institutional policies, and research risks while preserving common requirements for transparency, verification, confidentiality, human judgement, and equitable support.
Taken together, these recommendations form a recurring accountability and quality-assurance cycle. Expectations are agreed and documented; substantive AI use is made visible; outputs, sources, and decisions are verified; candidate and supervisor evaluations identify strengths, weaknesses, risks, and development needs; proportionate improvement actions are agreed and documented; arrangements are re-evaluated at relevant milestones; and elevated risks are referred to ethics, data governance, information security, or disciplinary expertise. At the institutional level, aggregated findings provide evidence for revising policy, training, resources, and support. This cycle prevents governance from becoming a one-time policy declaration and gives supervisory teams a practical basis for evaluating whether an AI-enabled practice strengthens learning, integrity, trust, and equitable participation. It also makes clear that efficiency cannot compensate for weakened judgement, reduced candidate autonomy, or unaccountable decision-making.

9. Research Gaps and Future Agenda

Research on GenAI in doctoral supervision remains emergent. Existing work establishes plausible opportunities and risks, but the evidence is uneven across disciplines, countries, doctoral models, and stages of candidature [22,29]. Future research should examine not only what systems can produce, but how their use changes learning, judgement, ethical practice, relationships, and institutional decision-making in the distinctive context of original, long-term research. Studies also need to distinguish reported use from actual practice and short-term satisfaction from longer-term development.
First, empirical research should document not only how candidates and supervisors actually use GenAI, including informal or undisclosed practices [52,65], but also why particular practices emerge, persist, change, or remain hidden. Interviews and surveys can examine motives, attitudes, perceived usefulness, confidence, workload pressures, language or accessibility needs, and expectations concerning acceptable use, while diaries, prompt or decision logs, and observations of supervision can show how these influences translate into everyday practice. Research should also examine relational and institutional drivers, including supervisor expectations, access to advanced tools and training, disciplinary norms, policy clarity, and perceptions of risk or sanction. Such designs would help distinguish individual preferences from contextual conditions that enable, constrain, or redirect GenAI use. Feedback studies should compare machine and supervisor comments not only for accuracy, but also for disciplinary nuance, dialogic value, emotional tone, and candidate response [19,58]. Research should examine how candidates combine multiple sources of feedback and whether critical comparison improves feedback literacy and writing development.
Second, the benefits, harms, and developmental trade-offs of sustained GenAI use require longitudinal study. Research should investigate how candidates and supervisors negotiate authorship, originality, disclosure, confidentiality, and acceptable assistance in real projects [47,49], while also examining when AI-supported practices enhance or undermine doctoral development. Potential benefits may include improved access to feedback, support for reflection, language or organisational assistance, and temporary scaffolding for difficult tasks; potential harms include overdependence, weakened independent reading or writing, reduced methodological judgement, inaccurate or biassed outputs, and displacement of supervisory dialogue. These outcomes should not be treated as inherent properties of the technology: the same practice may be beneficial or detrimental depending on its purpose, frequency, transparency, disciplinary context, stage of candidature, and the candidate’s capacity to evaluate and defend the resulting decisions. Multi-year designs could therefore examine whether AI operates as temporary scaffolding that becomes internalised or as a dependency that weakens independent scholarly judgement [7,66]. Milestone assessments, oral explanation, and process-based evidence should complement textual outputs because final products alone may not reveal how understanding, autonomy, and scholarly identity are developing.
Third, comparative and implementation research is needed across disciplines, institutions, countries, cultural contexts, and doctoral models [4,62]. The acceptability and consequences of AI-assisted coding, language editing, literature synthesis, or qualitative analysis are unlikely to be uniform. Studies should examine who benefits from access to advanced tools, licences, training, language support, and supervisor confidence [42,43]. They should also evaluate how institutional policies are interpreted, whether they encourage open dialogue or hidden use, and which forms of supervisor development and doctoral AI literacy training lead to responsible practice [40,64].
Relational research should additionally examine how GenAI use interacts with social capital within doctoral education, including candidates’ access to trusted academic relationships, reciprocal support, disciplinary networks, and informational resources. Of particular interest is whether AI-mediated practices broaden access to such relational resources or, conversely, reduce opportunities for the human interaction, academic socialisation, and network participation through which doctoral researchers develop scholarly belonging and support.
Future studies should also define outcomes that reflect doctoral development rather than relying only on satisfaction, speed, or text quality. Relevant outcomes include the candidate’s ability to explain decisions, detect errors, transfer learning to unaided tasks, respond to critique, protect data, and exercise independent methodological judgement. Supervisory and relational outcomes may include feedback quality, clarity of expectations, trust, workload distribution, access to supportive academic relationships and networks, perceived reciprocity and scholarly belonging, and the frequency with which difficult issues are discussed rather than hidden. Institutional evaluations should examine consistency, equity, and whether guidance enables responsible experimentation without normalising uncritical dependence.
These outcomes also provide a starting point for empirical operationalization of the five-dimensional framework. Possible indicators include disclosure and source-verification practices for ethical accountability; critical evaluation of AI-generated feedback and transfer to unaided tasks for feedback and learning; independent explanation and defence of research decisions for agency and autonomy; openness about AI use, clarity of expectations, and perceived trust for supervisory relationships; and consistency of guidance, access to approved systems, AI literacy provision, and equity of support for institutional governance. These indicators are illustrative rather than validated measures and should be refined and tested across disciplinary and institutional contexts.
Methodological innovation will be needed to study practices that are private, rapidly changing, and sometimes undisclosed. Self-reporting should be combined with artefacts such as revision histories, decision logs, feedback records, or supervised demonstrations, with appropriate consent and privacy safeguards. Participatory designs can involve candidates and supervisors in defining meaningful outcomes and interpreting the difference between helpful scaffolding and inappropriate substitution. Research should also report system versions, access conditions, prompt procedures, and institutional context so that findings can be compared and not attributed to ‘GenAI’ as though it were a stable, uniform intervention.
Table 7 organises these priorities by research area, guiding question, suggested methodology, and relevance to doctoral supervision. The agenda emphasises empirical, longitudinal, comparative, ethical, and explicitly relational designs so that future policy and pedagogy can be informed by evidence rather than by assumptions about either technological promise or inevitable harm.

10. Limitations

This review has several limitations. The search centred on Scopus and Web of Science and prioritised English-language literature. Relevant scholarship in regional databases, professional repositories, institutional collections, or other languages may therefore be underrepresented. The review was designed as a critical and integrative narrative synthesis rather than an exhaustive systematic review, and the inclusion of foundational work from adjacent fields necessarily involved judgement about conceptual relevance.
The database search sets overlapped substantially. Although query-level counts and the final analytical corpus are documented, the original workflow did not retain unique counts for cross-search deduplication, title and abstract screening, full-text assessment, or exclusions by reason. The review therefore reports this gap directly and does not present a retrospective PRISMA-style flow. Reproducing those stages numerically would require the original database exports and screening ledger; future updates should preserve both.
Source selection, extraction, thematic coding, and interpretation necessarily involved judgement about conceptual relevance and evidential weight. Predefined eligibility criteria, a structured extraction framework, duplicate checking of the final bibliography, evidence-type separation, and thematic mapping improve transparency but cannot eliminate selection or interpretive bias. The corpus also combines empirical studies, conceptual papers, reviews, methodological sources, and policy guidance. These materials support different claim types and cannot be reduced to a single numerical risk-of-bias score.
Finally, GenAI systems, institutional policies, publication requirements, and data-protection conditions are changing rapidly. Much of the available evidence is recent, self-reported, cross-sectional, or concentrated in specific disciplines and higher education systems. The conclusions should therefore be read as a structured synthesis of the available literature and a set of propositions for further empirical testing, not as universal rules or as a technical or legal evaluation of particular AI systems.

11. Conclusions

This review positions GenAI as a socio-technical influence on doctoral supervision rather than a neutral productivity tool. Its contribution lies in integrating ethical accountability, feedback and learning, agency and autonomy, supervisory trust and power, and institutional governance within one human-centred framework. The framework shows that GenAI can support writing, the literature engagement, planning, accessibility, and feedback preparation, but that each benefit is conditional on critical evaluation, visibility of substantive assistance, and protection of the developmental work through which candidates become independent scholars.
Taken together, the reviewed evidence and normative guidance support a set of observable and distributed practices for responsible GenAI use. Candidates should disclose substantive assistance, verify claims and sources, protect confidential material, and remain able to defend every intellectual decision. Supervisors should establish and revisit expectations, evaluate the candidate’s reasoning, contextualise any AI-supported feedback, and preserve the relational dimensions of mentoring. Institutions should provide risk-tiered policy, authorised systems, ethics and data-governance integration, recurrent training, and equitable access. These recommendations translate the ethical, developmental, and governance principles identified across the review into an accountability cycle that can be revisited as projects and technologies change.
The decisive criterion is educational and ethical rather than technological. GenAI adds value when it improves preparation, reflection, accessibility, and dialogue while strengthening the candidate’s capacity to reason and act independently. It becomes detrimental when textual fluency conceals weak understanding, when sensitive material is exposed, when bias or fabricated evidence enters the research process, or when human responsibility is displaced. Its place in doctoral supervision should therefore be judged by whether it supports the formation of independent, critical, transparent, and responsible researchers and sustains the trust on which effective supervision depends. The strongest evidence of successful integration will not be the presence of sophisticated tools, but candidates who can use or decline them deliberately, justify their choices, and continue to perform the relevant scholarly work without surrendering responsibility.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/encyclopedia6100209/s1, Supplementary Material S1: Search Strategy and Literature Mapping, including database search sets, retrieval counts, inclusion and exclusion logic, thematic literature mapping, and a SANRA-informed checklist for narrative review quality alignment. Table S1: Database search sets used to support source selection; Table S2: Eligibility criteria and rationale for source selection; Table S3: Literature mapping by thematic area; Table S4: SANRA-informed checklist for narrative review quality alignment; Figure S1: Workflow of source identification, eligibility assessment, and critical integrative synthesis.

Author Contributions

Conceptualization, C.G.G.P., H.E.L.G. and C.P.A.A.; methodology, C.G.G.P., L.Y.P.P., A.A.L. and H.E.L.G.; investigation, L.Y.P.P., A.A.L., L.M.H.R., Z.R.D.T. and F.M.R.W.; data curation, L.Y.P.P., A.A.L., L.M.H.R., Z.R.D.T. and F.M.R.W.; formal analysis, C.G.G.P., H.E.L.G., C.P.A.A., W.M.R.-S., C.M.-d.-S. and J.M.-A.; writing—original draft preparation, C.G.G.P. and H.E.L.G.; writing—review and editing, all authors; visualisation, H.E.L.G. and F.M.R.W.; supervision, H.E.L.G., C.P.A.A. and W.M.R.-S.; project administration, C.G.G.P. and H.E.L.G. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

No new data were created or analysed in this study. Data sharing is not applicable to this article. The search strategy, search sets, inclusion and exclusion criteria, the literature mapping, and SANRA-informed checklist supporting the review are provided in Supplementary Material S1.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Vähämäki, M.; Saru, E.; Palmunen, L.-M. Doctoral Supervision as an Academic Practice and Leader–Member Relationship: A Critical Approach to Relationship Dynamics. Int. J. Manag. Educ. 2021, 19, 100510. [Google Scholar] [CrossRef] [Scilit]
  2. Courtney, S.A.; Alexander, A.N. Doctoral Preparation in Mathematics Education: Graduates’ Perspectives on Program Features, Expertise Development, and Scholarly Learning Cultures. Int. J. Sci. Math. Educ. 2026, 24, 69. [Google Scholar] [CrossRef] [Scilit]
  3. Li, Y.; Xu, W.; Chen, J. PhD Student-Supervisor Relationship and Its Impacts: A Perspective of the Interpersonal Relationship Model. Front. Educ. 2025, 10, 1570137. [Google Scholar] [CrossRef] [Scilit]
  4. Kitano, N.; Aldous, C.; Race, D.; Clegg, K. The Taboo of Power Dynamics in Doctoral Supervision: Cross-Cultural Insights from Australia, South Africa and the South Pacific Region. High. Educ. Q. 2026, 80, e70155. [Google Scholar] [CrossRef] [Scilit]
  5. Becker, S.; Jacobsen, M.; Friesen, S. Four Supervisory Mentoring Practices That Support Online Doctoral Students’ Academic Writing. Front. Educ. 2025, 10, 1521452. [Google Scholar] [CrossRef] [Scilit]
  6. Huang, X.; Li, H.; Shcheglova, I. Exploring the Role of Generative Artificial Intelligence (GenAI) in Research Activities in Shaping Doctoral Students’ Well-Being. Euro J. Educ. 2026, 61, e70750. [Google Scholar] [CrossRef] [Scilit]
  7. Guo, J. The Augmented Scholar: How Generative AI Usage Influences Research Self-Efficacy among Chinese Doctoral Students. High. Educ. Res. Dev. 2026, 1–17. [Google Scholar] [CrossRef] [Scilit]
  8. Peng, Y. Toward an Understanding of Doctoral Students’ Feedback Literacy in Academic Publishing: An Ecological Perspective. Humanit. Soc. Sci. Commun. 2025, 12, 1819. [Google Scholar] [CrossRef] [Scilit]
  9. Cowling, M.; Crawford, J.; Allen, K.-A.; Wehmeyer, M. Using Leadership to Leverage ChatGPT and Artificial Intelligence for Undergraduate and Postgraduate Research Supervision. Australas. J. Educ. Technol. 2023, 39, 89–103. [Google Scholar] [CrossRef] [Scilit]
  10. Dai, Y.; Lai, S.; Lim, C.P.; Liu, A. ChatGPT and Its Impact on Research Supervision: Insights from Australian Postgraduate Research Students. Australas. J. Educ. Technol. 2023, 39, 74–88. [Google Scholar] [CrossRef] [Scilit]
  11. Iatrellis, O.; Bania, A.; Samaras, N.; Kosmopoulou, I.; Panagiotakopoulos, T. ChatGPT in Doctoral Supervision: Proposing a Tripartite Mentoring Model for AI-Assisted Academic Guidance. Int. J. Dr. Stud. 2025, 20, 009. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  12. Lai, S.; Liu, S.; Dai, Y.; Lim, C.P.; Liu, A. The Impacts and Tensions of Generative AI on Doctoral Students’ Supervisory and Peer Dynamics: An Activity Theory Analysis. Australas. J. Educ. Technol. 2025, 41, 5. [Google Scholar] [CrossRef] [Scilit]
  13. Brown, A.; Rossouw, J. Artificial Intelligence as a Reflexive Collaborator in Graduate Studies Supervision. Transform. High. Educ. 2026, 11, a657. [Google Scholar] [CrossRef] [Scilit]
  14. Bittle, K.; El-Gayar, O. Generative AI and Academic Integrity in Higher Education: A Systematic Review and Research Agenda. Information 2025, 16, 296. [Google Scholar] [CrossRef] [Scilit]
  15. Balalle, H.; Pannilage, S. Reassessing Academic Integrity in the Age of AI: A Systematic Literature Review on AI and Academic Integrity. Soc. Sci. Humanit. Open 2025, 11, 101299. [Google Scholar] [CrossRef] [Scilit]
  16. Yusuf, A.; Pervin, N.; Román-González, M. Generative AI and the Future of Higher Education: A Threat to Academic Integrity or Reformation? Evidence from Multicultural Perspectives. Int. J. Educ. Technol. High. Educ. 2024, 21, 21. [Google Scholar] [CrossRef] [Scilit]
  17. Wang, H.; Dang, A.; Wu, Z.; Mac, S. Generative AI in Higher Education: Seeing ChatGPT through Universities’ Policies, Resources, and Guidelines. Comput. Educ. Artif. Intell. 2024, 7, 100326. [Google Scholar] [CrossRef] [Scilit]
  18. Qi, J.; Tuxworth, J.; Gomes, C. Navigating Supervisory Conversations About Gen AI: Taboo and Supervisory Power Dynamics. High. Educ. Q. 2026, 80, e70151. [Google Scholar] [CrossRef] [Scilit]
  19. Jensen, L.X.; Bearman, M.; Boud, D.; Konradsen, F. Feedback Encounters in Doctoral Supervision: The Role of Generative AI Chatbots. Assess. Eval. High. Educ. 2025, 51, 849–862. [Google Scholar] [CrossRef] [Scilit]
  20. Khuder, B.; Cervin-Ellqvist, M.; Vander Borght, M.; Wingrove, P. Academic Socialisation through AI and Peer Feedback in Doctoral Writing: Information and Orientation in Feedback Ecologies. Assess. Eval. High. Educ. 2026, 1–20. [Google Scholar] [CrossRef] [Scilit]
  21. Urzúa, C.A.C.; Ranjan, R.; Saavedra, E.E.M.; Badilla-Quintana, M.G.; Lepe-Martínez, N.; Philominraj, A. Effects of AI-Assisted Feedback via Generative Chat on Academic Writing in Higher Education Students: A Systematic Review of the Literature. Educ. Sci. 2025, 15, 1396. [Google Scholar] [CrossRef] [Scilit]
  22. Thong, C.L.; Atallah, Z.; Islam, S.; Lim, W.; Cherukuri, A.K. AI-Powered Tools for Doctoral Supervision in Higher Education: A Systematic Review. J. Info. Know. Mgmt. 2025, 24, 2530001. [Google Scholar] [CrossRef] [Scilit]
  23. Whittemore, R.; Knafl, K. The Integrative Review: Updated Methodology. J. Adv. Nurs. 2005, 52, 546–553. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  24. Torraco, R.J. Writing Integrative Literature Reviews: Guidelines and Examples. Hum. Resour. Dev. Rev. 2005, 4, 356–367. [Google Scholar] [CrossRef] [Scilit]
  25. Grant, M.J.; Booth, A. A Typology of Reviews: An Analysis of 14 Review Types and Associated Methodologies. Health Info Libr. J. 2009, 26, 91–108. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  26. Tyndall, D.E.; Firnhaber, G.C.; Kistler, K.B. An Integrative Review of Threshold Concepts in Doctoral Education: Implications for PhD Nursing Programs. Nurse Educ. Today 2021, 99, 104786. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  27. Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 Statement: An Updated Guideline for Reporting Systematic Reviews. BMJ 2021, 372, n71. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  28. Baethge, C.; Goldbeck-Wood, S.; Mertens, S. SANRA—A Scale for the Quality Assessment of Narrative Review Articles. Res. Integr. Peer Rev. 2019, 4, 5. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  29. Susnjak, T.; McIntosh, T.; Liu, T.; Watters, P. A Design Science Blueprint for an Orchestrated AI Assistant in Doctoral Supervision. Technol. Pedagog. Educ. 2026, 1–27. [Google Scholar] [CrossRef] [Scilit]
  30. Naganuma, S.; Minematsu, T.; Matsueda, K.; Oshima, J. Doctoral Students’ Interdisciplinary Research Proposal Development Supported by Generative AI. Educ. Inf. Technol. 2026, 31, 3743–3779. [Google Scholar] [CrossRef] [Scilit]
  31. Ou, A.W.; Tai, K.W.H.; Wang, X. The Emergence of Academic Writers: Multilingual Doctoral Students’ Translanguaging and Transpositioning in AI-Mediated Academic Writing. J. Engl. Acad. Purp. 2026, 79, 101613. [Google Scholar] [CrossRef] [Scilit]
  32. Yao, G.; Fan, L. L2 Writers’ Critical AI Literacy in AI-Assisted Academic Writing: Scale Development and Validation. Int. Rev. Appl. Linguist. Lang. Teach. 2026. [Google Scholar] [CrossRef] [Scilit]
  33. Zhang, L.; Li, J.; Tsung, L. Activity Systems in Transition: GenAI-Mediated Thesis Writing Strategies and Contradictions among Chinese Doctoral Students in Australia. System 2026, 138, 104002. [Google Scholar] [CrossRef] [Scilit]
  34. Miao, F.; Holmes, W. Guidance for Generative AI in Education and Research; UNESCO: Paris, France, 2023; ISBN 978-92-3-100612-8. [Google Scholar]
  35. Parker, J.L.; Richard, V.M.; Acabá, A.; Escoffier, S.; Flaherty, S.; Jablonka, S.; Becker, K.P. Negotiating Meaning with Machines: AI’s Role in Doctoral Writing Pedagogy. Int. J. Artif. Intell. Educ. 2025, 35, 1218–1238. [Google Scholar] [CrossRef] [Scilit]
  36. Hoomanfard, M.; Shamsi, Y. Generative AI in Dissertation Writing: L2 Doctoral Students’ Self-Reported Use, AI-Giarism, and Perceived Training Needs. J. Engl. Acad. Purp. 2025, 78, 101570. [Google Scholar] [CrossRef] [Scilit]
  37. Kumar, S.; Gunn, A. Doctoral Students’ Reflections on Generative Artificial Intelligence (GenAI) Use in the Literature Review Process. Innov. Educ. Teach. Int. 2025, 62, 1395–1408. [Google Scholar] [CrossRef] [Scilit]
  38. Walters, W.H.; Wilder, E.I. Fabrication and Errors in the Bibliographic Citations Generated by ChatGPT. Sci. Rep. 2023, 13, 14045. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  39. Jin, Y.; Yan, L.; Echeverria, V.; Gašević, D.; Martinez-Maldonado, R. Generative AI in Higher Education: A Global Perspective of Institutional Adoption Policies and Guidelines. Comput. Educ. Artif. Intell. 2025, 8, 100348. [Google Scholar] [CrossRef] [Scilit]
  40. Weikang, L.; Shiyin, L.; Xiaomo, Q. The Impact of Artificial Intelligence Literacy on Doctoral Students’ Innovative Behaviour from the Perspective of Technology Affordance. Euro J. Educ. 2025, 60, e70245. [Google Scholar] [CrossRef] [Scilit]
  41. Wu, C.; Moorhouse, B.L.; Wan, Y.; Wu, M. Exploring PhD Students’ Utilization of Generative AI in Academic Writing for Publication Purposes: Insights for EAP. J. Engl. Acad. Purp. 2026, 79, 101612. [Google Scholar] [CrossRef] [Scilit]
  42. Pretorius, L.; Huynh, H.-H.; Pudyanti, A.A.A.R.; Li, Z.; Noori, A.Q.; Zhou, Z. Empowering International PhD Students: Generative AI, Ubuntu, and the Decolonisation of Academic Communication. Internet High. Educ. 2025, 67, 101038. [Google Scholar] [CrossRef] [Scilit]
  43. Cowe, H. Using AI to Help Generate Disability Scholarship. Disabil. Soc. 2026, 41, 1149–1153. [Google Scholar] [CrossRef] [Scilit]
  44. Sanz-Tejeda, A.; Domínguez-Oller, J.C.; Baldaquí-Escandell, J.M.; Gómez-Díaz, R.; García-Rodríguez, A. The Impact of Generative AI on Academic Reading and Writing: A Synthesis of Recent Evidence (2023–2025). Front. Educ. 2026, 10, 1711718. [Google Scholar] [CrossRef] [Scilit]
  45. Pratiwi, H.; Suherman; Hasruddin; Ridha, M. Between Shortcut and Ethics: Navigating the Use of Artificial Intelligence in Academic Writing Among Indonesian Doctoral Students. Euro J. Educ. 2025, 60, e70083. [Google Scholar] [CrossRef] [Scilit]
  46. COPE Council. Authorship and AI Tools; Committee on Publication Ethics: Hampshire, UK, 2024; Available online: https://publicationethics.org/guidance/cope-position/authorship-and-ai-tools (accessed on 9 July 2026).
  47. Bouaziz, K.; Eddouada, S. AI-Assisted Academic Writing and Publishing in Morocco: Doctoral Students’ Ethical Decision-Making Accounts. AWEJ 2026, 3, 268–286. [Google Scholar] [CrossRef] [Scilit]
  48. Wang, Y.; Zhao, L. Toward the Transparent Use of Generative Artificial Intelligence in Academic Articles. J. Sch. Publ. 2024, 55, 467–484. [Google Scholar] [CrossRef] [Scilit]
  49. Bjelobaba, S.; Waddington, L.; Perkins, M.; Foltýnek, T.; Bhattacharyya, S.; Weber-Wulff, D. Maintaining Research Integrity in the Age of GenAI: An Analysis of Ethical Challenges and Recommendations to Researchers. Int. J. Educ. Integr. 2025, 21, 18. [Google Scholar] [CrossRef] [Scilit]
  50. European Commission Living Guidelines on the Responsible Use of Generative AI in Research. Research and Innovation. Available online: https://research-and-innovation.ec.europa.eu/document/2b6cf7e5-36ac-41cb-aab5-0d32050143dc_en (accessed on 9 July 2026).
  51. Chelli, M.; Descamps, J.; Lavoué, V.; Trojani, C.; Azar, M.; Deckert, M.; Raynier, J.-L.; Clowez, G.; Boileau, P.; Ruetsch-Chelli, C. Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis. J. Med. Internet Res. 2024, 26, e53164. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  52. Akbar, M.N. Use of Artificial Intelligence Tools by Doctoral Students: A Mixed-Methods Explanatory-Sequential Investigation. J. Furth. High. Educ. 2025, 49, 995–1013. [Google Scholar] [CrossRef] [Scilit]
  53. Bearman, M.; Tai, J.; Henderson, M.; Esterhazy, R.; Mahoney, P.; Molloy, E. Enhancing Feedback Practices within PhD Supervision: A Qualitative Framework Synthesis of the Literature. Assess. Eval. High. Educ. 2024, 49, 634–650. [Google Scholar] [CrossRef] [Scilit]
  54. Madleňák, R.; Madleňáková, L.; Cvacho, V.; Gachulinec, D. Ethical Challenges of Artificial Intelligence in Higher Education: A Four-Pillar Student-Activity Framework for Institutional Governance. Educ. Sci. 2026, 16, 555. [Google Scholar] [CrossRef] [Scilit]
  55. Carless, D.; Jung, J.; Li, Y. Feedback as Socialization in Doctoral Education: Towards the Enactment of Authentic Feedback. Stud. High. Educ. 2024, 49, 534–545. [Google Scholar] [CrossRef] [Scilit]
  56. Bao, J.; Feng, D. Supervisory Feedback and Doctoral Students’ Academic Literacy Development: The Case of Writing for Publication. Teach. High. Educ. 2025, 30, 862–879. [Google Scholar] [CrossRef] [Scilit]
  57. Stracke, E.; Kumar, V. Encouraging Dialogue in Doctoral Supervision: The Development of the Feedback Expectation Tool. Int. J. Dr. Stud. 2020, 15, 265–284. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  58. Tensen, D.; Grainger, P.; Graham, W. Using AI to Generate Formative Feedback in Doctoral Education. Assess. Eval. High. Educ. 2026, 51, 476–492. [Google Scholar] [CrossRef] [Scilit]
  59. Khaw, L.L. Engineering PhD Students’ Perception and Experience with ChatGPT in Research Communication: Pedagogical Insights. Innov. Educ. Teach. Int. 2025, 63, 1484–1500. [Google Scholar] [CrossRef] [Scilit]
  60. Tensen, D.; Carey, M.D.; Grainger, P. Comparing ChatGPT-5 and Other GenAI Tools in Doctoral Confirmation: Variability, Persona Effects, and Alignment with Human Feedback. Assess. Eval. High. Educ. 2026, 1–17. [Google Scholar] [CrossRef] [Scilit]
  61. Nguyen, A.; Hong, Y.; Dang, B.; Huang, X. Human-AI Collaboration Patterns in AI-Assisted Academic Writing. Stud. High. Educ. 2024, 49, 847–864. [Google Scholar] [CrossRef] [Scilit]
  62. Wu, E. From Integrity to Identity: Course-Level Generative AI Governance and Scholarly Subject Formation in Graduate Education. Front. Educ. 2026, 11, 1827251. [Google Scholar] [CrossRef] [Scilit]
  63. Russell Group Principles on the Use of Generative AI Tools in Education. Russell Group. Available online: https://www.russellgroup.ac.uk/policy/policy-briefings/principles-use-generative-ai-tools-education (accessed on 9 July 2026).
  64. Parker, J.L.; Becker, K.P. Defining and Assessing AI Literacy for Researchers across the Research Lifecycle. Front. Educ. 2026, 11, 1827603. [Google Scholar] [CrossRef] [Scilit]
  65. English, R.; Nash, R.; Mackenzie, H. ‘A Rather Stupid but Always Available Brainstorming Partner’: Use and Understanding of Generative AI by UK Postgraduate Researchers. Innov. Educ. Teach. Int. 2026, 63, 193–207. [Google Scholar] [CrossRef] [Scilit]
  66. Xu, H.; Shen, W. From Tool to Partner: Generative AI Usage Patterns and Research Performance among Doctoral Students. Stud. High. Educ. 2026, 1–16. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Conceptual and integrative human-centred framework for generative AI in doctoral supervision. Four actors and five interdependent dimensions are situated within a recurrent cycle of expectation setting, evaluation, adjustment, and re-evaluation across doctoral progression.
Figure 1. Conceptual and integrative human-centred framework for generative AI in doctoral supervision. Four actors and five interdependent dimensions are situated within a recurrent cycle of expectation setting, evaluation, adjustment, and re-evaluation across doctoral progression.
Encyclopedia 06 00209 g001
Table 1. Positioning of the present review in relation to prior syntheses.
Table 1. Positioning of the present review in relation to prior syntheses.
Prior Synthesis Primary Focus Scope Relative to Doctoral Supervision Distinctive Positioning of the Present Review
Bittle and El-Gayar [14]; Balalle and Pannilage [15]AI/GenAI and academic integrity across higher educationPrimarily addresses student behaviour, misconduct, detection, and institutional integrity; doctoral supervision is not the principal unit of analysis.Extends integrity analysis to sustained supervisory relationships, authorship, disclosure, confidentiality, and distributed accountability.
Urzúa et al. [21]AI-assisted feedback and academic writing among higher education studentsCentres on writing outcomes and learner effects across university contexts rather than doctoral feedback ecologies or long-term supervisory relationships.Links AI-mediated feedback to feedback literacy, disciplinary judgement, doctoral autonomy, and relational mentoring.
Thong et al. [22]AI-powered tools and applications in doctoral supervisionMaps tools and applications in the doctoral context; relationships, responsibility, and governance are not used as the integrating analytical unit.Reframes generative AI as a socio-technical participant and integrates ethics, feedback, trust, power, agency, and governance.
Present reviewHuman-centred and socio-technical doctoral supervisionCritical and integrative synthesis spanning doctoral education, feedback, writing, ethics, relationships, and governance.Proposes an original four-actor, five-dimension framework centred on human judgement and scholarly responsibility.
Table 2. Methodological positioning of the present critical and integrative narrative review.
Table 2. Methodological positioning of the present critical and integrative narrative review.
Dimension Present Critical and Integrative Narrative Review Systematic Review Reported Using PRISMA
Primary purposeCritically integrate heterogeneous evidence to identify conceptual relationships, tensions, mechanisms, propositions, and evidence gaps.Answer a predefined review question through explicit identification and synthesis of eligible evidence.
Scope of the questionCan integrate interconnected conceptual, ethical, pedagogical, relational, and institutional dimensions.Usually defined more narrowly according to a specified review question and eligibility framework.
Eligible evidenceEmpirical studies, reviews, conceptual and methodological scholarship, and authoritative normative or policy sources, with different evidential roles.Evidence types are specified according to the review question and predefined eligibility criteria.
Search and selectionStructured and transparent searches support a conceptually bounded synthesis; foundational and authoritative sources may also be retained when analytically necessary.Explicit and reproducible procedures are used to identify and select eligible evidence within the predefined scope.
Evidence appraisalInterpretive weight is calibrated according to evidence type, methodological coherence, relevance, context, and contribution to the synthesis; heterogeneous sources are not treated as interchangeable.Critical appraisal or risk-of-bias procedures are selected according to the included study designs and purpose of the review.
SynthesisCritical, thematic, comparative, and conceptual integration across heterogeneous evidence types.Structured synthesis of included evidence, which may be qualitative or quantitative depending on the question and evidence base.
Typical outputConceptual relationships, explanatory interpretation, propositions, framework development, and identification of research gaps.A synthesised answer to the predefined review question, bounded by the characteristics and limitations of the included evidence.
Role of PRISMANot adopted as the reporting framework for the present narrative-integrative review.PRISMA provides reporting guidance for systematic reviews; it is not itself a review methodology.
Note. The comparison clarifies the methodological positioning of the present review rather than establishing mutually exclusive categories. Review approaches may share procedures such as structured database searching, predefined eligibility criteria, transparent extraction, and thematic synthesis; the distinction concerns the overall purpose, evidential scope, and form of synthesis.
Table 3. Evidential roles and calibrated formulations used in the synthesis.
Table 3. Evidential roles and calibrated formulations used in the synthesis.
Source or Evidence Type How the Source Is Used in the Present Review
Empirical studiesUsed to describe observed practices, associations, participant experiences, or comparisons of outputs and outcomes within the contexts reported by the original studies. Claims are expressed using calibrated formulations such as “studies report,” “findings indicate,” or “participants described.” These sources are not treated as supporting universal causal claims beyond their reported designs and contexts.
Review articlesUsed to characterise broader patterns, areas of convergence or disagreement, and gaps across existing bodies of literature. Claims are expressed using formulations such as “reviews synthesise,” “reviews identify,” or “the reviewed evidence suggests.”
Conceptual, theoretical, and
design-science articles
Used to identify propositions, conceptual relationships, explanatory models, or proposed mechanisms. Claims are expressed using formulations such as “proposes,” “conceptualizes,” or “models,” rather than as empirically demonstrated effects.
Methodological sourcesUsed to inform the design, transparency, quality alignment, and reporting of the review. These sources support methodological decisions rather than substantive claims about the effectiveness or consequences of GenAI in doctoral supervision.
Institutional, policy, and
international guidance
Treated as normative evidence concerning recommended responsibilities, safeguards, governance arrangements, and good practice rather than as empirical evidence of effectiveness. Claims are expressed using formulations such as “guidelines recommend,” “guidance advises,” or “institutions recommend.”
Present critical and integrative reviewUsed to interpret relationships, tensions, responsibilities, implications, and evidence gaps across heterogeneous source types. Synthesised claims are expressed using formulations such as “this review interprets,” “the synthesis suggests,” or “the framework proposes.”
Table 4. Potential uses of generative AI in doctoral supervision.
Table 4. Potential uses of generative AI in doctoral supervision.
Area of Supervision AI-Supported Use Potential Benefit Main Risk Supervisory Safeguard
Doctoral writingEditing, restructuring, language refinement, translationImproved clarity and support for multilingual writingLoss of scholarly voice or superficial fluencyRequire students to explain accepted and rejected AI suggestions
Literature engagementSearch terms, summaries, conceptual mapping, discussion preparationBetter preparation for literature conversationsInaccurate summaries, fabricated references, epistemic narrowingVerify claims against peer-reviewed sources
Feedback practicesPreliminary comments, self-review, clarification of supervisor feedbackFaster formative support and better meeting preparationGeneric or misleading feedbackSupervisor validates and contextualises feedback
Research planningRefining research questions, timelines, agendas, methodological promptsMore focused supervision meetingsMethodological oversimplificationSupervisor retains final methodological judgement
Administrative coordinationMeeting notes, reminders, email drafts, milestone planningReduced routine workloadConfidentiality or data governance risksUse only non-sensitive information and approved tools
Table 5. Ethical boundaries in AI-augmented doctoral supervision.
Table 5. Ethical boundaries in AI-augmented doctoral supervision.
Ethical Issue Example in Doctoral Supervision Principal Risk Recommended Practice
Academic integrityAI-generated text submitted as the student’s own reasoningMisrepresentation of intellectual workDistinguish support from outsourcing
AuthorshipAI substantially shapes thesis sections or manuscriptsAmbiguous contribution and accountabilityHuman authors remain fully responsible
OriginalityAI generates arguments, interpretations, or literature synthesisWeakening of independent contributionUse AI to support reflection, not replace analysis
DisclosureHidden AI-assisted writing or feedback interpretationLoss of trust and unclear assessment of competenceDeclare substantive AI use
Privacy and
confidentiality
Uploading unpublished chapters, transcripts, or sensitive dataData protection and research ethics breachesAvoid unauthorised tools for confidential material
Bias and hallucinationAI-generated summaries or invented referencesMisleading evidence and epistemic distortionVerify all AI outputs against reliable sources
OverdependenceAI revises every draft or interprets every commentDeskilling in reading, writing, and judgementUse AI as scaffold, not authority
Table 6. Recommendations for responsible AI-augmented doctoral supervision.
Table 6. Recommendations for responsible AI-augmented doctoral supervision.
Responsible Actor(s) Operational Recommendation Implementation or Evidence of Enactment
Doctoral researchersVerify claims, quotations, and references.Check outputs against original and preferably primary sources before use.
Doctoral researchersMaintain a proportionate record.Retain prompts, outputs, major revisions, and reasons for important decisions when intellectual or data risk warrants it.
Doctoral researchersDisclose substantive AI use.State the tool, purpose, scope, and nature of assistance when required by institutional, thesis, research ethics, or journal policy.
Doctoral researchersDistinguish language support from intellectual contribution.Use AI for expression or organisation without outsourcing analysis, interpretation, or judgement.
Doctoral researchersRemain able to defend every decision.Demonstrate independent understanding of the sources, methods, analysis, and conclusions.
Doctoral researchersComplete proportionate milestone self-evaluation of substantive AI use.Document the tool and purpose, how it contributed to the work, what outputs were accepted or rejected, what was independently verified, remaining uncertainties, and areas in which additional guidance or development is required.
Supervisors/supervisory teamsSet expectations at candidature start.Record permitted, conditional, and prohibited uses in the supervisory agreement; review them at relevant milestones.
Supervisors/supervisory teamsReview expectations as the project changes.Revisit rules when methods, data sensitivity, publication plans, research stages, or tools change.
Supervisors/supervisory teamsExamine the candidate’s reasoning.Ask how AI was used, which outputs were accepted or rejected, and what was independently verified.
Supervisors/supervisory teamsProtect unpublished and confidential material.Use only authorised systems and never upload identifiable or sensitive material without appropriate approval.
Supervisors/supervisory teamsDisclose substantive AI use in supervisory work.State the tool, purpose, scope, and nature of assistance when AI materially shapes feedback, examples, summaries, or other supervisory materials.
Supervisors/supervisory teamsDocument supervisory evaluation of AI-supported doctoral practice.Assess the candidate’s understanding, capacity to explain and defend decisions, adequacy of disclosure and verification, relevant integrity or data risks, and areas requiring further development.
Doctoral researchers and supervisors/supervisory teamsAgree and document improvement actions.Record identified strengths and weaknesses, agreed actions, required support or training, responsibilities for implementation, and the milestone at which progress will be re-evaluated.
Institutions/doctoral schoolsAdopt, formally authorise, and maintain a doctoral-specific AI policy.Address writing, literature work, methods, data, feedback, supervision, publication, and research ethics; document approval by the responsible senior academic or research authority; identify policy ownership; and publish the current version, review date, and routes for interpretation, advice, or appeal.
Institutions/doctoral schoolsClassify AI uses by task and risk.Define permitted, conditional, and prohibited uses according to intellectual, ethical, and data risk.
Institutions/doctoral schoolsMaintain operational quality-assurance resources.Provide approved-system registers, risk-based guidance, discipline-sensitive examples, disclosure and record templates, data-protection guidance, and accessible reference materials.
Institutions/doctoral schoolsProvide recurring AI-literacy training.Cover bias, hallucination, disclosure, source verification, confidentiality, data protection, responsible prompting, and approved tools using discipline-relevant cases.
Institutions/doctoral schoolsIntegrate AI into research governance.Add proportionate AI questions to ethics applications, data-management plans, supervisory agreements, milestone reviews, and other relevant research-governance processes.
Institutions/doctoral schoolsEnsure equitable access to approved support.Provide licences where justified, accessible training, disability support, and appropriate alternatives for candidates without equivalent resources or who cannot use particular systems.
Institutions/doctoral schoolsProvide targeted developmental or corrective support.Use refresher training, case-based workshops, revised guidance, specialist consultation, or referral to ethics, data-governance, information-security, or disciplinary expertise according to the identified need.
Institutions/doctoral schoolsReview aggregated quality-assurance evidence.Use appropriately anonymised patterns from milestone evaluations, recurrent uncertainties, training needs, equity concerns, ethics cases, and support requests to revise policy, resources, training, and approved-system requirements.
Table 7. Future research agenda for generative AI in doctoral supervision.
Table 7. Future research agenda for generative AI in doctoral supervision.
Research Area Key Question Suggested Methodology Relevance for Doctoral Supervision
Actual AI use and driversHow, why, and under what conditions do doctoral researchers and supervisors use, avoid, disclose, or conceal GenAI in everyday doctoral work?Interviews, surveys, diary studies, prompt and decision logs, supervision observations, and mixed-methods designs.Identifies actual practices and the individual, relational, disciplinary, technological, and institutional factors that enable, constrain, or redirect GenAI use.
AI-mediated feedbackHow does AI-generated feedback compare with supervisor feedback?Comparative feedback analysis, student response studies.Supports feedback literacy and writing development.
Ethical boundariesHow are disclosure, authorship, originality, and confidentiality negotiated?Policy analysis, case studies, ethics review analysis.Clarifies acceptable and unacceptable AI assistance.
Benefits, harms, and developmental trade-offsUnder what conditions does sustained GenAI use support doctoral learning, access, feedback, and autonomy, and under what conditions does it contribute to dependency, weakened judgement, relational displacement, or other adverse effects?Longitudinal and mixed-methods studies, repeated milestone assessments, process tracing, comparative cohort studies, and oral or performance-based evaluation.Clarifies for whom, when, and under what conditions GenAI functions as productive scaffolding rather than as substitution for independent scholarly development.
Trust, relational dynamics, and social capitalHow does GenAI use influence trust, reciprocal support, scholarly belonging, and access to academic relationships, networks, and informational resources within doctoral supervision?Longitudinal qualitative studies, social-network approaches, interviews, supervision observations, and mixed-methods relational designs.Clarifies whether AI-mediated practices expand or constrain the relational resources and human connections that support doctoral development, academic socialisation, and effective supervision.
EquityWho benefits from AI-supported supervision?Comparative institutional studies.Identifies access gaps and unequal digital competencies.
Supervisor trainingWhat training helps supervisors guide AI use responsibly?Programme evaluation, mixed-methods studies.Supports AI literacy and responsible supervision.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Guanilo Pareja, C.G.; Pareja Pera, L.Y.; Arias Lizares, A.; Huanca Rojas, L.M.; López Gómez, H.E.; Díaz Tavera, Z.R.; Almonte Andrade, C.P.; Ramirez Wong, F.M.; Román-Santana, W.M.; Mata-de-Salcedo, C.; et al. Generative Artificial Intelligence in Doctoral Supervision: Ethical Boundaries, Feedback Practices, and Supervisory Relationships. Encyclopedia 2026, 6, 209. https://doi.org/10.3390/encyclopedia6100209

AMA Style

Guanilo Pareja CG, Pareja Pera LY, Arias Lizares A, Huanca Rojas LM, López Gómez HE, Díaz Tavera ZR, Almonte Andrade CP, Ramirez Wong FM, Román-Santana WM, Mata-de-Salcedo C, et al. Generative Artificial Intelligence in Doctoral Supervision: Ethical Boundaries, Feedback Practices, and Supervisory Relationships. Encyclopedia. 2026; 6(10):209. https://doi.org/10.3390/encyclopedia6100209

Chicago/Turabian Style

Guanilo Pareja, Carla Giuliana, Lidia Ysabel Pareja Pera, Andrés Arias Lizares, Lupe Marilu Huanca Rojas, Henri Emmanuel López Gómez, Zoila Rosa Díaz Tavera, Clara Patricia Almonte Andrade, Fernando Martin Ramirez Wong, Wanda Marina Román-Santana, Carmen Mata-de-Salcedo, and et al. 2026. "Generative Artificial Intelligence in Doctoral Supervision: Ethical Boundaries, Feedback Practices, and Supervisory Relationships" Encyclopedia 6, no. 10: 209. https://doi.org/10.3390/encyclopedia6100209

APA Style

Guanilo Pareja, C. G., Pareja Pera, L. Y., Arias Lizares, A., Huanca Rojas, L. M., López Gómez, H. E., Díaz Tavera, Z. R., Almonte Andrade, C. P., Ramirez Wong, F. M., Román-Santana, W. M., Mata-de-Salcedo, C., & Martínez-Alonzo, J. (2026). Generative Artificial Intelligence in Doctoral Supervision: Ethical Boundaries, Feedback Practices, and Supervisory Relationships. Encyclopedia, 6(10), 209. https://doi.org/10.3390/encyclopedia6100209

Article Metrics

Back to TopTop