Beyond Agentic Translation: A Design-Based Framework for Human-AI Collaboration in Translator Education
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsThank you for the opportunity to review this manuscript. I found the topic timely and the four-dimensional framework to be the strongest part of the paper. The connection between cross-disciplinary knowledge, AI orchestration, strategic judgement, and critical AI evaluation gives the work a solid foundation. I also appreciated that the authors are clear that this is a protocol paper and that the empirical work is still forthcoming.
I do think a few areas need attention before publication. First, some of the claims about what AI can or cannot do feel stronger than the evidence supports. The paper is actually more nuanced in other places, so I would encourage that same caution throughout, especially in the displacement inventory. It would also help to make clearer which ratings are supported by existing evidence and which are informed projections that still need validation.
The co-design process could also be clarified. If the four dimensions are already fixed and participants are only helping shape the curriculum, assessments, and implementation, then it may be more accurate to describe that as co-design of the operationalization rather than co-design of the framework itself.
I would also reconsider the use of the term “representative sample” for five pilot participants. For a small proof-of-concept study, purposeful selection may be a better fit. Along the same lines, the hypotheses may be better framed as pre-registered benchmarks or evaluation criteria, since the authors themselves acknowledge that the sample is too small for meaningful inferential claims.
One of the more important revisions is keeping the protocol language consistent throughout. At times, the Discussion and Conclusion sound as if pilot evidence already exists, even though the study has not yet been conducted. I would suggest clearly separating what is already supported by the literature, what the framework proposes, and what the future pilot is intended to test.
I also think the “human” side of human-AI collaboration could be developed a little more. The framework does a strong job describing what people need to know and do, but there is less attention to professional identity, presence, trust, responsibility, agency, and the relational side of working with AI. Even a brief discussion of where those ideas fit would strengthen the paper.
Finally, I would soften some of the language around reliability and validity with an N of 5, tighten some repetition, and complete a careful proofread. For example, the Limitations section says there are four limitations but then presents five.
Overall, I see a lot of potential here. I do not think the study needs to be redesigned. The main need is to make the claims more cautious and clearly distinguish between what is established, what is proposed, and what still needs to be tested.
Author Response
Manuscript ID: education-4573470
Title: Beyond Agentic Translation: A Design-Based Framework for Human-AI Collaboration in Translator Education
Dear Reviewer 1,
Thank you for your thorough and constructive review. Your seven areas of concern have been addressed in the revised manuscript, and we provide a point-by-point response below. Substantive changes in the revised manuscript (manuscript.v5.docx) are flagged in the response below by section number.
Reviewer 1, Point 1
"Some of the claims about what AI can or cannot do feel stronger than the evidence supports... It would also help to make clearer which ratings are supported by existing evidence and which are informed projections that still need validation."
Response. We agree and have revised Table 2 (The Displacement Inventory) so that every displacement-likelihood cell is annotated [E] (empirically supported) or [P] (informed projection). Of the 18 sub-tasks, 6 are rated [E] and 12 are rated [P]. The §2.6 introductory paragraph now states that "Ratings marked [P] should be read as explicitly provisional working hypotheses pending triangulation through the Phase 1.5 mini-Delphi survey, particularly for sub-tasks where published empirical evidence is sparse." The Phase 1.5 mini-Delphi with 10–15 practising translators and LSP managers is now specified in §3.3 as the validation step for the [P]-rated sub-tasks.
Reviewer 1, Point 2
"The co-design process could also be clarified. If the four dimensions are already fixed and participants are only helping shape the curriculum, assessments, and implementation, then it may be more accurate to describe that as co-design of the operationalization rather than co-design of the framework itself."
Response. Agreed. The distinction is now stated in three places: the Abstract inserts the parenthetical "(whose operationalisation is developed through a co-design process rather than the framework structure itself)"; §3.3 explicitly states that "The 4D framework's ontological structure (D1–D4 as dimensions of distributed-cognition human node capacity) is theoretically fixed; its operationalisation details … are subject to refinement, extension, and contextualisation by the 5-member co-design team during Phase 2"; and §5.1's first workshop output has been reworded from "a refined four-dimensional competence framework" to "a refined operationalisation of the four-dimensional competence framework".
Reviewer 1, Point 3
"I would also reconsider the use of the term 'representative sample' for five pilot participants. For a small proof-of-concept study, purposeful selection may be a better fit."
Response. Done. §3.2 now reads "a purposefully selected sample that captures variation in prior AI experience and year of study within the constraints of a small proof-of-concept design." The §3.4 Threats to Validity section explicitly frames the small N = 5 as a proof-of-concept with purposeful selection, not as a study with a representative sample.
Reviewer 1, Point 4
"The hypotheses may be better framed as pre-registered benchmarks or evaluation criteria, since the authors themselves acknowledge that the sample is too small for meaningful inferential claims."
Response. Done throughout. §5.3 has been retitled from "Pre-Registered Hypotheses" to "Pre-Registered Evaluation Criteria"; the three registered items are now numbered EC1, EC2, EC3 (rather than H1, H2, H3); the introductory paragraph states that the criteria function as benchmarks for proof-of-concept success rather than as confirmatory statistical tests; and the term "pre-registered evaluation criteria" (or "pre-registered benchmarks") has replaced "hypotheses" wherever the registered items are mentioned throughout the manuscript body, including §3.1, §3.3, §5.3, §6.2, the RQ3 framing, and the Supplementary Materials list.
Reviewer 1, Point 5
"Keeping the protocol language consistent throughout... Discussion and Conclusion sound as if pilot evidence already exists... clearly separating what is already supported by the literature, what the framework proposes, and what the future pilot is intended to test."
Response. A systematic audit has been conducted. Specific changes:
(a) §6.3: the sentence formerly reading "the pilot evidence provides preliminary information" has been changed to "the planned pilot is designed to provide preliminary information"; the industry-relevance claim has been rephrased from "correspond to" to "are hypothesised to correspond to … are hypothesised to be increasingly sought in new hires".
(b) §6.4: the fourth limitation has been reframed as a genuine constraint ("the framework's educational effectiveness remains untested and is contingent on the planned pilot's results").
(c) §7.2: the implications paragraph now opens with "each is offered as a hypothesis to be evaluated through the planned pilot and subsequent multi-site implementations rather than as an established finding".
(d) Abstract: the closing sentence clearly distinguishes the present paper's contributions (the framework, the methodology, the protocol) from the empirical evidence to be reported in a subsequent paper.
Reviewer 1, Point 6
"The 'human' side of human-AI collaboration could be developed a little more. Professional identity, presence, trust, responsibility, agency, and the relational side of working with AI."
Response. We have expanded on three of the requested dimensions in three locations:
(a) §2.2 (Agency): a paragraph now addresses calibrated trust explicitly, drawing on the trust-in-automation literature [49, 50] and locating trust calibration as a competence supported jointly by D3 and D4, rather than as a disposition.
(b) §6.1 (Professional identity): a paragraph on translator professional identity has been added, arguing that identity shifts from "language transfer specialist" to "human-AI collaboration professional" and that identity formation should be designed around situated decision-making rather than tool familiarity.
(c) §2.2 (Responsibility / power): the existing paragraph on Leonardi's (2025) [28] framework of agency and power distribution has been retained; it directly addresses the responsibility and agency dimensions of human-AI collaboration.
We acknowledge that a comprehensive treatment of presence and other affective dimensions of human-AI teaming remains a direction for future research, as stated in the response letter.
Reviewer 1, Point 7
"Soften some of the language around reliability and validity with an N of 5, tighten some repetition, and complete a careful proofread. For example, the Limitations section says there are four limitations but then presents five."
Response. Three specific actions have been taken:
(a) §4.5 reliability/validity language has been softened: all such references now include the qualifier "preliminary" and the N = 5 constraint is explicit ("The pilot generates preliminary reliability evidence (inter-rater agreement, internal consistency) and preliminary content and construct validity evidence … subject to confirmation in larger subsequent implementations").
(b) §6.4 Limitations count has been corrected: the section now opens "Five limitations of the present design should be acknowledged" and lists five numbered items, including a fifth limitation on the custom-platform / author-developer conflict.
(c) Proofreading: a complete proofread has been conducted, including standardisation of hyphenation, correction of grammatical issues, and cross-checking of citation–reference correspondence.
Summary
We thank Reviewer 1 for the constructive feedback. The seven areas of concern have been addressed through the systematic revisions described above. The revised manuscript (manuscript.v5.docx) is now well-calibrated in its claims, has clearer co-design scope, uses appropriate language for the small pilot sample, maintains protocol-stage language consistency, gives meaningful treatment to the human side of human-AI collaboration, and has been carefully proofread.
Sincerely,
Dr. John Qiong WANG
Guangxi Minzu University
john.wangqiong@hotmail.com
Reviewer 2 Report
Comments and Suggestions for AuthorsI would like to start by acknowledging the relevance and timeliness of this manuscript. The discussion about how AI is reshaping professional education, particularly translator education, is highly relevant, and the effort to propose a framework for human-AI collaboration represents an interesting contribution.
The manuscript is well organized and demonstrates a strong effort to connect AI, professional competencies, curriculum redesign, and design-based research. The proposed framework has potential, especially because it moves beyond a simple discussion of technological adoption and focuses on how human competencies may be reconfigured in AI-mediated professional contexts.
However, some aspects would benefit from further refinement.
First, I believe the manuscript needs to establish a clearer distinction between what is already occurring in translation practice, what represents an emerging tendency associated with AI development, and what remains a future-oriented hypothesis. The authors acknowledge that adoption of agentic translation is uneven across contexts, but this distinction should be maintained more consistently throughout the manuscript.
Second, the contribution of the proposed framework could be discussed more explicitly. The manuscript presents four relevant dimensions, but it would benefit from a clearer explanation of what new analytical perspective this framework provides and how it advances current discussions on professional competencies, AI literacy, and curriculum transformation.
Third, since this is a design protocol paper, some conclusions should be presented with greater caution. The manuscript provides a strong rationale for the proposed framework, but its educational value and broader applicability will depend on the empirical evidence generated in future phases of the research.
Finally, I appreciate the transparency regarding the author's involvement in developing a custom AI platform used in the planned pilot. This is an important aspect that has been appropriately acknowledged. Nevertheless, the procedures for maintaining independence in evaluation and minimizing possible bias should remain clearly described throughout the research process.
Overall, I consider this a promising manuscript with relevant potential contribution. The main revisions needed are related to strengthening the theoretical positioning of the framework, maintaining a careful distinction between current evidence and future expectations, and ensuring that claims remain aligned with the empirical stage of the research.
Comments on the Quality of English LanguageThe manuscript is generally understandable and presents a consistent academic style. However, some sentences are very long and contain multiple conceptual layers, which occasionally affects clarity and readability. A careful language revision would help improve precision, reduce repetition, and make the argument easier to follow for an international audience.
Author Response
Manuscript ID: education-4573470
Title: Beyond Agentic Translation: A Design-Based Framework for Human-AI Collaboration in Translator Education
Dear Reviewer 2,
Thank you for your thoughtful review and for recognising the relevance of the manuscript's contribution to discussions of how AI is reshaping professional education. Your four areas of refinement have been addressed in the revised manuscript (manuscript.v5.docx). A point-by-point response is provided below.
Reviewer 2, Point 1
"The manuscript needs to establish a clearer distinction between what is already occurring in translation practice, what represents an emerging tendency associated with AI development, and what remains a future-oriented hypothesis. The authors acknowledge that adoption of agentic translation is uneven across contexts, but this distinction should be maintained more consistently throughout the manuscript."
Response. Agreed. The three-register distinction (currently-occurring, emerging, future-oriented) has been made explicit and is now maintained in five locations:
(a) §1.1 introduces the temporal distinction directly: "It is important to acknowledge … that the paradigm shift is uneven across language pairs, domains, and professional settings. Many working translators in 2026 still do substantial post-editing; the 'orchestrator' role is emerging but not yet universal."
(b) §2.1 reiterates the same point at the paradigm level: "These paradigms coexist in current practice rather than superseding one another sequentially; the rate of agentic translation adoption varies across the profession … while others continue to operate primarily with NMT or even CAT workflows."
(c) §2.6 operationalises the distinction at the displacement-inventory level: Table 2 annotates each cell as [E] (empirically supported by published evidence) or [P] (informed projection), and the introductory paragraph marks [P] ratings as "explicitly provisional working hypotheses pending triangulation through the Phase 1.5 mini-Delphi survey".
(d) §6.1 and §7.2 frame future-oriented claims as hypotheses: "each is offered as a hypothesis to be evaluated through the planned pilot and subsequent multi-site implementations rather than as an established finding".
(e) The Abstract closing sentence separates present contributions (framework, methodology, protocol) from future empirical evidence ("to be reported in a subsequent paper following the pilot's completion").
The current-vs-emerging-vs-future distinction is therefore now operative in the introduction, the theoretical framework, the displacement inventory, the discussion, the conclusion, and the abstract.
Reviewer 2, Point 2
"The contribution of the proposed framework could be discussed more explicitly. The manuscript presents four relevant dimensions, but it would benefit from a clearer explanation of what new analytical perspective this framework provides and how it advances current discussions on professional competencies, AI literacy, and curriculum transformation."
Response. Agreed. The framework's contribution is now stated explicitly in three locations, each pitched to a different readership concern:
(a) §2.5 opens with the framework's primary epistemic contribution: "The framework's primary epistemic contribution is the shift in the unit of analysis from the individual translator to the human node in a distributed cognitive system; the four integration points below operationalise this shift." It then identifies three specific challenges prior work has not addressed: (i) the four dimensions are not typically co-trained in existing translator programmes; (ii) the integration bridges a gap in the EMT 2022 competence framework by providing operationalisations for the agentic translation era (operationalised in the new Table 1b); and (iii) the framework addresses the epistemic dimension by shifting the unit of analysis to the human node in a distributed cognitive system.
(b) §2.4 includes Table 1b, which explicitly maps each 4D dimension to the corresponding EMT 2022 sub-competence and specifies the "delta" (what the 4D framework adds to EMT 2022). The table makes the framework's additive contribution to the existing EMT competence model visible to the translation-studies reader.
(c) §6.1 reinforces the analytical-perspective contribution theoretically by anchoring it in distributed-cognition (Hutchins, 1995; Hollan, Hutchins, & Kirsh, 2000), which has not previously been applied to translator competence in the agentic translation era.
Reviewer 2, Point 3
"Since this is a design protocol paper, some conclusions should be presented with greater caution. The manuscript provides a strong rationale for the proposed framework, but its educational value and broader applicability will depend on the empirical evidence generated in future phases of the research."
Response. Agreed. Caution has been systematically introduced at the three points where the protocol-vs-finding distinction was at risk of blurring:
(a) §6.3 (Practical Contribution): the formerly assertive industry-relevance claim has been softened to "The 4D competencies specified in the framework are hypothesised to correspond to the competencies that localisation service providers and in-house translation departments are increasingly seeking … validation of this hypothesis through job-posting analysis or industry survey is a priority for future research."
(b) §6.4 (Limitations): the fourth limitation has been reframed as a genuine constraint ("Because the present study reports no empirical data, the framework's educational effectiveness remains untested and is contingent on the planned pilot's results; the framework is therefore presented as a design foundation whose empirical warrant will be assessed through the pre-registered evaluation criteria and subsequent multi-site implementations"), not as a description of paper scope.
(c) §7.2 (Implications): each implication is now framed as a hypothesis to be evaluated ("each is offered as a hypothesis to be evaluated through the planned pilot and subsequent multi-site implementations rather than as an established finding").
(d) The Abstract closing sentence explicitly states the boundary: "The empirical evidence on the framework's feasibility, acceptability, and educational effects will be reported in a subsequent paper following the pilot's completion; the present paper's contributions are the framework, the reproducible co-design methodology, and the transferable design protocol."
Reviewer 2, Point 4
"I appreciate the transparency regarding the author's involvement in developing a custom AI platform used in the planned pilot. This is an important aspect that has been appropriately acknowledged. Nevertheless, the procedures for maintaining independence in evaluation and minimizing possible bias should remain clearly described throughout the research process."
Response. The conflict-of-interest disclosure is retained, and the bias-mitigation procedures have been strengthened and made explicit in three locations:
(a) §3.4 (Threats to Validity): the author-researcher overlap is addressed by three named procedural safeguards: (i) raters are blinded to which agent platform (custom or commercial) the student used for each workflow; (ii) inter-tool reliability will be reported by comparing rubric scores for workflows executed on the custom platform versus commercial platforms; (iii) an external reviewer, unaffiliated with the platform development, will audit a random 20% sample of D2 assessments to confirm that scoring is not systematically influenced by the choice of tool.
(b) §4.2 (D2: AI Orchestration Competence): the custom platform's architecture is described in detail (planner module / orchestrator module / specialised agents / memory module / verification module) to enable reviewers and readers to assess its capabilities independently. The custom platform is positioned as one component of the toolset used in the pilot alongside commercial platforms (GPT-4, Claude, Gemini, DeepL); the platform is not the object of evaluation; the D2 assessment measures learner capability, not tool performance; learners' use of the custom platform is treated identically to their use of commercial platforms in assessment scoring.
(c) §4.5 (Assessment Approach): the independent rating procedures are specified ("Assessments for D2, D3, and D4 will be independently evaluated by two trained raters who are not members of the co-design team. Inter-rater agreement will be measured using weighted Cohen's Kappa, with a target of κ ≥ 0.75 across all dimensions.").
(d) The Conflict of Interest statement at the end of the manuscript has been expanded to incorporate the three procedural safeguards and to note that the AI college colleague on the co-design team provides an independent technical perspective.
Summary
We thank Reviewer 2 for the constructive feedback. The four areas of refinement have been addressed through systematic revisions to the distinction between current, emerging, and future-oriented claims; to the explicitness of the framework's analytical contribution; to the cautionary framing of conclusions; and to the explicitness of bias-mitigation procedures. The revised manuscript (manuscript.v5.docx) is now stronger in its theoretical positioning, more careful in its claims, and more transparent about the procedures for minimising potential bias.
Sincerely,
The Author
Round 2
Reviewer 2 Report
Comments and Suggestions for AuthorsThe authors have satisfactorily addressed the points raised in my previous review. I consider my previous concerns adequately addressed and recommend acceptance.

