Intermediate Representations in Human–AI Creative Collaboration: A Framework for Distributed Creative Cognition
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsThis paper presents a potentially valuable argument concerning the role of sketching in human–AI creative collaboration. Its main strength is in reconceptualising sketches as evolving intermediate representations rather than merely preliminary drawings. The integration of embodied cognition, distributed cognition, distributed intelligence, and boundary-object theory provides a promising basis for the proposed framework. The two figures are also useful in making the conceptual model accessible.
The paper’s main weakness, however, is the unresolved relationship between its conceptual and empirical dimensions. At several points, the paper appropriately describes itself as an exploratory, hypothesis-generating conceptual contribution rather than an experimental demonstration. Nevertheless, it is structured as an empirical article, with participants, data sources, an analytical approach, Results, and claims about recurring or consistent patterns. The evidence and methodological detail currently provided are not sufficient to support that empirical presentation.
As such, the author(s) should decide to either: 1) develop the article as an empirical design-based research study, which would require substantially expanding the methodology and results, including clearly formulated research questions; details of the six course iterations; participant and assignment breakdowns; ethical procedures; the size and composition of the dataset; analytical stages; theme-development procedures; reflexivity; and evidence supporting each reported theme; or 2) reframe the article as a conceptual or reflective-practice paper, in which case the classroom observations would be presented as illustrative cases that motivated the framework, rather than as empirical findings. The present Results section could then be titled as “Classroom Observations” or “Illustrative Patterns.”
The theoretical framework is convincing, but the paper should separate observation from interpretation more consistently. For example, claims concerning student ownership, enriched understanding, and the cognitive effects of deliberate slowing are plausible, but they require either direct participant evidence or more cautious language.
The literature review also needs further development. The foundational sources are relevant, but a contemporary article on generative AI cannot rely almost exclusively on theories and studies published before the emergence of current generative systems. Recent research should be used not merely to establish that AI is educationally important, but to position the specific contribution of the sketch-first framework in relation to existing work on human–AI co-creativity, creative agency, AI-supported design processes, prompt engineering, reflective learning, and multimodal interaction.
The limitations section is mainly fine but relevant limitations should also inform the earlier discussion. These include the single-course setting, the instructor’s dual role as designer and researcher, the absence of a clearly defined comparison group, possible selection effects associated with an elective course, disciplinary differences among participants, and variation in the more than 30 AI systems reportedly used.
Comments on the Quality of English LanguageThe paper is generally readable and professionally written, but it should undergo language editing for grammatical accuracy, sentence structure, concision, and consistency. There are several awkward or incorrect formulations, repeated statements, agreement errors, and minor typographical problems. Although these do not prevent comprehension, they weaken the presentation.
Author Response
Comment 1: The paper’s main weakness is the unresolved relationship between its conceptual and empirical dimensions.
Response 1:
I agree with this comment and have substantially revised the manuscript to clarify the relationship between its conceptual and empirical dimensions. The revised paper is now explicitly framed as a conceptual paper, with the classroom experiences serving as the motivating context for development of the framework rather than as empirical evidence of causal or comparative effects. This distinction is stated in the Abstract (p. 1, lines 16–27) and made explicit in the Introduction, where the manuscript now states that the classroom experiences are “not presented as a controlled comparison or as evidence that sketching causes particular learning outcomes,” but instead motivate the conceptual question developed in the paper (p. 2, lines 60-68).
The revised manuscript also separates the reflective classroom context from the claims of the framework. Section 3 describes the classroom experiences that prompted the conceptualization while explicitly acknowledging variation and counterexamples (pp. 5-6, lines 230–244), and Figure 1 is identified as an “interpretive synthesis” rather than an experimentally validated process model (p. 6, lines 245–255).
Finally, claims requiring empirical validation have been reframed as propositions for future research rather than findings of the present paper. Section 5 identifies five testable propositions and describes possible comparative studies, including the need to seek disconfirming cases (pp. 8–9, lines 314–342). The limitations now explicitly state that the framework has not been validated through controlled comparison and that the classroom experiences should not be generalized beyond their context without further study (p. 9, lines 350–358)
Comment 2: The literature review also needs further development.
Response 2:
I agree with this comment and have substantially expanded the literature review to include recent considerations. The revised manuscript retains the foundational literature on embodied cognition, distributed cognition, distributed intelligence, sketching, and boundary objects (pp. 2–3, lines 83–119), but now places these perspectives in dialogue with recent research on human–AI co-creativity, AI literacy and prompt engineering, multimodal design interaction, the timing of AI participation, and creative agency. A new section, 2.5 Contemporary human–AI co-creativity with intermediate representation, has been added specifically to develop this synthesis (pp. 3–5, lines 120–221).
The expanded review draws on recent work including Rezwana and Maher (2023) on interaction structures in human–AI co-creativity; Darvishi et al. (2024) on AI assistance and student agency; Knoth et al. (2024) and Oppenlaender et al. (2023) on AI literacy and prompt engineering; Paananen et al. (2024), Edwards et al. (2024), and Baudoux et al. (2025) on generative AI and multimodal interaction in design; Yang et al. (2026) on the timing of AI participation; and Guo et al. (2025) and Rafner et al. (2025) on agency during human–AI collaboration.
Importantly, these sources are used not simply to broaden the reference base but to critically position the proposed framework. The revised discussion distinguishes prompt competence from the cognitive formation of intention before prompting (p. 4, lines 150–158), distinguishes sketches as multimodal inputs from sketches as cognitive intermediate representations (pp. 4–5, lines 179–186), and uses recent empirical findings to examine the timing of AI participation and the continuing negotiation of human agency (p. 5, lines 187–215).
Comment 3: The limitations section is mainly fine but relevant limitations should also inform the earlier discussion.
Response 3:
I agree with this comment and have revised the manuscript so that relevant limitations and qualifications are incorporated into the earlier conceptual discussion rather than appearing only in the limitations section. In the Introduction, the classroom experiences are explicitly identified as motivating context rather than a controlled comparison or evidence that sketching causes particular learning outcomes (p. 2, lines 60–68).
The revised discussion of classroom practice also incorporates variation and counterexamples directly into development of the framework. Section 3 now notes that the observed experiences did not occur uniformly: students could abandon an initial sketch, prefer an AI interpretation, or find that an early representation constrained rather than expanded exploration. The manuscript therefore explicitly states that the framework does not assume that sketching always improves collaboration (p. 6, lines 238–244). Figure 1 is also identified as an interpretive synthesis rather than a prescribed or experimentally validated process model (p. 6, lines 245–255).
Limitations are also incorporated into the theoretical claims themselves. The discussion of human agency states explicitly that the proposed relationship between intermediate representations and agency is a mechanism requiring empirical investigation rather than a demonstrated effect (p. 8, lines 292–299). Similarly, the discussion of productive friction identifies its possible educational value while acknowledging that these possibilities require direct study (p. 8, lines 300–312).
Finally, the research agenda explicitly calls for future studies to seek disconfirming cases in which sketching adds little, constrains exploration, or is abandoned in favor of an AI-generated alternative, so that the boundary conditions of the framework can be established rather than simply confirmed (p. 9, lines 343–349).
Reviewer 2 Report
Comments and Suggestions for AuthorsThis design-based educational research makes a timely, theoretically rich contribution to art and design education amid generative AI integration. The paper advances a novel conceptual framework centred on sketching as evolving intermediate representations, challenging the dominant discourse that frames human–AI creative collaboration as starting with text prompting.
Here are some revision recommendations:
- The theoretical contextualisation would benefit from a brief critical synthesis section acknowledging counterarguments: some scholars argue generative AI renders manual sketching obsolete for professional visual production. Explicitly engaging with this opposing position would deepen the contextual framing and sharpen the paper’s central counterclaim. The link between boundary object theory and AI interpretive agents is novel but under-contextualised. Very few prior boundary object studies extend the framework to non-human computational actors; a short paragraph reviewing limited existing work on computational boundary objects would better position this paper’s theoretical extension;
- Formal, testable hypotheses are not explicitly enumerated as distinct text items. While tentative predictive claims appear in the Discussion and Future Research sections, consolidating 3–4 formal a priori hypotheses in the Methods chapter would align the paper with standard empirical research conventions and clarify the analytical focus of classroom observations;
- Data coding procedures lack specificity: the qualitative reflective analysis describes searching for recurring behavioural patterns but does not detail coding schema, inter-rater reliability protocols for instructor field notes and student reflections, or thematic saturation checks. Adding a short paragraph on qualitative coding standards will strengthen methodological rigourï¼›
- The argument for “productive friction” via sketching would benefit from deeper engagement with student reflective excerpts as supporting qualitative evidence. The paper references written reflections as a data source but embeds almost no direct student quotes within the Discussion to illustrate claims about slowed, reflective cognition. Integrating brief anonymised student reflection extracts will make interpretive arguments more persuasive;
- Results lack structured comparative contrast between sketch-first and implicit prompt-first student behaviours. The manuscript notes divergent approaches between the two cohorts of students but does not systematically tabulate or describe these contrasting behavioural styles side-by-side; a brief comparative summary table would clarify the core behavioural differences driving the paper’s central argumentï¼›
- If possible,try to add comparative analysis of sketch-first vs prompt-first student behavioural differences in the Results section, either via structured text summary or brief table.
Author Response
Comment 1: The theoretical contextualisation would benefit from a brief critical synthesis section acknowledging counterarguments.
Response 1:
I agree with this comment therefore I have developed a new subsection: 2.5 Contemporary human–AI co-creativity with intermediate representation. The section is not brief only because other reviewers suggested it should be a significant part of the paper once I focused more on a conceptual paper than an empirical one. I hopefully removed all wording that suggested the contribution was empirical. Section 2.5 appears on pages 3-5, lines 120-221.
The expanded review draws on recent work including Rezwana and Maher (2023) on interaction structures in human–AI co-creativity; Darvishi et al. (2024) on AI assistance and student agency; Knoth et al. (2024) and Oppenlaender et al. (2023) on AI literacy and prompt engineering; Paananen et al. (2024), Edwards et al. (2024), and Baudoux et al. (2025) on generative AI and multimodal interaction in design; Yang et al. (2026) on the timing of AI participation; and Guo et al. (2025) and Rafner et al. (2025) on agency during human–AI collaboration.
Importantly, these sources are used not simply to broaden the reference base but to critically position the proposed framework. The revised discussion distinguishes prompt competence from the cognitive formation of intention before prompting (p. 4, lines 150–158), distinguishes sketches as multimodal inputs from sketches as cognitive intermediate representations (pp. 4–5, lines 179–186), and uses recent empirical findings to examine the timing of AI participation and the continuing negotiation of human agency (p. 5, lines 187–215).
Comment 2: Formal, testable hypotheses are not explicitly enumerated as distinct text items.
Response 2:
I agree with this comment and have revised the manuscript to explicitly enumerate five testable propositions derived from the conceptual framework. These are presented as distinct items P1–P5 in Section 5 Educational Implications and Research Agenda (pp. 8–9, lines 326–342). The propositions address the effects of pre-AI externalization of intention, representation-first versus prompt-first workflows, ambiguity in intermediate representations, provenance and reflective judgment, and productive friction.
The manuscript also now describes how these propositions could be investigated empirically, including comparative sketch-first and prompt-first conditions, process traces, successive representations, reflective accounts, revision behavior, and measures of agency or ownership (p. 9, lines 343–349). The proposed research explicitly includes seeking disconfirming cases in order to establish the boundary conditions of the framework rather than simply confirm it.
Comment 3: Data coding procedures lack specificity
Response 3:
I agree that the previous version did not provide sufficient specificity to support interpretation of the classroom experiences as systematically coded empirical data. In revising the manuscript, however, I addressed this concern by clarifying the status of those experiences rather than by introducing a retrospective coding procedure. The revised manuscript is explicitly framed as a conceptual paper, and the classroom experiences are used as illustrative and motivating context rather than as a systematically analyzed empirical dataset. This distinction is stated in the Abstract (p. 1, lines 14–27) and Introduction (p. 2, lines 60–68), where the manuscript now specifies that the classroom experiences are not presented as a controlled comparison or as evidence that sketching causes particular learning outcomes.
Section 3 has consequently been reframed from an empirical results analysis to “From Classroom Practice to a Conceptual Framework.” It describes the reflective teaching experiences that motivated the framework while acknowledging that these experiences were variable and did not occur uniformly (pp. 5–6, lines 222–250). Figure 1 is similarly identified as an interpretive synthesis rather than an experimentally validated process model (p. 6, lines 245–255).
Rather than retrospectively imposing a coding procedure on classroom experiences that were not originally collected for that purpose, the revised paper identifies the framework's claims as propositions requiring prospective empirical testing. Section 5 therefore specifies possible future comparative designs and forms of evidence, including process traces, successive representations, reflective accounts, revision behavior, and measures of agency or ownership (p. 9, lines 343–349).
Comment 4: The argument for “productive friction” via sketching would benefit from deeper engagement with student reflective excerpts as supporting qualitative evidence.
Response 4:
I agree that the previous version did not provide sufficient qualitative evidence to support “productive friction” as an empirically demonstrated effect of sketching. In revising the manuscript, I addressed this concern by more clearly distinguishing the conceptual framework from the classroom experiences that motivated it, rather than by adding selected student excerpts as retrospective qualitative evidence. The revised manuscript therefore presents productive friction as a proposed educational mechanism requiring empirical investigation, rather than as a finding established by the classroom observations. In Section 4.5, the manuscript now states that the additional time introduced by sketching “may be educationally useful” and that the proposed effects on intention formation and provenance “require direct study” (p. 8, lines 300–306).
The revised discussion also considers a potential counterargument: increasingly capable multimodal AI systems may reduce the instrumental need for sketching. The framework therefore distinguishes communicative efficiency from cognitive value and proposes that constructing and revising an intermediate representation may remain educationally valuable even when the AI system does not require it as input (p. 8, lines 307–312).
Finally, productive friction is now explicitly formulated as a testable proposition (P5): “Productive friction introduced by representation-building activities may deepen observation and reflection even when it reduces the speed of artifact production” (pp. 8–9, lines 340–342). The subsequent research agenda identifies the kinds of prospective evidence that could be collected to test propositions of this kind, including process traces, successive representations, reflective accounts, and revision behavior (p. 9, lines 343–349).
Comment 5: Results lack structured comparative contrast between sketch-first and implicit prompt-first student behaviours
Response 5:
I agree that the previous version did not provide a sufficiently structured empirical basis for comparing sketch-first and prompt-first student behaviors. In revising the manuscript, I addressed this concern by removing the implication that the classroom experiences constitute a comparative empirical study. The revised manuscript is framed as a conceptual paper, with classroom experiences serving as the motivating context for the framework rather than as evidence of differences between experimental or systematically observed conditions. The Introduction now explicitly states that these experiences are “not presented as a controlled comparison or as evidence that sketching causes particular learning outcomes” (p. 2, lines 60–68).
Section 3 has accordingly been reframed as “From Classroom Practice to a Conceptual Framework” rather than as a Results section. The classroom discussion identifies recurring experiences associated with sketch-first activities while also noting that these experiences did not occur uniformly and that students could abandon sketches, prefer AI interpretations, or find that an initial representation constrained exploration (pp. 5–6, lines 222–250).
Rather than retrospectively constructing a sketch-first/prompt-first comparison from these classroom experiences, the revised manuscript now identifies such a comparison as an important subject for prospective empirical investigation. P1 proposes comparing how learners evaluate AI outputs when an emerging intention is externalized before AI interaction versus when they begin with prompting, while P2 proposes examining differences in patterns of iteration between representation-first and prompt-first workflows (p. 8, lines 326–333). The research agenda then explicitly proposes comparative studies of sketch-first and prompt-first conditions using process traces, successive representations, reflective accounts, revision behavior, and measures of agency or ownership (pp. 8–9, lines 343–349).
Comment 6: If possible,try to add comparative analysis of sketch-first vs prompt-first student behavioural differences in the Results section
Response 6:
I agree that a direct comparison of sketch-first and prompt-first student behaviors would provide an important empirical test of the proposed framework. However, the classroom activities that motivated the framework were not designed as controlled or systematically observed sketch-first and prompt-first conditions. I therefore chose not to construct a retrospective comparative analysis that the available classroom experiences could not adequately support. The revised manuscript instead explicitly identifies the classroom material as motivating context for a conceptual framework rather than as evidence of comparative effects (p. 2, lines 60–68).
Reviewer 3 Report
Comments and Suggestions for AuthorsThe manuscript presents an interesting argument: sketching can function as an intermediate cognitive representation that shapes creative collaboration with generative AI before ideas become prompts.
Nevertheless, the current version has a substantial mismatch between the strength of its empirical claims and the evidence and methods reported. The manuscript is presented as an empirical Article, with Materials and Methods and Results sections, but the analysis is described only as reflective observation by the instructor. The paper does not provide enough information to evaluate how the reported patterns were identified, whether alternative interpretations were considered, or how consistently the evidence supports the conclusions. Major revision is therefore necessary.
Major comments
-
Clarify whether this is an empirical study or a conceptual paper
The manuscript alternates between presenting itself as design-based educational research and describing itself as an interpretive framework derived from reflective teaching practice. These are not equivalent forms of inquiry.
If retained as an empirical Article, the paper needs explicit research questions, a documented research design, systematic data collection and analysis, empirical evidence in the Results section, and appropriate ethics statements.
Alternatively, the manuscript could be reframed as a conceptual or opinion paper informed by an illustrative course case. In that version, the classroom material should be presented as a motivating vignette rather than as evidence establishing recurrent differences between sketch-first and prompt-first approaches.
-
The comparative claims are not supported by the reported design
Several statements compare students who began by sketching with students who began primarily by prompting. The manuscript also states that sketch-first assignments produced “substantially more iterative behaviour” and that students maintained “stronger ownership.” However, the Methods section does not describe comparison groups, assignment allocation, baseline conditions, outcome definitions, or a procedure for comparing these behaviours.
It is especially unclear where the prompt-first evidence originated because the reported protocol states that most exercises intentionally required or encouraged sketching before AI interaction. The authors should either:
-document the comparison conditions, cohort structure, assignments, observations, and analytical procedure; or
-remove comparative and quasi-causal language and describe these points as tentative instructor interpretations.
Terms such as “consistently changed,” “substantially more,” and “stronger ownership” currently imply evidence that has not been presented.
-
The qualitative analytical procedure requires substantial development
Section 2.5 does not provide a reproducible or auditable qualitative method. The authors should explain:
-how field notes, reflections, discussions, sketches, and iterative works were selected and organised;
-the volume of material analysed;
-whether all 65 students and all six course offerings contributed data;
-whether coding was inductive, deductive, or both;
-how codes developed into the six reported patterns;
-whether data from different sources were triangulated;
-whether negative or contradictory cases were sought;
-whether another researcher reviewed the coding or interpretations;
-how the instructor-researcher’s dual role and potential confirmation bias were addressed.
The manuscript currently says that patterns “emerged repeatedly” and remained “remarkably stable,” but it does not define recurrence or provide an evidentiary trail supporting those judgments.
-
The Results section needs actual evidence
The Results section consists almost entirely of generalised assertions. No anonymised student quotations, reflection excerpts, sketches, AI outputs, prompt revisions, observational extracts, or case trajectories are presented.
The authors should include several concrete examples showing how a student’s initial observation became a sketch, how the sketch informed an AI interaction, what discrepancy was identified, and how a later sketch or prompt changed. A small number of carefully analyzed cases would be more convincing than repeated general statements.
The paper should also report disconfirming cases. Did some students abandon their sketches, accept AI outputs uncritically, or find that sketching constrained exploration? Such cases would improve the credibility and boundaries of the framework.
-
Ethics approval and informed consent must be addressed
The study analyses information obtained from 65 students, including classroom discussions, written reflections, sketches, AI-generated work, and instructor observations. However, the manuscript contains no Institutional Review Board/Ethics Committee statement and no informed-consent statement.
The authors must state whether approval or a formal exemption was obtained, identify the responsible body and approval details, explain how consent was handled, and describe how students’ privacy and the power imbalance inherent in instructor-led research were managed. If the data were originally collected as normal coursework and later repurposed for research, that retrospective transition requires particular clarification.
This is a potentially decisive issue. If the necessary approval, exemption, or consent cannot be documented, the material may not be publishable as human-participant research.
-
Provide essential contextual and procedural information
The manuscript should report cohort sizes and dates for each of the six offerings, student level, course duration, institutional context, assignment sequence, and relevant participant characteristics. It currently calls the course a graduate elective in one place but does not consistently describe the educational level elsewhere.
The role of AI also requires clarification. More than 30 systems were apparently used, but they are not identified. The authors should explain:
-which principal systems and versions were used;
-whether students uploaded sketches directly or translated them into text prompts;
-whether image-to-image, multimodal, or text-to-image interaction was involved;
-how system access and choice differed between cohorts;
-whether changing models across 2024–2026 could have affected the observations.
This is essential because the argument depends on how an AI system actually encountered the intermediate representation.
-
Strengthen the theoretical contribution and literature review
The synthesis of four theoretical perspectives is promising, but the manuscript relies on a very small and predominantly historical reference base. The paper needs engagement with recent scholarship on human–AI co-creativity, AI literacy, design cognition, reflective practice with generative systems, human agency, and multimodal interaction.
The authors should also clarify what is genuinely new about the proposed framework. External representations as components of cognition are already central to the theories cited. The manuscript’s distinctive contribution appears to be the placement of generative AI within this representational system, but this point needs sharper differentiation from existing work.
The boundary-object claim also needs qualification. It should be explained whether AI is being treated metaphorically as a participant, as a computational interpreter, or as part of the sociotechnical infrastructure. These possibilities carry different theoretical assumptions.
-
Moderate claims extending beyond the evidence
The evidence comes from a single art-and-design course and an instructor’s reflective analysis. Claims about education across mathematics, engineering, science, music, and other disciplines should therefore be framed as hypotheses for future research, not as implications demonstrated by this study.
Similarly, claims concerning student agency and ownership require either direct evidence from students or more cautious language. Observable references to sketches during critique do not necessarily demonstrate psychological ownership or retained creative authority.
Minor comments
-Add explicit research questions at the end of the Introduction.
-Define “intermediate representation” and “distributed creative cognition” operationally.
-Explain what changed during each design-based research cycle; course redesign alone does not establish a design-based research methodology.
-Correct “Across multiple offering” to “Across multiple offerings.”
-Sections 5.5 and 5.5 are duplicated; the limitations section should be renumbered 5.6.
-Figure 2 cites Varela as 1996, whereas the reference list gives 1992.
-Figures should specify the source, author creation, interpretation, etc.
-Correct the punctuation error in “Thompson, E..” in the Varela reference.
-Standardise the references and in-text citations to the journal’s required style.
-The AI-writing disclosure should identify the ChatGPT product/version and provide the required acknowledgement if its use went beyond ordinary language editing.
-Add the standard funding, data-availability, institutional-review, informed-consent, acknowledgment, and conflict-of-interest statements.
-The abstract should identify the study as exploratory and instructor-led and avoid implying stronger empirical confirmation than the design supports.
-Consider reducing repetition in the Discussion and Conclusion; several paragraphs restate the same claim that sketching externalises partially formed ideas.
-The final speculation about intermediate representations facilitating communication among AI systems is interesting but insufficiently connected to the study and could be shortened or removed.
Overall assessment
The manuscript contains a worthwhile educational insight and could make a useful conceptual contribution. Its figures are clear, and the proposed sketch-mediated cycle offers a potentially valuable basis for future empirical research. However, publication as a research Article requires a much more transparent methodology, concrete empirical evidence, appropriate ethics documentation, and substantially more cautious claims.
I recommend reconsideration after major revision. If the empirical and ethical requirements cannot be satisfied, I recommend that the manuscript be fundamentally reframed as a conceptual/opinion contribution, using the course only as an illustrative context.
Comments on the Quality of English Language-Across multiple offering → Across multiple offerings
-to assist with organization, editorial revision, and improve relevance → to assist with organization and editorial revision and to improve its relevance
-In equal measures → To a similar extent
-scientific-based thinking → scientific thinking or evidence-based reasoning
-sketch mediated → sketch-mediated
-Varela et al, 1992 → Varela et al., 1992
-Section numbering repeats 5.5
-Figure 2 cites Varela as 1996, while the reference list says 1992
-AI systems themselves and “biological and artificial intelligences" sound unnecessarily speculative
Author Response
Comment 1: Clarify whether this is an empirical study or a conceptual paper
Response 1:
I agree with this comment as a necessary clarification and have substantially revised the manuscript to clarify the relationship between its empirical and conceptual dimensions. The revised paper is now explicitly framed as a conceptual paper, with the classroom experiences serving as the motivating context for development of the framework rather than as empirical evidence of causal or comparative effects. This distinction is stated in the Abstract (p. 2, lines 16–27) and made explicit in the Introduction, where the manuscript now states that the classroom experiences are “not presented as a controlled comparison or as evidence that sketching causes particular learning outcomes,” but instead motivate the conceptual question developed in the paper (p. 3, lines 70–79).
The revised manuscript also separates the reflective classroom context from the claims of the framework. Section 3 describes the classroom experiences that prompted the conceptualization while explicitly acknowledging variation and counterexamples (p. 7, lines 249–264), and Figure 1 is identified as an “interpretive synthesis” rather than an experimentally validated process model (p. 7–8, lines 265–275).
Finally, claims requiring empirical validation have been reframed as propositions for future research rather than findings of the present paper. Section 5 identifies five testable propositions and describes possible comparative studies, including the need to seek disconfirming cases (pp. 9–10, lines 342–375). The limitations now explicitly state that the framework has not been validated through controlled comparison and that the classroom experiences should not be generalized beyond their context without further study (p. 10, lines 376–384)
Comment 2: The comparative claims are not supported by the reported design
Response 2:
I agree with this comment. The classroom activities were not designed as a controlled comparison between sketch-first and prompt-first conditions, and the previous version therefore allowed the classroom observations to carry more comparative weight than the design could support. I have revised the manuscript to remove that implication. The revised Introduction explicitly states that the classroom experiences are “not presented as a controlled comparison or as evidence that sketching causes particular learning outcomes” and instead identifies them as the motivating context for development of the conceptual framework (p. 2, lines 60–68).
The classroom material has consequently been reframed in Section 3, “From Classroom Practice to a Conceptual Framework,” rather than presented as comparative results. The revised discussion also acknowledges nonuniform and potentially disconfirming experiences, including cases in which students abandoned an initial sketch, preferred an AI interpretation, or found that an early representation constrained exploration (p. 6, lines 238–244). Figure 1 is explicitly described as an interpretive synthesis rather than a prescribed or experimentally validated process model (p. 6, lines 245–255).
Comparative relationships are now presented as testable propositions rather than findings. In particular, P1 and P2 identify differences between pre-AI externalization/representation-first and prompt-first workflows that should be tested empirically (p. 8, lines 326–333). The research agenda explicitly proposes prospective comparative studies of sketch-first and prompt-first conditions and identifies appropriate forms of evidence for evaluating such differences (p. 9, lines 343–349).
Comment 3: The qualitative analytical procedure requires substantial development.
Response 3:
I agree with this comment. The previous version did not describe a sufficiently systematic qualitative analytical procedure to support treating the classroom experiences as the results of a qualitative study. In revising the manuscript, I addressed this issue by clarifying the methodological status of the classroom material rather than retrospectively constructing a qualitative analytical procedure that was not part of the original course activities. The revised manuscript is now explicitly framed as a conceptual paper, with the classroom experiences serving as illustrative and motivating context rather than as a formally analyzed qualitative dataset. This distinction is stated in the Abstract (p. 1, lines 11–17) and made explicit in the Introduction, where the manuscript states that the classroom experiences are not presented as a controlled comparison or as evidence of particular learning outcomes (p. 2, lines 60–68).
Consistent with this reframing, Section 3 is now titled “From Classroom Practice to a Conceptual Framework” and describes the reflective teaching experiences that motivated development of the framework rather than presenting them as qualitative findings. The section also explicitly acknowledges variation and counterexamples in those experiences (pp. 5–6, lines 222–244), and Figure 1 is described as an interpretive synthesis rather than an experimentally validated process model (p. 6, lines 245–255).
The revised manuscript therefore does not claim that a formal qualitative analytical procedure was conducted. Instead, the conceptual claims derived from these reflective experiences are formulated as propositions for subsequent empirical investigation. The research agenda identifies forms of evidence that could support such future work, including process traces, successive representations, reflective accounts, revision behavior, and measures of agency or ownership (p. 9, lines 343–349).
Comment 4: The Results section needs actual evidence
Response 4:
I agree with this comment. A Results section making empirical claims would require systematically collected and analyzed evidence, which the reflective classroom experiences reported in the previous version were not designed to provide. I have therefore substantially reframed the manuscript rather than attempting to strengthen the Results section with retrospective evidence. The revised paper is explicitly identified as a conceptual paper, and the classroom experiences are described as illustrative and motivating context rather than as empirical evidence of causal or comparative effects (p. 1, lines 11–17; p. 2, lines 60–68).
Accordingly, the previous Results framing has been removed. Section 3 is now “From Classroom Practice to a Conceptual Framework” and explains how reflective teaching experiences motivated development of the framework without presenting those experiences as systematically analyzed findings. The section also acknowledges observations that do not uniformly support the framework, including cases in which students abandoned an initial sketch, preferred an AI interpretation, or found that an early representation constrained exploration (pp. 5–6, lines 222–244). Figure 1 is likewise explicitly described as an interpretive synthesis rather than an experimentally validated process model (p. 6, lines 245–255).
Claims that would require empirical evidence are now presented as testable propositions rather than results (pp. 8–9, lines 326–342). The subsequent research agenda describes how these propositions could be evaluated through prospective comparative studies using process traces, successive representations, reflective accounts, revision behavior, and measures of agency or ownership (p. 9, lines 343–349).
Comment 5: [Regarding ethics] If the data were originally collected as normal coursework and later repurposed for research, that retrospective transition requires particular clarification.
Response 5:
I agree that such a retrospective transition would require explicit clarification. In revising the manuscript, however, I have clarified that the classroom materials and experiences are not being presented as a dataset retrospectively repurposed for empirical research. The revised manuscript is explicitly framed as a conceptual paper in which reflective teaching experiences provide the motivating context for development of the framework. The Abstract states that the classroom experiences are illustrative rather than causal evidence (p. 1, lines 11–17), and the Introduction explains that the framework grew from reflective teaching practice across six course offerings while explicitly stating that these experiences are not presented as a controlled comparison or as evidence that sketching causes particular learning outcomes (p. 2, lines 60–68).
Consistent with this clarification, the revised manuscript does not present student coursework as a formally analyzed qualitative dataset, does not report coded student responses or comparative student outcomes, and does not make empirical claims based upon retrospective analysis of those materials. Section 3 instead describes the reflective classroom context from which the conceptual framework emerged and explicitly acknowledges variability and counterexamples in those experiences (pp. 5–6, lines 222–244).
Claims requiring systematic student data are consequently presented as propositions for prospective empirical investigation rather than as findings from the coursework. Section 5 identifies these propositions and describes possible future comparative studies and appropriate forms of evidence for testing them (pp. 8–9, lines 326–349).
Comment 6: Provide essential contextual and procedural information
Response 6:
I agree with this comment and have expanded the manuscript to provide clearer contextual and procedural information about the origins and development of the conceptual framework. The revised Introduction now explains that the framework emerged through reflective teaching practice across six offerings of an elective art and design course and identifies the kinds of classroom interactions that prompted the inquiry, including students’ formation of intention, responses to AI-generated outputs, revision processes, and perceptions of creative ownership (p. 2, lines 60–68).
Section 3, “From Classroom Practice to a Conceptual Framework,” provides additional context for the instructional setting and describes how sketching, AI interaction, revision, and reflection contributed to the development of the conceptual model (p. 5, lines 222–237). The section also reports variability in these experiences, including instances in which students abandoned an initial sketch, preferred an AI interpretation, or found that an early representation constrained exploration (pp. 5–6, lines 238–244).
At the same time, I have clarified the procedural boundaries of the paper. Because the classroom experiences were reflective teaching experiences rather than data generated through a formal qualitative or comparative research protocol, I have not retrospectively imposed procedures such as coding, experimental grouping, or systematic comparison. Figure 1 is consequently identified as an interpretive synthesis rather than a prescribed or experimentally validated process model (p. 6, lines 245–255), and claims requiring systematic empirical evidence are presented as propositions for future investigation (pp. 8–9, lines 326–349).
Comment 7: Strengthen the theoretical contribution and literature review
Response 7:
I agree with this comment and have substantially revised both the theoretical contribution and the literature review. The revised literature review now integrates the foundational perspectives of embodied cognition, distributed cognition, distributed intelligence, sketching, and boundary objects (pp. 2–3, lines 83–119) with a substantially expanded review of recent research on human–AI co-creativity, AI literacy and prompt engineering, multimodal design interaction, timing of AI participation, and human agency. A new Section 2.5, “Contemporary human–AI co-creativity with intermediate representation,” develops this synthesis in detail (pp. 3–5, lines 120–221).
The expanded review incorporates recent work on interaction structures in human–AI co-creativity, student agency, AI literacy and prompt engineering, generative AI in design ideation, multimodal sketch-based interaction, timing of AI participation, and agency during creative collaboration. Importantly, these sources are used to establish the theoretical distinction at the center of the revised paper. The manuscript distinguishes prompt competence from the cognitive formation of intention before prompting (p. 4, lines 150–158) and distinguishes the use of sketches as multimodal inputs to AI from their role as cognitive intermediate representations for the human creator (pp. 4–5, lines 168–186).
The theoretical contribution has consequently been sharpened around the role of intermediate representations within a distributed creative cognitive system. Rather than arguing simply that sketching is preferable to prompting, the revised framework proposes that an intermediate representation can externalize emerging intention, preserve ambiguity, provide a persistent object for revision and negotiation, and help organize the timing and form of AI participation. The critical synthesis at the end of the literature review makes explicit that the framework does not propose replacing prompting or asserting the superiority of manual techniques; instead, it treats the timing and form of AI participation as variables in the organization of creative cognition (p. 5, lines 216–221).
Sections 4 and 5 further develop this theoretical contribution through mechanisms concerning intention formation, representational negotiation, ambiguity, human agency, and productive friction (pp. 7–8, lines 270–312), followed by five propositions that make the framework available for subsequent empirical testing (pp. 8–9, lines 326–342).
Comment 8: Moderate claims extending beyond the evidence
Response 8:
I agree with this comment and have revised the manuscript throughout to moderate claims that extended beyond what the classroom experiences could support. Most importantly, the revised paper is explicitly framed as a conceptual paper rather than an empirical demonstration of the effects of sketching. The Abstract now describes the classroom experiences as illustrative rather than causal evidence (p. 1, lines 11–17), and the Introduction explicitly states that they are “not presented as a controlled comparison or as evidence that sketching causes particular learning outcomes” (p. 2, lines 60–68).
Claims arising from the classroom experiences have also been qualified within the development of the framework itself. Section 3 now acknowledges that the observed patterns did not occur uniformly and includes potentially disconfirming cases in which students abandoned an initial sketch, preferred an AI interpretation, or found that an early representation constrained exploration. The manuscript therefore explicitly states that the framework does not assume that sketching always improves human–AI collaboration (p. 6, lines 238–244). Figure 1 is similarly identified as an interpretive synthesis rather than a prescribed or experimentally validated process model (p. 6, lines 245–255).
The theoretical discussion has likewise been revised to distinguish proposed mechanisms from demonstrated effects. For example, the relationship between intermediate representations and human agency is explicitly characterized as a proposed mechanism requiring empirical investigation (p. 8, lines 292–299), while productive friction is described in terms of what it “may” contribute and is explicitly identified as requiring direct study (p. 8, lines 300–312).
Finally, claims that extend beyond the available classroom evidence have been reformulated as five testable propositions rather than findings (pp. 8–9, lines 326–342). The research agenda specifies prospective approaches for testing these propositions and explicitly calls for disconfirming cases to establish the framework's boundary conditions (p. 9, lines 343–349). The limitations section further states that the framework has not been validated through controlled comparison and should not be generalized beyond its context without further study (p. 9, lines 350–358).
Comment 9: Minor Revision Comments [various]
Response 9: I very much appreciate all the minor revision comments that pointed out typos, grammatical errors, awkward language, and redundancies. I did my best to make all the changes suggested because I found them to be appropriate and necessary for correction.
Round 2
Reviewer 1 Report
Comments and Suggestions for AuthorsThe revised paper addresses the main concerns raised in the previous review and is considerably more coherent in its present conceptual form. A few minor points could further strengthen the final version, such as careful final proofreading to remove remaining typographical and formatting problems and ensure consistency in citation and reference formatting. It may also be useful to review occasional formulations that risk attributing human-like cognitive processes to AI systems; for example, statements about representations changing “the thinking that occurs on both sides of the interaction” could be phrased more precisely in terms of human cognition and computational processing. Finally, the five propositions for future research are a valuable addition, but their modality could be made more consistent. For example, P4 uses the stronger formulation “will support,” whereas most of the other propositions appropriately use “may.”
Comments on the Quality of English LanguageThe revised paper is generally clear and professionally written, and the language has improved considerably. However, I would still recomment one final language-editing pass, which would address occasional awkward formulations, grammatical inconsistencies, typographical errors, and minor problems of sentence structure and expression. These issues do not hinder comprehension but should nevertheless be corrected before publication.
Author Response
Comments 1:
A few minor points could further strengthen the final version, such as careful final proofreading to remove remaining typographical and formatting problems and ensure consistency in citation and reference formatting. It may also be useful to review occasional formulations that risk attributing human-like cognitive processes to AI systems; for example, statements about representations changing “the thinking that occurs on both sides of the interaction” could be phrased more precisely in terms of human cognition and computational processing.
Response 1:
I agree with all the suggestions and have made the corresponding revisions. I conducted a final proofreading and language-editing pass throughout the manuscript, correcting typographical, grammatical, formatting, and citation inconsistencies. I also reviewed formulations describing AI processes to avoid unnecessarily anthropomorphic language. In particular, the statement that an intermediate representation can change “the thinking that occurs on both sides of the interaction” has been revised to distinguish human cognitive activity from computational processing. Finally, I reviewed the modality of the five propositions for consistency and changed P4 from “will support” to “may support,” consistent with the conceptual and prospective status of these propositions.
And for addressing the Quality of English
I agree with the recommendation. I conducted a final language-editing pass throughout the manuscript and corrected remaining awkward formulations, grammatical inconsistencies, typographical errors, and minor sentence-structure and formatting issues. Specifically:
Typographical & Punctuation Issues:
- Section Headers – Fixed spacing issues for consistency and compliance.
- Reference List Inconsistencies:
- Hutchins (1996): Missing period after year bracket: (1996) Cognition... (1996). Cognition...
- Pea (1993): Trailing comma after year: (1993), A Heuristic... (1993). A Heuristic...
- Rezwana & Maher (2023): Trailing comma: (2023), Designing... (2023). Designing...
- Star & Griesemer (1989): Abbreviated author structure and missing full stop: Star, S. and Griesemer, J (1989) Star, S., & Griesemer, J. (1989).
- Suwa & Tversky (2009): Missing period at the end of the entry.
Sentence Structure & Phrasing Optimizations:
- Section 4.5 (Paragraph 1) — Awkward Phrasing:
- Current: "...differentiate into encapsulated agentic services."
- Fix: "differentiate into specialized agentic services."
- Section 6 (Paragraph 3) — Run-on Syntax:
- Current: "As intelligent systems become collaborators across educational disciplines, researchers will need to understand how intermediate representations shape the relationship between human intention and computational generation, during a period of time when the computational systems may be changing rapidly and differentiate into encapsulated agentic services."
- Fix: Broke the long sentence into two sentences with a period after generation
Author Response File:
Author Response.pdf

