1. Introduction
Generative AI is increasingly entering product design practice and has shown substantial supporting capabilities in requirement understanding, concept generation, visual expression, and scheme iteration [
1]. By accelerating information processing, design exploration, and formal development [
2], it can improve the efficiency of early-stage design and broaden the range of possible solutions [
3,
4,
5]. However, visual realism and generative fluency do not necessarily imply engineering feasibility, regulatory compliance, or value appropriateness [
6]. This limitation becomes especially important in highly constrained contexts such as medical devices, healthcare services, and other safety-critical design tasks, where human judgment remains essential for assessing risk, feasibility, and stakeholder trade-offs [
7]. Therefore, in ethics-sensitive product development, gains in generative efficiency cannot replace the need for human oversight, explicit priority setting, and transparent decision support [
8].
As AI becomes more deeply embedded in design activities, human-AI co-design has gradually become an important research focus. Existing studies have applied Generative AI tools such as ChatGPT-4.1 and Midjourney V6.1 to tasks including requirement interpretation, visual generation, style exploration, and creative expansion [
9]. These studies have improved local stages of the design process, particularly in ideation, visual communication, and workflow acceleration. As shown in
Table 1, however, their main contributions remain concentrated on generative support, local optimization, and efficiency improvement. Comparatively, less attention has been paid to how design priorities are translated into structured generation conditions, how multiple candidate schemes are compared on explicit grounds, and how human and AI roles are organized across the full process of requirement analysis, scheme generation, and final selection [
10].
This limitation becomes more salient in ethics-sensitive product development. Such tasks involve not only functional realization and formal expression, but also usability, privacy protection, service accessibility, regulatory constraints, and value trade-offs among different stakeholders [
16]. Existing approaches may accelerate scheme production, but they often fail to connect three key aspects within one coherent process: where human judgment should remain central, how human-defined priorities enter AI-supported generation, and how alternative schemes are compared and selected in an explainable manner [
17]. When role organization is unclear, requirement translation is unstable, and evaluation support is weak [
18], AI may increase output speed without adequately supporting accountable collaborative design [
19].
Accordingly, this study focuses on the following problem: how to organize human-AI co-design in ethics-sensitive product development so that requirement expression, candidate scheme generation, and final evaluation form a continuous process. On this basis, the study addresses three research questions. RQ1: How can human and AI roles be organized across requirement analysis, candidate generation, and final selection so that key judgment stages remain human-led? RQ2: How can human-defined requirement priorities be translated into structured generation conditions and comparable evaluation dimensions? RQ3: In a case-based quasi-experimental setting, what preliminary evidence can be observed regarding design efficiency, value drift, and human oversight under the proposed framework?
To address these questions, this study proposes GAGT, an explainable HCI-based decision support framework for human-AI co-design. The framework organizes requirement analysis, candidate generation, and result selection into a continuous process. Within this process, the Analytic Hierarchy Process (AHP) [
20] is used to structure design requirements and determine their priorities, Grey Relational Analysis (GRA) [
21] is used to compare candidate schemes across multiple dimensions, and the Technique for Order Preference by Similarity to Ideal Solution (TOPSIS) [
22] is used to support final ranking. Within this framework, human participants mainly undertake requirement confirmation, priority judgment, review at key checkpoints, and final scheme selection, while AI mainly supports information organization, candidate scheme generation, and quantitative comparison. In this way, the framework seeks to establish a clearer division of labor between generative support and judgment support while maintaining a relatively explicit relationship among design requirements, generated results, and evaluation criteria [
23]. The framework is further applied to the design of a community medical vehicle to examine its applicability in an ethics-sensitive context.
The contribution of this study is threefold. First, at the conceptual level, it frames human-AI co-design as a continuous decision support problem centered on role organization, requirement translation, and explainable evaluation. Second, at the methodological level, it develops a staged combination of AHP, GRA, and TOPSIS to support priority structuring, candidate comparison, and final ranking within one workflow. Third, at the empirical level, it provides preliminary observations on design efficiency, value drift, and human oversight through a case-based quasi-experimental application in community medical vehicle design. Given the small-sample case setting, these findings are more appropriately interpreted as context-specific exploratory results.
2. Related Theories and Methods
2.1. Generative Artificial Intelligence (GenAI)
Generative AI is increasingly used in design activities to support the production of textual, visual, and multimodal design materials, thereby assisting concept development, sketch generation, and early-stage communication [
24,
25]. In product design, its value lies not only in rapid output generation, but also in its ability to expand the range of alternatives available to designers within a limited time [
26].
From the perspective of design support, GenAI mainly demonstrates three capabilities. First, it supports information processing. With the assistance of large language models, user descriptions, literature materials, contextual constraints, and design semantics can be organized, extracted, and summarized more efficiently, providing more focused inputs for subsequent design activities [
27]. Second, it supports candidate scheme generation. Through prompt-based generation and iterative adjustment, designers can obtain multiple formal directions, stylistic features, or design variants within a short period of time, thereby facilitating creative expansion and design exploration [
28]. Third, it supports visual expression and iterative refinement. GenAI accelerates the translation from conceptual descriptions to visible outputs, making it easier for designers to test, compare, and revise ideas through repeated feedback cycles [
29].
However, these advantages mainly strengthen generative support rather than decision support. Whether a generated scheme is feasible, compliant, and appropriate to the design context still depends on human judgment [
7]. In addition, prompts do not necessarily correspond to stable design priorities. Design requirements often involve functional expectations, user preferences, contextual constraints, and value orientations at the same time. Without further structuring, these elements may be weakened, shifted, or unevenly represented during the generation process [
30]. Moreover, although GenAI can expand the range of possible schemes, it does not by itself provide a clear basis for comparing alternatives or selecting a final result. Therefore, in the present study, GenAI is understood primarily as a tool for generative support in human-AI co-design, while priority expression, scheme comparison, and final selection still require additional decision support mechanisms.
2.2. Human-AI Co-Design
As Generative AI becomes more deeply embedded in design activities, human-AI co-design has become an important research direction. Existing studies mainly explore how AI participates in creative generation, information processing, and workflow advancement, and how humans and AI form collaborative relationships across different design stages [
31].
Current research on human-AI co-design can be broadly grouped into three strands. The first emphasizes creative generation, using prompt input, image generation, and iterative feedback to support concept exploration and scheme expansion [
32]. The second is interaction-oriented, focusing on how prompt refinement, feedback adjustment, and iterative mechanisms can improve the usability and controllability of AI-assisted design processes [
33]. The third adopts a system- or framework-oriented perspective, seeking to improve output quality and process efficiency through clearer collaborative structures and more explicit process linkage [
34]. Together, these studies indicate that AI is no longer limited to peripheral assistance, but is gradually entering more central stages of design activity.
Even so, from the perspective of decision support, current human-AI co-design research still shows three shared limitations. First, role organization remains underdeveloped. Many studies do not clearly specify where human judgment should remain central, where AI should provide support, and how the division of labor should be maintained across the workflow [
35]. Second, requirement translation remains unstable. Although existing methods can convert textual input into generative instructions, prompts do not necessarily preserve relatively stable human-defined priorities across multiple rounds of generation [
36]. Third, evaluation support is often weak. Many frameworks emphasize generation and iteration, but provide relatively limited support for explainable comparison among multiple candidate schemes and for transparent grounds of final selection [
37]. As shown in
Table 2, although existing studies differ in emphasis, they still do not fully connect role organization, requirement translation, and evaluation support within one continuous process.
These limitations become more salient in ethics-sensitive product development. In such contexts, design tasks require not only scheme generation, but also a relatively stable way for human-defined priorities to enter the design process, while ensuring that scheme comparison and final selection rest on relatively explicit grounds. When role organization is unclear, requirement translation is unstable, and evaluation support is insufficient, AI may improve generative efficiency without providing complete support for accountable collaborative design. Therefore, although existing human-AI co-design research has made progress in creative expansion and interaction optimization, it still lacks a relatively coherent decision support logic linking role organization, requirement translation, and final evaluation [
38].
2.3. Design Decision-Making Methods
In human-AI co-design, GenAI can improve the efficiency of candidate scheme generation, but it also increases the complexity of comparison and final selection. Especially in ethics-sensitive product development, design work involves not only generating alternatives, but also expressing priorities, comparing candidate schemes, and supporting final convergence. In this context, decision-making methods are relevant not only because they support ranking, but also because they can provide a clearer basis for requirement structuring, alternative comparison, and result selection [
39].
Existing decision-making methods differ in their functional emphasis. Some are more suitable for constraint screening and boundary control, some for priority determination, some for candidate scheme comparison, and others for result ranking. These methods correspond to different stages of design decision-making, but when used independently, they usually cover only part of the process. As shown in
Table 3, individual methods have targeted advantages in specific stages, yet they do not by themselves form a continuous support process covering requirement expression, scheme comparison, and final selection.
Among these methods, non-compensatory approaches such as ELECTRE are useful for screening constraints and controlling boundaries, but they offer limited support for requirement structuring and result ranking when used alone. AHP is effective for requirement decomposition and priority determination, but cannot by itself complete candidate scheme comparison. GRA is suitable for multi-scheme comparison under multiple indicators, yet it does not replace front-end priority expression or back-end ranking. TOPSIS supports result ranking and convergence, but it depends on predefined indicators and weights. Aggregation-based methods such as weighted-sum approaches can provide comprehensive scoring, but their explainability is often limited in complex trade-off situations. These observations suggest that different methods are useful at different stages, but no single method adequately supports the entire decision support chain required in human-AI co-design.
To address the limitations of individual methods, existing studies have attempted to combine multiple decision-making approaches. For example, AHP-TOPSIS is often used to connect weight determination with result ranking, while AHP-GRA combines priority expression with candidate comparison. Hybrid approaches involving fuzzy methods, neural-network-based ranking, and other integrated evaluation strategies are also frequently adopted to address uncertainty in complex evaluation tasks [
42,
43,
44,
45,
46]. In general, these combinations can strengthen evaluation structure, improve result stability, and increase computational efficiency.
However, from the perspective of human-AI co-design, existing hybrid methods still focus mainly on evaluation computation itself. They pay relatively limited attention to how human-defined priorities enter the generation process, how AI-generated schemes can be compared on explicit and explainable grounds, and how a clearer division of judgment between humans and AI can be maintained throughout the workflow. As summarized in
Table 4, existing hybrids strengthen local evaluation capability, but they still do not fully connect the three key dimensions emphasized in this study: role organization, requirement translation, and final evaluation.
Based on the above review, the issue in human-AI co-design is not simply how to rank results more accurately, but how to organize priority expression, candidate scheme comparison, and final selection as a relatively continuous process. From a functional perspective, AHP is suitable for structuring design priorities, GRA is suitable for comparing candidate schemes across multiple dimensions, and TOPSIS is suitable for supporting result ranking. When assigned to different stages of the workflow, these methods help establish a clearer support relationship from requirement expression and scheme comparison to result convergence.
Accordingly, the present study adopts a staged combination of AHP, GRA, and TOPSIS as the methodological basis of the proposed framework. In this arrangement, AHP supports front-end priority structuring, GRA supports mid-stage scheme comparison, and TOPSIS supports back-end ranking and convergence. The significance of this combination lies not in claiming a universally optimal hybrid, but in providing a more coherent decision support arrangement for the specific problem addressed here, namely, how to connect role organization, requirement translation, and explainable evaluation within one human-AI co-design workflow.
3. Structure and Operational Mechanism of the GAGT Framework
GAGT (
https://github.com/xieyu0705/GAGT-framework, accessed on 9 April 2026) is an explainable decision support framework for human-AI co-design. It consists of three connected modules: requirement analysis, creative generation, and decision optimization. These modules organize requirement expression, candidate scheme generation, and result evaluation into a continuous process, so that front-end requirements, middle-stage candidate schemes, and back-end evaluation results remain linked within the same workflow. In this process, AI mainly supports information organization, candidate generation, and quantitative analysis, while human participants mainly undertake requirement confirmation, priority judgment, key-stage review, and final selection. Through this arrangement, the framework maintains a relatively clear relationship among design requirements, generated results, and evaluation criteria in ethics-sensitive product development.
3.1. Framework Structure and Module Division
GAGT consists of a requirement analysis module, a creative generation module, and a decision optimization module. These three modules correspond respectively to requirement expression, candidate scheme formation, and result selection, and remain connected through input-output relationships, as shown in
Figure 1.
The requirement analysis module is located at the front end of the framework. Its main function is to transform raw task information into requirement expressions that can enter subsequent stages. Its inputs include user requirements, contextual materials, constraint information, and related textual materials, and its outputs are the confirmed requirement set and the indicator basis for subsequent weighting and evaluation. This module provides the analytical basis for the workflow.
The creative generation module is located in the middle stage of the framework. It takes structured requirements and generation conditions as inputs and supports the formation and expansion of candidate schemes through Generative AI. Its output is a candidate scheme set. The role of this module is to expand the design space and provide multiple alternatives for subsequent comparison.
The decision optimization module is located at the back end of the framework. It takes candidate schemes, evaluation indicators, and weight information as inputs, and supports scheme comparison, ranking, and screening through multi-criteria decision-making methods. Its outputs are comparison results, ranking results, and the selected scheme. The role of this module is to provide an analytical basis for result convergence and final review.
Viewed as a whole, the three modules do not operate independently. Instead, they unfold sequentially around requirement expression, scheme formation, and result selection. The requirement analysis module provides the analytical basis, the creative generation module produces candidate schemes, and the decision optimization module supports scheme comparison and final selection. Their functional relationships are summarized in
Table 5.
3.2. Workflow and Information Transmission Mechanism
The operation of GAGT consists of five steps: requirement input, priority structuring, candidate scheme generation, scheme comparison, and result ranking. Each step takes the output of the previous stage as its input, allowing requirement expression, candidate generation, and result evaluation to remain connected within one workflow, as shown in
Figure 2.
In this workflow, H1-H3 denote the key human review nodes, corresponding respectively to requirement confirmation, generation control, and result review, while A1–A3 denote the key AI-supported nodes, corresponding respectively to data processing, scheme generation, and quantitative analysis. These nodes operate across the requirement analysis, creative generation, and decision optimization stages, forming a relatively clear transmission path from raw requirements to final scheme selection.
In the requirement input step, raw task information first enters the framework, including user requirements, contextual materials, constraints, and related textual information. Node A1 is responsible for organizing, extracting, and summarizing unstructured information so as to form preliminary requirement items. Node H1 then screens, revises, and confirms these items, producing a requirement set for subsequent analysis, as shown in
Figure 3.
In the priority structuring step, the confirmed requirement set is further organized into a hierarchical structure. Node H2 undertakes the main judgment task at this stage by comparing the relative importance of indicators according to the design objectives and constraints. AHP is embedded here to generate requirement weights and establish the weighted requirement structure. The output of this step is the priority structure used in subsequent generation and evaluation, as shown in
Figure 4.
In the candidate scheme generation step, the weighted requirement structure enters the creative generation stage. At this stage, node H2 provides generation directions and control conditions based on the priority structure formed in the previous step, thereby defining the main requirements for subsequent scheme generation, as shown in
Figure 5. On this basis, node A2 generates multiple candidate schemes accordingly, forming a differentiated and comparable candidate set, as shown in
Figure 6. The purpose of this step is not to directly determine a final result, but to provide multiple alternatives for subsequent comparison and ranking.
In the decision optimization stage, the candidate scheme set and the weighted requirement structure jointly enter node A3 for subsequent analysis. GRA is first used to compare the responsiveness of candidate schemes across multiple evaluation dimensions, producing the comparison results, and TOPSIS is then applied on this basis to generate the final ranking of candidate schemes. After that, node H3 reviews the ranking results in light of the specific task context and completes the final scheme selection. The output of this stage is the ranking result together with the final selected scheme, as shown in
Figure 7.
Viewed as an integrated workflow, AHP, GRA, and TOPSIS are embedded respectively in the steps of priority structuring, scheme comparison, and result ranking. At the same time, H1–H3 and A1–A3 undertake different tasks across the workflow, allowing front-end requirement expression, middle-stage candidate generation, and back-end result selection to remain continuously connected. In the present framework, this connection is maintained through workflow sequencing, node allocation, and stage-based review. The main workflow elements are summarized in
Table 6.
Through this workflow, requirement expression, candidate scheme generation, and result evaluation are no longer disconnected, but form a continuous relationship through node division and method embedding.
3.3. Human-AI Division of Labor and Decision Support Relationship
The human-AI division of labor in GAGT is organized around task types rather than a simple separation between “human” and “AI.” Within this framework, AI mainly undertakes information processing, candidate scheme generation, and quantitative analysis, while human participants mainly undertake requirement confirmation, priority judgment, key-stage review, and final selection. This arrangement distributes generative support and judgment support across different nodes while maintaining a relatively clear correspondence among design requirements, generated results, and evaluation criteria.
More specifically, A1, A2, and A3 correspond respectively to data processing, candidate scheme generation, and comparison and ranking analysis, whereas H1, H2, and H3 correspond respectively to requirement confirmation, generation control, and final review. At the front-end stage, A1 supports text organization, semantic extraction, and requirement summarization, while H1 confirms task objectives, screens requirement items, and revises requirement boundaries. At the middle stage, A2 supports candidate generation and iterative expansion, while H2 determines generation directions, reviews intermediate outputs, and adjusts generation conditions. At the back-end stage, A3 undertakes comparison and ranking analysis, while H3 reviews the analytical results and completes final scheme selection according to the specific task context. The corresponding division of labor is summarized in
Table 7.
Overall, the human-AI division of labor in GAGT forms a collaborative relationship centered on decision support. Rather than assigning humans and AI to isolated roles, the framework distributes different functions across the workflow so that requirement expression, candidate scheme generation, and result selection remain connected within the same process. In this way, the framework provides a clearer basis for linking generative support with human review in ethics-sensitive product development.
4. Case Study
4.1. Validation Scenario and Experimental Design
This study selected the design of a community medical vehicle as the validation scenario in order to examine the application of GAGT in an ethics-sensitive product design task. A community medical vehicle serves primary healthcare contexts, and its design involves not only spatial organization, equipment configuration, and operational processes, but also privacy protection, service accessibility, and coordination among multiple stakeholders. Compared with general product design tasks, this scenario is characterized by clearer constraints and a more complex requirement structure, making it suitable for observing the relationships among requirement expression, candidate generation, and result screening in human-AI co-design.
To compare process differences under different design conditions, three experimental conditions were established: human-only design (C1), AI-autonomous generation (C2), and GAGT-supported human-AI co-design (C3). In C1, designers independently completed requirement interpretation, concept generation, and scheme expression. In C2, AI generated candidate schemes under a unified input condition, while designers mainly undertook result organization and necessary screening. In C3, GAGT was introduced into the stages of requirement analysis, candidate generation, and decision optimization, and the design outcome was produced through collaboration between designers and AI. All three conditions used the same design brief, contextual materials, and baseline constraints to maintain a consistent basis for comparison. This study adopted a small-sample, case-based, quasi-experimental design, and result interpretation focuses primarily on process differences and observed outcome characteristics.
Participants included industrial designers, medically related stakeholders, and community residents. In the experimental stage, the same five industrial designers participated in both C1 and C3, while three industrial designers participated in C2. In the formal evaluation stage, 15 medically related stakeholders and 15 community-resident stakeholders were invited to rate the design outputs of the three conditions under a unified evaluation framework. Because C1 and C3 involved the same five designers and were relatively closer in terms of process structure, the comparison between these two conditions was treated as the primary matched comparison for examining differences between human-only design and GAGT-supported design. By contrast, C2 was used mainly as a descriptive reference condition for generative efficiency and risk-related characteristics. Since the overall sample size remained limited and some designers participated in more than one condition, learning and carryover effects cannot be fully excluded despite the interval arrangement. Accordingly, the interpretation of condition differences in this study remains exploratory, as shown in
Table 8 and
Table 9.
4.2. Experimental Implementation and Evaluation Procedure
Before the experiment began, all three groups were provided with the same design brief, contextual materials, constraint descriptions, and reference information, and were required to complete the concept design output for the community medical vehicle within the specified time period. Differences among the three conditions lay primarily in the process pathway rather than in the task basis itself. In C1, designers independently completed requirement interpretation, scheme generation, iterative modification, and final expression. In C2, AI generated candidate schemes based on the given input, while designers mainly undertook result organization and screening. In C3, designers first confirmed the requirements on the basis of the task materials and established the indicator structure used for subsequent evaluation, then generated candidate schemes under structured requirement conditions, and finally conducted within-group ranking of the schemes entering the comparison stage in order to support final screening.
The formal evaluation stage adopted a unified 16-indicator framework covering four dimensions: medical functionality, technological integration, user experience, and sustainability. The outputs of all three conditions entered the same evaluation framework and were rated by medically related stakeholders and community-resident stakeholders. In procedural terms, all three conditions were conducted around the same design task and entered subsequent evaluation under the same indicator system so as to maintain a consistent comparison basis. The within-group ranking in C3 was used mainly for within-group screening, whereas cross-condition comparison continued to rely on stakeholder rating results.
In terms of indicator construction, this chapter adopts the evaluation system established through the AHP results presented earlier. The four criterion layers are medical functionality, technological integration, user experience, and sustainability, with corresponding weights of 0.579, 0.233, 0.121, and 0.067. On this basis, 16 indicators were further specified for subsequent scheme rating, comparison, and result analysis. This indicator system served both as the basis of the formal evaluation and as a shared reference for within-group comparison and ranking in C3, as shown in
Table 10.
Overall, the experimental procedure maintained consistency in task basis and evaluation criteria in order to observe how different design conditions affected process progression and result formation. At the same time, because the sample size was limited and the three experimental conditions were not fully symmetrical, the following analysis focuses primarily on performance differences observed within the case context.
4.3. Experimental Results and Analysis
To compare differences among the three experimental conditions in terms of design efficiency, result consistency, and risk control, the core results are summarized in
Table 11. The overall performance of the three sets of schemes across the 16 indicators is shown in
Figure 8, and the within-group ranking results of candidate schemes in C3 are shown in
Figure 9. Given the limitations of sample size and experimental configuration, the results reported in this section are used mainly to describe performance differences across conditions rather than to support strict causal interpretation. For the matched C1–C3 comparison, exact Wilcoxon signed-rank tests were applied to full-cycle time, concept iterations, and revision cycles. All three metrics showed consistent directional differences in favor of C3 (one-sided
p = 0.0313 for each metric). For robustness checking, the same contrasts were also examined as independent observations, under which Cliff’s δ reached 1.00 for all three metrics, indicating complete directional separation within the observed sample. These statistics are reported to characterize the present case and should be interpreted as exploratory rather than population-level evidence.
In terms of design efficiency, C2 showed the shortest design cycle at 7.3 days, followed by C3 at 10.0 days, while C1 showed the longest cycle at 22.8 days. At the same time, C2 still exhibited relatively high downstream adjustment costs, with an average of 11.2 concept iterations and 10.83 revision rounds. By comparison, C3 showed lower values than C1 in full-cycle time, concept iterations, and revision cycles, suggesting that, under the present case conditions, structured requirement expression and subsequent comparison and screening helped reduce repeated revisions. Relative to C1, the design cycle of C3 was shortened by approximately 56.1%. The matched C1–C3 comparison further showed consistent directional differences across all three process metrics (exact Wilcoxon one-sided p = 0.0313 for each metric).
In terms of value drift and risk control, C2 showed the highest Value Drift Index (VDI) at 30.7 and recorded 8 HIPAA violations. C1 showed the lowest VDI at 6.0 and no HIPAA violations, while C3 showed a VDI of 16.1 and likewise no HIPAA violations. These results indicate that the purely generative pathway, although efficient, was relatively weaker in requirement retention and compliance control. Human-only design remained comparatively stable in drift and risk control, but at a higher efficiency cost. C3 showed a relatively more balanced pattern between efficiency and risk control.
In terms of design space coverage and external evaluation consistency, C3 achieved a Design Space Coverage (DSC) of 70.50%, higher than 63.82% in C1 and 55.27% in C2. Its stakeholder rating consistency (ICC) was 0.88, also higher than 0.65 in C1 and 0.36 in C2. Combined with
Figure 8, these results indicate that C3 showed relatively more balanced overall performance across the 16 indicators and higher consistency in external evaluation.
Overall, C2 was more consistent with a generative pattern characterized by high speed, high drift, and high risk; C1 more closely resembled a human design pattern characterized by low drift, low risk, and low efficiency; and C3 showed a collaborative pattern characterized by a moderate cycle, lower drift, and higher consistency. In the present case, the role of GAGT lies less in simply accelerating generation than in organizing requirement expression, scheme comparison, and result screening into a relatively clear decision-support pathway. Since this study is based on a small-sample, quasi-experimental, case-based validation, these findings are better understood as preliminary observations.
5. Discussion and Outlook
A comparison of the case results shows that the role of GAGT in the community medical vehicle design task is reflected mainly in two aspects. First, it integrates requirement expression, candidate scheme comparison, and result screening into a unified process, so that scheme development is supported by clearer stage-to-stage linkage. Compared with the human-only condition, C3 showed lower values in full-cycle time, concept iterations, and revision cycles, indicating that, when structured requirement expression is coordinated with subsequent comparison and screening, the design process becomes more focused and repetitive adjustments are reduced. Second, while improving generative efficiency, the framework also maintains attention to value drift control and external evaluation consistency. Compared with the AI-autonomous condition, C3 showed relatively more stable results in indicators such as VDI, HIPAA violations, and ICC, suggesting that, in this case, human review nodes and a unified evaluation structure appear to have supported more controlled result convergence.
From the perspective of result formation, the efficiency advantage observed in C3 is related both to workflow organization and to the speed advantage of AI at the candidate generation stage. Its lower value drift and higher external evaluation consistency are also closely associated with the unified evaluation structure, the designers’ continued involvement in organizing candidate schemes, and the final screening process. These findings indicate that GAGT provides a continuous process organization for human–AI co-design: the front end clarifies design priorities through requirement structuring, the middle stage expands the design space through candidate scheme generation, and the back end promotes result convergence through comparison and screening. With such a workflow, human–AI co-design develops from localized generative support into a more complete process with clearer stage logic and decision grounds.
Compared with existing related studies, GAGT shows a more distinctive process-oriented organization. In studies of creative generation and visual expansion, the role of AI is more often reflected in rapid generation and formal exploration. In integrated evaluation studies, multiple methods are typically used to improve the structure and practical applicability of evaluation. GAGT connects these two lines of concern by introducing structured requirement expression before candidate generation to clarify design priorities, and by applying comparison and ranking mechanisms after candidate formation to support scheme screening. Within this framework, AHP, GRA, and TOPSIS respectively undertake the tasks of priority expression, scheme comparison, and result convergence. What emerges is a continuous decision-support pathway centered on requirement clarification, generation, comparison, and screening. This pathway provides designers with clearer grounds for judgment and also establishes a more explicit process basis for human–AI collaboration in ethics-sensitive design tasks.
At the same time, several limitations remain in this case analysis. First, the study was conducted under a small-sample, case-based, quasi-experimental setting, and the three experimental conditions were not fully symmetrical in either sample size or process configuration. Accordingly, the results are more appropriately interpreted within the present case context. Second, because some designers participated in more than one experimental condition, potential learning and carryover effects cannot be fully excluded even though a time interval was arranged. Third, although the outputs of all three conditions were externally evaluated under the same indicator system, C3 still included a within-group screening stage before formal comparison, which means that the final results were shaped to some extent by the specific experimental procedure. In addition, although the workflow, node allocation, and indicator structure are described relatively clearly, implementation-level reproducibility remains limited, since prompt logs, screening rules, scoring governance, and some statistical analysis details still require further documentation. Therefore, the current findings are more appropriately understood as case-specific and exploratory observations, and their interpretation should remain within the boundaries of the present research context.
On this basis, future research may proceed in several directions. First, larger samples and more balanced experimental conditions are needed to improve comparability across conditions. Second, further validation can be conducted in a wider range of ethics-sensitive product design tasks in order to examine the applicability of the framework across different design contexts. Third, more complete documentation of prompt logs, screening procedures, scoring governance, and analytical details is needed to further improve reproducibility and methodological rigor. Overall, GAGT can be understood as a preliminary process framework for a specific human-AI co-design context, supporting requirement structuring, scheme comparison, and result convergence. Its further development still depends on richer case validation and more complete methodological documentation.