Next Article in Journal
Dynamics, Statistical Analysis, and Spectral Line-Shape of Semiconductor Lasers Subject to Optical Feedback
Previous Article in Journal
Pharyngeal Airway Changes After Mandibular Advancement with Clear Aligners in Growing Patients with Class II Malocclusion Due to Mandibular Retrusion: A Retrospective Controlled Pilot Study
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

User Experiences with Prompt-Supported Text-to-Image and Text-to-Video Task Conditions for Visual Creation in a Traditional Chinese Cultural Context: An Exploratory Study

Institute of Industrial Design, School of Mechanical Engineering, Shandong University, Jinan 250061, China
*
Authors to whom correspondence should be addressed.
Appl. Sci. 2026, 16(19), 9429; https://doi.org/10.3390/app16199429 (registering DOI)
Submission received: 16 August 2026 / Revised: 9 September 2026 / Accepted: 18 September 2026 / Published: 22 September 2026
(This article belongs to the Section Computing and Artificial Intelligence)

Featured Application

This study documents behavioural process measures and subjective experiences observed in prompt-supported text-to-image and text-to-video task conditions for visual creation in a traditional Chinese cultural context and informs future controlled evaluation of culturally oriented human–AI creation tools.

Abstract

Generative artificial intelligence (GenAI) is increasingly used in visual creation within a traditional Chinese cultural context, yet user experiences with prompt-supported text-to-image (T2I) and text-to-video (T2V) tasks remain underexplored. This exploratory study examined behavioural process measures and subjective experiences under prompt-supported T2I and T2V task conditions. Thirty-six participants, including 18 in the design-background group and 18 in the non-design-background group, completed two T2I and two T2V tasks using a fixed alternating sequence and condition-specific theme pools. Measures included total observed task duration, iteration count, NASA-TLX-based workload, overall flow-related score, process satisfaction, outcome satisfaction, UEQ-inspired semantic-differential evaluations, and participant-reported cultural and aesthetic impressions. Under the implemented fixed sequence and condition-specific task configurations, the T2V tasks were associated with longer total observed task duration and higher subjective workload than the T2I tasks, whereas no statistically detectable difference was found in outcome satisfaction. Within the recorded tasks, ending with direct adoption of a system-generated suggestion was associated with shorter total observed task duration and fewer iterations; outcome satisfaction was summarised descriptively by terminal strategy. Exploratory analyses also identified participant-background differences in several subjective evaluations, while participant-reported cultural and aesthetic impressions were summarised descriptively. Because the two conditions differed in task content, output configuration, and sequence position, these findings characterise the implemented T2I and T2V task conditions and should not be interpreted as modality-specific effects. By integrating behavioural records with subjective and culturally situated evaluations, this study provides an empirical account of user experiences in the implemented prompt-supported T2I and T2V task conditions and informs future controlled studies of culturally oriented human–GenAI creation.

1. Introduction

In recent years, generative artificial intelligence (GenAI) has rapidly advanced in image and video generation, and has increasingly been integrated into applied design workflows as a tool for human–AI co-creation. In cultural and creative industries, traditional Chinese patterns represent a typical application scenario because they involve both visual aesthetics and culturally meaningful symbolic structures. Compared with conventional image design and video production workflows, GenAI systems lower the technical threshold for visual creation and allow both professional and non-professional users to generate, evaluate, and refine visual outputs through natural language prompts. This shift has moved design interaction from direct parameter manipulation toward intent-driven generation, with prompts serving as an important interface through which users express intentions and interact with generative systems [1,2].
However, applying GenAI to traditional pattern design remains challenging because such tasks require not only visual novelty but also structural controllability and cultural fidelity. Traditional patterns usually follow stable formal principles, including symmetry, compositional order, repetition, and rhythmic arrangement. These principles constitute a formal language that can be described through rule-based methods such as Shape Grammar [3]. By contrast, diffusion-based generation models are effective at producing diverse visual styles, but they may be less stable when strict constraints are imposed on pattern proportions, spatial arrangements, and repeated structures [4]. This limitation becomes more pronounced in text-to-video (T2V) generation, where cross-frame consistency, motion plausibility, and semantic coherence must be maintained over time. Such instability may affect users’ confidence in the generated results; more generally, trust is central to appropriate reliance on automated systems [5]. Therefore, evaluating how GenAI supports controllable, culturally meaningful visual creation involving traditional patterns and related cultural imagery is important for optimizing applied generative design tools [6].
Previous studies have examined GenAI-assisted creation from several perspectives, including prompt engineering, design outcome evaluation, interactive interface support, and differences in user background [7,8,9,10]. Related work has also proposed structured prompting, multi-agent prompt coordination, and intent-alignment mechanisms to reduce barriers for non-expert users and improve generative-system reliability [11,12]. However, most existing studies have focused on general creative tasks, static image generation, or algorithmic prompt optimization. In culturally constrained creation, limited evidence has documented how users behave and report their experiences under specific T2I and T2V task configurations, how observed prompt-use behaviours relate to task records, and how recruited groups with different design backgrounds evaluate the experience.
Specifically, three gaps remain. First, empirical work has rarely examined T2I and T2V task conditions within a shared cultural-creation context while explicitly accounting for differences in task content and output configuration [13,14,15]. Second, although prompt rewriting and optimization are widely discussed [11,12], and reliability problems in generative systems are well documented [16], less is known about how users actually adopt self-written prompts, directly adopted AI suggestions, and manually revised suggestions during iterative creation, how these observed behaviours are associated with total observed task duration and iteration count, and how outcome satisfaction is distributed across the recorded terminal prompting strategies. Third, participant-background comparisons in GenAI-assisted cultural creation remain limited, and claims about professional expertise should be distinguished from comparisons between the recruited design-background and non-design-background groups [17].
In light of the issues mentioned above, this study uses an exploratory single-platform framework to examine two specific task conditions: text-to-image (T2I) and text-to-video (T2V). Participants in the design-background and non-design-background groups completed motif-oriented and culturally related creation tasks on the same platform. Behavioural process measures and subjective evaluations were recorded, and prompt-use categories were derived from observed task-log operations rather than assigned experimentally. The analysis therefore treats T2I and T2V as task conditions encompassing their associated task content, output settings, and sequence positions, and the resulting comparisons are interpreted as exploratory and configuration-specific. Within this scope, the study addresses the following research questions (RQs):
RQ1: What differences in behavioural process measures and subjective experience are observed between the specific T2I and T2V task conditions examined?
RQ2: How do total observed task duration, iteration count, and outcome satisfaction vary across observed terminal prompting strategies?
RQ3: What exploratory differences are observed between the design-background and non-design-background groups within each task configuration?
RQ4: What cultural and aesthetic impressions did participants report for the implemented task conditions?
The study contributes: (1) descriptive evidence on behavioural process measures and subjective experiences under two specific, single-platform T2I and T2V task configurations; (2) exploratory associations of observed terminal prompting strategy with total observed task duration and iteration count, alongside descriptive outcome-satisfaction patterns; and (3) exploratory participant-background comparisons and descriptive reporting of participant-reported cultural and aesthetic impressions. These contributions are intended to inform hypotheses and measurement choices for future controlled studies rather than to establish modality-specific causal effects or validated design recommendations.

2. Related Work

2.1. Prompt-Driven Generative Creation

GenAI is shifting design interaction from direct form manipulation toward semantic and intention-oriented creation, in which users express visual ideas and task requirements through natural-language prompts. Translating abstract aesthetic intentions into clear and actionable descriptions remains an important part of this interaction. Prior work has therefore examined how prompt wording, structure, and specificity can support the communication of design intentions and constraints [18]. Prompting is also related to how users organize creative exploration, consider alternatives, and evaluate generated outcomes. Studies of human-guided AI creative problem-solving and product-design exploration have examined different prompt-based search approaches, goal orientations, and editing strategies [19,20]. Research on T2I interaction has further documented iterative processes involving initial formulation, generation, inspection, diagnosis, and prompt revision [21,22]. Taken together, these studies position prompting as an interaction practice through which users progressively articulate intentions and respond to generated results, although the process may vary with task goals, user decisions, and system affordances.

2.2. Interaction Support Tools and Prompting

Research in human–computer interaction has examined how system-level interaction support can complement users’ prompt-writing activities and provide alternative ways of expressing intentions [23]. Direct-manipulation interfaces can connect natural-language prompts with editable interface elements, allowing users to adjust requests without relying exclusively on unrestricted text entry. Contextual prompt-refinement controls, templates, parameter options, and structured fields can further make possible prompt modifications more explicit during interaction [24]. In T2I creation, paint-like interactions and other visual guidance tools have also been developed to help users steer generation and explore alternatives through direct manipulation [25]. Related interface-design research has examined how the organization of prompt-entry areas, controls, and feedback influences users’ interaction with text-to-image systems [26]. Across these approaches, free-text prompting can be supplemented by templates, parameters, contextual options, or editable candidate descriptions. In this study, GenAI prompt assistance, templated fields, and editable candidate suggestions are collectively described as forms of structured prompt support.

2.3. Intention Alignment and Generative Reliability

Although interactive support can enhance the expression process to some degree, unstable intent alignment remains a common challenge in generative creation. Users often feel that the system fails to grasp their intent or notice significant variations in results caused by minor prompt changes, reflecting ambiguity in prompts and uncertainty in model mappings [27]. To address this, some studies have suggested human-machine collaborative adaptation strategies that improve consistency in intent expression through user feedback and joint correction mechanisms [28]. To improve reliability, one approach boosts the usefulness and task relevance of generated content by integrating feedback from multiple experts and aggregating and filtering various candidate outputs [12]. Another type of research focuses on improving interaction mechanisms by adding ambiguity resolution, step-by-step clarification, and alignment processes during the iterative cycle, helping users gradually refine their design intent and decrease generation uncertainty caused by prompt ambiguity [29]. Meanwhile, to reduce non-expert users’ dependence on manual prompt engineering, recent research has started exploring methods such as multi-agent prompt optimization and role-based collaboration to improve the model-side usability and stability of prompts [30]. Together, these studies suggest that intention alignment and reliability are important considerations when seeking practical, usable, and controllable generative-design outcomes.

2.4. Participant Background and Evaluation Differences

Debate remains regarding whether generative systems can reduce the gap between experts and non-expert users in creative design tasks. Past research shows that large language models can reach performance levels close to those of experts on certain tasks [31]. Comparative studies of AI-generated and human- or expert-produced content or guidance have likewise reported differences in outcomes and evaluations [32,33]. At the same time, appropriate trust calibration remains important when generative systems are used in professional settings [34]. Prior studies suggest that expertise has been associated with differences in design reasoning, including strategy use and convergence behaviour [17], as well as with differences in visual and aesthetic judgment [35,36]. In parallel, expectation-confirmation research indicates that satisfaction judgments may depend not only on perceived performance but also on users’ prior expectations [37,38]. Together, these lines of research suggest that user evaluations in AI-assisted creative processes may reflect multiple participant-related factors beyond operational skill alone.

2.5. T2I and T2V in Culturally Oriented Visual Creation

Compared with static-image generation, video generation additionally requires maintaining coherence across frames and controlling temporal changes [14,39]. In culturally oriented visual creation involving traditional motifs and related imagery, these technical concerns coexist with broader issues concerning cultural representation and the treatment of cultural expressions in AI-supported production [40]. Existing user-experience research has largely examined text-to-image creation and users’ perceptions of generated images [41]. Direct empirical comparisons of implemented T2I and T2V task conditions within the same cultural design framework remain limited, especially when prompting behaviour and participant background are considered together.
In summary, existing research has examined generative creation from various angles, including prompt skills, interactive support, intent alignment, and user-background differences. In visual creation within traditional Chinese cultural contexts, however, limited empirical work has jointly documented behavioural process measures, subjective experience, observed prompting behaviour, and participant-reported cultural and aesthetic impressions across the implemented T2I and T2V task conditions. The present study therefore examines these patterns within one platform-specific, fixed-order setting while treating the resulting comparisons as exploratory and configuration-specific.

3. Materials and Methods

3.1. Participants

Participants were recruited through purposive sampling supplemented by open recruitment. Purposive sampling was used to ensure the inclusion of participants with the background characteristics required by the study [42,43]. Forty-six individuals participated either in preliminary theme screening or in the formal experiment. Their roles were distinguished by the study procedure and by design-related educational or occupational background; these labels do not imply professional certification, homogeneous expertise, or representativeness of the population.
The theme-screening panel (Group P) comprised 10 design postgraduate students and faculty members with 5–10 years of research experience in traditional Chinese decorative motifs (M = 29.8 years, SD = 6.2; 9 females). It participated only in theme screening and was excluded from the formal analyses. The formal experiment included a design-background group (Group E; M = 24.9 years, SD = 3.0; 14 females) and a non-design-background group (Group L; M = 26.6 years, SD = 4.1; 5 females), with 18 participants in each group. Participants were classified as having a design background (Group E: design graduate students or participants with professional design experience) or a non-design background (Group L: participants from mechanical engineering, electronics, or biomedical engineering).
All formal participants were native Chinese speakers with normal or corrected-to-normal vision, no colour blindness or colour-vision deficiency, no self-reported history of neurological disorders, and right-handedness. Each had previously used at least one GenAI tool, but the specific tool types, frequency, duration, and self-perceived proficiency of prior use were not systematically quantified. None had used the Kling platform before the experiment. All participants provided written informed consent and received compensation after completing the study.

3.2. Stimuli

Separate candidate-theme pools were developed for the T2I and T2V task conditions based on pattern catalogues, relevant literature, and design-practice experience. Each pool initially contained 30 candidate themes. Group P rated every candidate on four dimensions: task feasibility (suitability for completion within the limited task period), creative potential (potential to generate diverse solutions), semantic clarity (clarity and usability of the theme description), and cultural representativeness (the extent to which the theme reflected recognisable traditional Chinese motifs or cultural connotations). Ratings used a five-point Likert scale [44], from 1 (very low) to 5 (very high). Within each pool, candidates were ranked using the unweighted mean across the four screening dimensions, and the 10 highest-ranked themes were retained; ties were resolved using the lower pooled rating SD. The screening was used to select suitable themes within each task condition and was not intended to match or establish equivalence between the T2I and T2V theme pools. No fixed numerical threshold or inter-rater-agreement criterion was used. The final pools therefore contained 10 T2I and 10 T2V themes (Table 1).
The thematic subjects in both pools were traditional Chinese decorative motifs, and the pools shared the same broad cultural context, although they were not matched one-to-one. The T2I tasks presented motif subjects in static visual forms, whereas the T2V tasks presented them through motion-oriented prompts. Across the two pools, the subjects included plant, animal, mythological, figural, cloud, and flame motifs. In the T2V pool, motion cues reflected the visual characteristics and cultural associations of the corresponding subjects, such as lotus blooming, koi swimming, a dragon soaring through clouds, and flying apsaras dancing slowly. The two pools therefore represented static and dynamic forms of motif-based visual creation within the same broad cultural context.

3.3. Procedure

All experimental tasks were performed on the Kling GenAI platform [45]. T2I used Image version 2.1, with each generation iteration producing four candidate images by default at a 1:1 aspect ratio and 2K resolution. T2V used Video 2.5 Turbo in high-quality mode, with each generation iteration producing one approximately 5 s, 1080p video by default at a 16:9 aspect ratio. These two task conditions differed in candidate number, output unit, duration, aspect ratio, and evaluation demands. These settings constituted the platform-specific implementation of the two task conditions; accordingly, comparisons between them refer to the implemented T2I and T2V task conditions rather than to generation modality in isolation. The embedded prompt-support module, labelled DeepSeek-R1 in the interface, returned three candidate prompt suggestions per invocation. Participants were free to select any suggestion, combine or revise suggestions, or proceed without adopting one. DeepSeek-R1 has been described by Guo et al. [46]. Kling AI and its prompt-assistance module served as a platform-specific experimental environment. It reduced cross-platform interface and logic differences but did not equalize the two task conditions. The platform did not expose random seeds, low-level sampling parameters, hidden system instructions, or a complete versioned update history. The hardware setup used a 22-inch Dell LCD monitor (Dell Technologies, Round Rock, TX, USA; 60 Hz; maximum brightness 220 cd/m2) for consistent visual presentation.
The experiment comprised a practice session followed by the formal experiment, conducted from 4 to 11 November 2025. Before the formal tasks, all 36 participants completed one T2I and one T2V practice task using topics outside the formal theme pools. The practice tasks followed the same duration and stopping criteria as the formal tasks and familiarized participants with the task procedure and prompt-support interface. Practice-task records were neither analysed nor used to determine the formal themes or generation parameters. In the formal experiment, each participant selected four themes in total—two distinct themes from the T2I pool and two distinct themes from the T2V pool—and completed one task for each selected theme. No theme was repeated by the same participant within either task condition. Theme selection was participant-driven rather than random; consequently, theme frequencies were not expected to be balanced across participants, participant-background groups, or task occurrences. All participants completed the four formal tasks in a fixed alternating T2I–T2V–T2I–T2V sequence so that the two tasks of the same condition were not performed consecutively. The sequence was not fully counterbalanced; therefore, temporal position could not be completely separated from task condition.
The study used a participant-defined stopping rule informed by the concept of satisficing [47], allowing participants to end a task once they considered the current result sufficiently acceptable for submission within the given time. This mechanism was used as a satisficing-based stopping rule rather than as evidence that participants always reached an optimal or fully satisfactory outcome. If they were still unsatisfied at the time limit, the last result they generated was considered the final output. After each task and after completing all tasks, participants were asked to fill out a post-task questionnaire and a comprehensive post-experiment evaluation form, respectively (Figure 1).

3.4. Measures

All questionnaire items were administered in Chinese. The measures comprised behavioural process records and subjective evaluations. Total observed task duration was recorded on site by the experimenter as the elapsed time from the start of each formal task to the participant’s decision to stop. The recorded duration included active user engagement through prompt entry or revision, output inspection and evaluation, and the stopping decision, together with system generation and rendering latency. These components were not timed separately. Iteration count was the number of complete prompt-to-generation cycles. Each iteration was classified based on the recorded prompt-use operation as S1 (self-written prompting: use of the original theme or participant-authored text without adopting a system-generated suggestion), S2 (direct adoption of a system-generated suggestion without manual textual editing), or S3 (adoption or partial adoption of a system-generated suggestion followed by manual addition, deletion, replacement, combination, or rewriting). Prompt-use operations were recorded during the sessions and subsequently classified and reviewed using the predefined operational definitions. For each iteration, the retained platform history contained the submitted generation prompt, generation timestamp, and visible model-version label. These categories represented observed prompt-use behaviours rather than experimentally assigned conditions. For the exploratory task-level analyses, terminal prompting strategy was defined as the strategy recorded in the final iteration of each task. It represented the final recorded operation rather than the complete prompting trajectory and did not necessarily correspond to the iteration that produced the participant-selected output.
Subjective experience was assessed on 7-point scales. The post-task questionnaire, completed immediately after each task, included NASA-TLX-based workload items, flow-related items, and post-task satisfaction items. The six workload items assessed cognitive demand, time pressure, effort, perceived self-performance, frustration, and physical burden [48]. Q4 operationalized perceived self-performance through participants’ satisfaction with their own generated result, asking, “How satisfied were you with the result you generated in this task? (self-performance),” on a scale from 1 (very low) to 7 (very high). Its outcome-focused wording therefore overlaps conceptually with outcome satisfaction. For the workload composite, Q4 was reverse-scored as 8 minus the original response and averaged with Q1, Q2, Q3, Q5, and Q6, so that higher scores indicated greater subjective workload. The separate outcome-satisfaction item retained its original scoring and was analysed separately from the workload composite. Flow-related experience was assessed using eight items covering immersion, challenge-skill fit and sense of control, and enjoyment/reduced self-consciousness [49]. The mean of Q7–Q14 formed the overall flow-related score, with immersion and challenge-skill fit/control additionally examined as subdimensions. Post-task satisfaction items included common items for process satisfaction, outcome satisfaction, and perceived understanding of creative intention. A task-condition-specific subjective item assessed aesthetic quality, harmony, and detail for T2I, whereas a different item assessed narrative coherence and immersion for T2V. The two items were used only for their corresponding task condition; structurally non-applicable values for the other condition were treated as missing and excluded from item-level summaries.
The post-experiment questionnaire was completed once after all four tasks. Separate T2I and T2V evaluations used eight 7-point adjective pairs: complex–simple, confusing–clear, inefficient–efficient, boring–interesting, unattractive–attractive, not novel–novel, bad–excellent, and conventional–innovative [50,51,52]. For each pair, 1 corresponded to the left-hand adjective and 7 to the right-hand adjective. Item-level descriptive statistics were calculated separately for the eight adjective pairs within each task condition. Three task-related items assessed participant-reported T2I traditional-pattern style, T2I cultural essence, and T2V dynamic beauty, while a fourth general item assessed perceived preservation/innovation value. The T2V dynamic-beauty item was treated as an aesthetic impression, and these items reflected participant perceptions rather than objective cultural fidelity, iconographic accuracy, symmetry, repetition, proportion, or output quality. Six categorical items compared the two task conditions in preference, perceived difficulty, interest, creativity stimulation, surprising results, and AI co-creation, with a neutral option available for each item. Two optional open-ended items elicited reasons for task preference and difficulties encountered with the two functions; all responses were reviewed and grouped descriptively by task type and issue.

3.5. Data Analysis

The questionnaires were administered through Wenjuanxing, and the resulting data were analysed using SPSSAU [53,54]. For RQ1, participant-level paired-samples t-tests were used to compare the T2I and T2V task conditions across 9 behavioural and subjective measures. To assess whether the task-condition comparison depended on the inclusion of Q4, the participant-level paired comparison of subjective workload was repeated as a sensitivity analysis using the mean of Q1, Q2, Q3, Q5, and Q6. Mean observed duration per iteration was retained only as a descriptive summary and was not included in the inferential comparisons. The UEQ-inspired adjective pairs were summarised individually. Mean differences were defined as T2V minus T2I. As a descriptive check of the participant-level records, Pearson correlation coefficients were calculated between mean total observed task duration and mean iteration count separately for the T2I and T2V task conditions. For RQ2, iteration-level prompting-strategy frequencies, task-level strategy combinations, the occurrence of strategy changes, and outcome satisfaction by terminal prompting strategy were summarised descriptively. Mixed-effects models for total observed task duration and iteration count were fitted to all 144 task records, with random intercepts for participant and task-condition-specific theme. Fixed effects included task condition, participant background, within-condition occurrence, and terminal prompting strategy. The first three terms served as adjustment covariates; the focal RQ2 contrasts were S1 versus S2 and S3 versus S2. Log-transformed total observed task duration was analysed using a Gaussian linear mixed-effects model with an identity link, fitted by maximum likelihood; fixed effects are reported with 95% Wald confidence intervals. Iteration count was analysed on its original scale using a Gaussian linear mixed-effects model with an identity link and restricted maximum likelihood (REML) estimation; the corresponding coefficients therefore represent adjusted mean differences in the number of iterations. Model convergence and residual behaviour were assessed using the estimation output, residual-versus-fitted plots, normal Q–Q plots, and standardized residuals. Within-condition occurrence indicated whether a task was the first or second occurrence of that condition and was not treated as a correction for the overall fixed sequence. T2I, the design-background group, the first occurrence, and S2 served as reference categories. For RQ3, Welch’s independent-samples t-tests were used for 10 exploratory within-condition comparisons between the design-background and non-design-background groups across five measures. For RQ4, the study-specific cultural and aesthetic items were summarised descriptively at the item level. Condition-specific internal consistency was estimated from participant-level item means, with 95% confidence intervals based on 10,000 bootstrap resamples of participants. Unless otherwise specified, the significance level was set at 0.05. Effect estimates were accompanied by 95% confidence intervals where applicable. Holm correction was applied separately to the nine RQ1 comparisons, the two reference-category contrasts (S1 vs S2 and S3 vs S2) within each behavioural RQ2 model, and the ten RQ3 comparisons.
An a priori sample-size calculation was conducted in G*Power 3.1 for a repeated-measures ANOVA with a within–between interaction [55,56]. Assuming f = 0.25, α = 0.05, power (1 − β) = 0.80, two groups, two repeated-measures levels, a correlation of 0.50 among repeated measures, and ε = 1.00, the minimum required sample size was 34 participants. Thirty-six participants completed the formal experiment. The calculation was used solely to guide recruitment and did not ensure adequate power for every analysis reported in the study; participant-background comparisons and task-level prompting-strategy analyses were therefore treated as exploratory.

4. Results

4.1. Behavioural Results

4.1.1. Total Observed Task Duration and Number of Iterations

Under the implemented task configurations, mean total observed task duration was 5.11 min (SD = 2.45) for T2V and 3.14 min (SD = 1.81) for T2I. The paired mean difference was 1.97 min (95% CI [1.05, 2.90]), t(35) = 4.34, p < 0.001, Holm-adjusted p = 0.001, Cohen’s dz = 0.72 (Figure 2a). No task reached the 15 min ceiling; the maximum observed durations were 11.88 min for T2I and 13.93 min for T2V.
The mean iteration count was 1.85 (SD = 0.95) for T2I and 1.69 (SD = 0.71) for T2V; the paired difference was −0.15 iterations (95% CI [−0.43, 0.12]), t(35) = −1.12, unadjusted p = 0.270, Holm-adjusted p = 1.000. As a descriptive derived measure, mean observed duration per iteration was calculated for each participant by dividing total observed task duration by total iteration count. The resulting values were 1.74 min (SD = 0.79) for T2I and 3.11 min (SD = 1.39) for T2V (Figure 2b).
At the participant level, the cumulative relationship between mean total observed task duration and mean iteration count was summarised descriptively by Pearson’s r (T2I, r = 0.648; T2V, r = 0.687; Figure 2c).

4.1.2. Distribution and Use of Prompt Strategies

At the iteration level, 255 valid iterations were recorded. S2 was used most frequently (129, 50.6%), followed by S3 (98, 38.4%) and S1 (28, 11.0%). At the first iteration of each task, S1, S2, and S3 were used in 17 (11.8%), 94 (65.3%), and 33 (22.9%) tasks, respectively. Across all 144 tasks, S1 appeared at least once in 18 tasks (12.5%), S2 in 105 (72.9%), and S3 in 64 (44.4%). The corresponding T2I- and T2V-specific distributions are reported in Table 2.
Most tasks used a single prompting strategy throughout (105 tasks, 72.9%), with S2 being the most common single-strategy pattern (67 tasks, 46.5%). The remaining 39 tasks (27.1%) involved at least one strategy change, most commonly the S2 + S3 combination (27 tasks, 18.8%) (Table 3). In four tasks, the prompt associated with the participant-selected output differed from the prompt submitted in the final iteration.
Single-iteration tasks accounted for 56 of the 75 S2-ending tasks (74.7%), compared with 15 of the 60 S3-ending tasks (25.0%) and 1 of the 9 S1-ending tasks (11.1%). Mixed-effects models were then fitted to all 144 tasks to examine associations between terminal prompting strategy and behavioural process measures. Random intercepts were specified for participant and task-condition-specific theme, with task condition, participant background, within-condition occurrence, and terminal strategy included as fixed effects. Relative to S2, the estimated task-duration ratios were 2.55 for S1 (95% CI [1.82, 3.56]) and 2.17 for S3 (95% CI [1.83, 2.56]). The log-duration model converged with the participant random-intercept variance estimated at the boundary (<0.001); the task-condition-specific theme and residual variances were 0.034 and 0.216, respectively, and residual screening identified one observation with an absolute standardized residual greater than 3. Tasks ending with S1 and S3 were also associated with 1.048 (SE = 0.368, 95% CI [0.327, 1.769]) and 0.863 (SE = 0.154, 95% CI [0.561, 1.165]) additional iterations, respectively. For the iteration-count model, the estimated random-intercept variances were 0.327 for participant and less than 0.001 for task-condition-specific theme, with a residual variance of 0.516. The S1 category comprised only nine tasks, compared with 75 for S2 and 60 for S3, which limited the precision of the S1 contrasts. All terminal-strategy associations were therefore interpreted as exploratory, particularly those involving S1. Descriptive statistics for total observed task duration and iteration count by terminal prompting strategy are presented in Table 4.

4.2. Subjective Experience Results

Condition-specific internal consistency was estimated from participant-level item means. For T2I and T2V, respectively, α (95% CI) was 0.759 [0.602, 0.837] and 0.829 [0.671, 0.899] for the six-item workload composite; 0.810 [0.696, 0.870] and 0.868 [0.751, 0.925] for the five-item sensitivity composite excluding Q4; and 0.858 [0.702, 0.922] and 0.872 [0.765, 0.920] for the overall flow-related score. The corresponding α values for immersion, challenge–skill fit/control, and enjoyment/reduced self-consciousness were 0.904/0.873, 0.747/0.774, and 0.604/0.613 for T2I/T2V. The two-item enjoyment/reduced-self-consciousness grouping was retained only descriptively.

4.2.1. Subjective Workload and Flow Experience

Under the implemented task configurations, subjective workload was higher in the T2V condition (M = 2.99, SD = 0.89) than in the T2I condition (M = 2.63, SD = 0.73). The paired difference was 0.36 (95% CI [0.11, 0.61]), t(35) = 2.96, unadjusted p = 0.005, Holm-adjusted p = 0.044, Cohen’s dz = 0.49 (Figure 3a). The five-item sensitivity composite excluding Q4 yielded M = 2.99 (SD = 0.99) for T2V and M = 2.51 (SD = 0.78) for T2I, with a paired difference of 0.48 (95% CI [0.22, 0.73]), t(35) = 3.78, and unadjusted p < 0.001.
Overall flow-related scores were M = 5.19 (SD = 0.84) for T2I and M = 5.33 (SD = 0.82) for T2V; the paired difference was 0.14 (95% CI [−0.04, 0.32]), t(35) = 1.63, unadjusted p = 0.113, and Holm-adjusted p = 0.790. For immersion, the paired difference was 0.02 (95% CI [−0.24, 0.28]), t(35) = 0.18, unadjusted p = 0.857, and Holm-adjusted p = 1.000. For challenge–skill/control, the paired difference was 0.18 (95% CI [−0.06, 0.41]), t(35) = 1.53, unadjusted p = 0.134, and Holm-adjusted p = 0.807 (Figure 3b). The enjoyment/reduced-self-consciousness pair was interpreted descriptively only.
For process satisfaction, T2I yielded M = 5.43 (SD = 1.27) and T2V M = 5.53 (SD = 1.14); the paired difference was 0.10 (95% CI [−0.30, 0.50]), t(35) = 0.50, unadjusted p = 0.623, Holm-adjusted p = 1.000 (Figure 3c). For outcome satisfaction, T2I yielded M = 5.28 (SD = 1.40) and T2V M = 5.24 (SD = 1.20); the paired difference was −0.04 (95% CI [−0.57, 0.49]), t(35) = −0.16, unadjusted p = 0.874, Holm-adjusted p = 1.000. Perceived understanding of creative intention was M = 4.89 (SD = 1.39) for T2I and M = 4.78 (SD = 1.38) for T2V; the paired difference was −0.11 (95% CI [−0.57, 0.35]), t(35) = −0.49, unadjusted p = 0.629, Holm-adjusted p = 1.000. The non-equivalent task-condition-specific items were reported only descriptively: perceived aesthetic quality/harmony/detail for T2I, M = 5.15 (SD = 1.24), and perceived narrative coherence/immersion for T2V, M = 5.33 (SD = 1.10).
Outcome satisfaction by terminal prompting strategy was summarised descriptively: S1, M = 5.00 (SD = 2.12); S2, M = 5.51 (SD = 1.30); and S3, M = 4.98 (SD = 1.59).

4.2.2. Overall Experience and Cultural and Aesthetic Impressions

Across the 36 participants, item-level means (SDs) for T2I and T2V, respectively, were: complex–simple, 5.28 (1.32) and 5.03 (1.23); confusing–clear, 5.39 (1.27) and 5.33 (1.35); inefficient–efficient, 5.53 (1.16) and 5.44 (1.38); and boring–interesting, 5.83 (0.97) and 5.75 (1.23). For the remaining pairs, the corresponding T2I and T2V values were: unattractive–attractive, 5.44 (1.21) and 5.69 (1.39); not novel–novel, 5.39 (1.18) and 5.75 (1.30); bad–excellent, 5.28 (1.21) and 5.44 (1.34); and conventional–innovative, 5.61 (1.23) and 5.81 (1.04).
Participant-perceived ratings were M = 4.94 (SD = 1.69) for T2I traditional-pattern style, M = 4.83 (SD = 1.56) for T2I cultural essence, and M = 4.53 (SD = 1.58) for T2V dynamic beauty. The general preservation/innovation item had a mean rating of 5.22 (SD = 1.64).
Across the 36 participants, T2V was selected more often than T2I for overall preference (47.2% vs. 33.3%), perceived difficulty (55.6% vs. 25.0%), interest (63.9% vs. 11.1%), stimulation of creativity (55.6% vs. 27.8%), surprising results (47.2% vs. 41.7%), and perceived co-creation with AI (47.2% vs. 16.7%). The remaining responses indicated no clear preference or that the two task conditions were perceived as similar.
Twenty-three participants provided explanations for their task preferences. Reasons given for preferring T2V included its dynamic and vivid presentation, interest, surprise, narrative richness, and perceived controllability. Reasons given for preferring T2I included a closer perceived fit with expectations or traditional cultural characteristics, shorter perceived generation time, and the perceived maturity of image-generation technology. Nineteen participants also described difficulties involving the communication of creative intent, mismatches between prompts and generated outputs, sensitivity to wording changes, overly elaborate or unclear AI-generated prompt suggestions, and challenges in refining motion or physical plausibility in some T2V outputs.

4.3. Differences by Participant Background

After Holm adjustment across the 10 RQ3 comparisons, no statistically significant participant-background difference was detected in subjective workload under either task condition. The non-design-background group reported higher overall flow-related scores in both T2I and T2V, higher process satisfaction in both conditions, higher outcome satisfaction in T2I, and higher perceived understanding of creative intention in both conditions. The T2V outcome-satisfaction contrast did not remain statistically significant after adjustment (Holm-adjusted p = 0.129) (Figure 4; Table 5).

5. Discussion

5.1. Observed Differences Between the T2I and T2V Task Conditions

Under the fixed sequence, separate theme pools, different default output settings, and participant-defined stopping procedure used in this study, the T2V task condition was associated with longer total observed task duration and higher subjective-workload ratings than the T2I condition. No statistically detectable difference was found in outcome satisfaction. The observed differences therefore characterize the implemented task configurations, in which task content, candidate number, output duration and aspect ratio, sequence position, and condition-specific evaluation content varied jointly. Generation modality cannot be separated from these factors, and relative efficiency cannot be compared at a common quality level.
The two conditions also differed in their implemented task configurations. The T2I tasks centred on static single-frame outputs, whereas the T2V tasks additionally involved motion and temporal change, resulting in different interaction and evaluation characteristics across the two conditions. These configuration differences provide important context for the observed patterns in total observed task duration and workload. Because total observed task duration was recorded at the task level, the relative contribution of individual process stages cannot be distinguished.
Flow, narrative transportation, and immersion have been conceptualized as forms of absorption or focused engagement in different contexts [57,58,59]. These perspectives provide general theoretical background for considering subjective experiences during dynamic creation. The present results, however, do not establish an immersive advantage of T2V. In particular, the enjoyment/reduced-self-consciousness item pair showed low internal consistency, and the task-condition-specific subjective items assessed different content in the two task conditions.
These measures were therefore retained for transparency but were not used as evidence of a compensatory experiential benefit that offset the greater time or workload associated with the T2V tasks.
The participant-defined stopping procedure also affects the interpretation of total observed task duration and satisfaction. Participants ended each task when they considered the current result sufficiently acceptable, rather than when a common externally assessed quality threshold had been reached. Total observed task duration therefore reflects both the unfolding creation process and the participant’s evolving judgment of whether further iteration was worthwhile. This interpretation is consistent with satisficing accounts of decision-making under limited time and effort [47,60], although the present study did not directly measure participants’ acceptance thresholds, perceived opportunity costs, or reasons for stopping. Consequently, the absence of a statistically detectable difference in outcome satisfaction should not be interpreted as evidence that the two task conditions produced equivalent experiences or outcomes.

5.2. Observed Prompting Strategies and Recorded Process Measures

Tasks ending with S2 were associated with shorter total observed task duration and fewer iterations than tasks ending with S3. The S1 category comprised only nine tasks, limiting the precision of estimates involving this category. Outcome satisfaction was summarised descriptively across the terminal-strategy categories. Because the strategies were observed rather than experimentally assigned, and because terminal S2 involved direct adoption without further prompt editing before task completion, these patterns may reflect both prompt-use behaviour and how participants concluded the task. They are therefore interpreted as exploratory process associations.
From an interaction perspective, S2 and S3 represented different ways of using the embedded prompt support: S2 involved direct adoption, whereas S3 involved manual revision of the suggested prompt. This distinction is consistent with prompt-support approaches that structure user input through interface-level guidance [61]. In the recorded tasks, these prompt-use patterns were associated with different profiles of total observed task duration and iteration count, although the contribution of individual interaction stages could not be distinguished.
The descriptive strategy trajectories also showed that 39 tasks involved at least one change in strategy, indicating that participants did not always rely on a single prompting approach throughout the task. Several specific paths were sparsely represented, so these trajectory patterns were retained as descriptive process information rather than compared inferentially. In addition, terminal prompting strategy captured only the final recorded prompt-use operation and did not represent the full prompting trajectory or necessarily the strategy associated with the selected output; the small number of tasks ending with S1 further limited the precision of contrasts involving that category.

5.3. Exploratory Participant-Background Comparisons

After Holm adjustment, several subjective ratings differed between the two recruited participant-background groups, whereas no statistically detectable difference was found in subjective workload under either task condition. The non-design-background group reported higher overall flow-related scores in both T2I and T2V, higher process satisfaction in both conditions, higher outcome satisfaction in T2I, and higher perceived understanding of creative intention in both conditions. The T2V outcome-satisfaction contrast did not remain statistically significant after multiplicity adjustment.
The pattern was selective rather than uniform across the measured dimensions. The between-group contrasts were concentrated in self-reported experiential evaluations, whereas subjective workload showed no statistically detectable group difference in either task condition. This pattern suggests that the recruited groups differed in some aspects of their evaluations of the implemented task experiences, rather than exhibiting a consistent difference across the user-experience measures as a whole.
Participant background was used as an observed grouping characteristic rather than an experimentally manipulated factor. The design-background and non-design-background labels therefore describe the recruited groups rather than directly measuring expertise. The observed pattern may also reflect differences in gender composition and unquantified variation in prior GenAI experience. Accordingly, these findings are interpreted as exploratory, sample-specific participant-background differences rather than as effects attributable specifically to design education, professional design experience, or expertise.

5.4. Cultural and Aesthetic Impressions and Implications

Traditional Chinese patterns rely not only on recognisable cultural symbols but also on formal relationships among symmetry, repetition, contour, proportion, and compositional rhythm. These structural features contribute to the perceptual organisation and cultural recognisability of traditional motifs [62]. In GenAI-assisted creation, users may therefore evaluate an output not only according to its immediate visual appeal, but also according to whether its form, symbolism, and overall style appear consistent with their understanding of traditional patterns.
The T2I traditional-pattern-style and cultural-essence ratings were treated as participant-reported cultural impressions, whereas the T2V dynamic-beauty rating was treated separately as an aesthetic impression of the generated motion. The descriptive ratings were above the midpoint of the 7-point scale for T2I traditional-pattern style, T2I cultural essence, and T2V dynamic beauty. The preservation/innovation item likewise reflected a generally favourable participant view of the potential role of GenAI in culturally oriented creation. Taken together, these ratings describe participants’ cultural and aesthetic impressions of the implemented task conditions.
From a design perspective, these considerations point to several directions for culturally constrained generative systems. Recent work has explored the use of generative AI in cultural-artifact reconstruction and intangible-cultural-heritage safeguarding [63,64]. More controllable generation methods, such as spatial conditional controls, may support targeted visual refinement [65]. Future controlled studies could evaluate whether structural, semantic, and interface-level forms of support help users inspect and refine culturally oriented visual outputs.

5.5. Limitations and Future Work

Several limitations should be noted. First, all tasks were conducted on a single commercial platform using specific model versions, which limits generalisation to other systems. The exact internal platform state of each experimental session cannot be fully reconstructed because some platform-level information was not available to users. Moreover, the T2I and T2V conditions were not fully matched in theme pools, task order, candidate number, output form, duration, aspect ratio, or condition-specific evaluation content. Because the participant-driven theme-selection process was not modeled, the task-condition, participant-background, and terminal-strategy associations may partly reflect the selected theme mix. The fixed T2I–T2V–T2I–T2V sequence did not permit task-condition effects to be separated from temporal-position effects, including possible learning or fatigue. Future studies should use counterbalanced or randomized task orders, matched themes and output settings, a standardized candidate-selection structure, and a common externally defined termination criterion to distinguish these sources of variation more clearly.
Second, total observed task duration combined active user engagement with system generation and rendering latency, which were not recorded separately. Each T2I iteration produced four candidate images, whereas each T2V iteration produced one video, so iteration counts were not equivalent across conditions. Prompting strategy was observed rather than experimentally assigned, and the small number of tasks ending with S1 limited the precision of the related estimates. Future work should use stage-specific process logging and experimentally assigned prompt-support conditions.
Third, the sample was relatively small, young, highly educated, and limited to native Chinese-speaking cultural insiders. The two participant-background groups also differed in gender composition. Although all formal participants had previously used at least one GenAI tool, the specific tool types, frequency, duration, and self-perceived proficiency of prior use were not systematically quantified, and other potentially relevant individual characteristics were not systematically measured. Future studies should recruit larger and more balanced samples and characterize participant background, GenAI familiarity, relevant creative-tool experience, and domain familiarity more systematically.
Finally, the UEQ-inspired items and the cultural and aesthetic impression items were study-specific self-reports rather than independent assessments of usability or cultural fidelity. Future studies should combine user ratings with blinded evaluation based on predefined structural, aesthetic, and cultural criteria.

6. Conclusions

This exploratory study examined behavioural process measures and subjective experiences in prompt-supported T2I and T2V tasks for visual creation in a traditional Chinese cultural context. Within the specific fixed-order task configurations implemented here, differences were observed primarily in total observed task duration and perceived workload. These differences should be interpreted as characteristics of the implemented task conditions rather than as effects attributable specifically to generation modality. Task-condition, participant-background, and terminal-strategy associations may also partly reflect the participant-selected theme mix. Observed prompting behaviours were associated with total observed task duration and iteration count, with outcome satisfaction showing modest descriptive variation across terminal-strategy categories. Participant ratings indicated generally favourable cultural and aesthetic impressions of the implemented task conditions. Overall, the study provides an empirical account of behavioural process measures and subjective experiences observed in the implemented prompt-supported T2I and T2V task conditions. These findings offer a reference for future controlled studies using more comparable task designs, counterbalanced orders, experimentally assigned prompt support, and finer-grained process measures.

Author Contributions

Conceptualization, J.H. and Y.L.; methodology, J.H.; formal analysis, J.H.; investigation, J.H. and W.W.; data curation, J.H. and W.W.; writing—original draft preparation, J.H.; writing—review and editing, Y.L. and F.S.; supervision, Y.L. and F.S.; project administration, F.S.; funding acquisition, Y.L. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the Shandong Province Key Research and Development Program, Soft Science, grant number 2025RZB0704.

Institutional Review Board Statement

The study was conducted in accordance with the Declaration of Helsinki and approved by the Ethics Committee of the School of Basic Medical Sciences, Shandong University (protocol code ECSBMSSDU(L)2024-1-006; approved on 15 June 2024).

Informed Consent Statement

Informed consent was obtained from all subjects involved in the study.

Data Availability Statement

Deidentified data and analytical documentation may be made available by the corresponding authors upon reasonable academic request, subject to applicable ethical and privacy requirements.

Conflicts of Interest

The authors declare no conflicts of interest. The funders had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results.

References

  1. Li, J.; Cao, H.; Lin, L.; Hou, Y.; Zhu, R.; El Ali, A. User Experience Design Professionals’ Perceptions of Generative Artificial Intelligence. In Proceedings of the CHI Conference on Human Factors in Computing Systems; Association for Computing Machinery: New York, NY, USA, 2024; pp. 1–18. [Google Scholar] [CrossRef] [Scilit]
  2. Feng, Y.; Wang, X.; Wong, K.K.; Wang, S.; Lu, Y.; Zhu, M.; Wang, B.; Chen, W. PromptMagician: Interactive Prompt Engineering for Text-to-Image Creation. IEEE Trans. Vis. Comput. Graph. 2024, 30, 295–305. [Google Scholar] [CrossRef] [Scilit]
  3. Stiny, G.; Gips, J. Shape grammars and the generative specification of painting and sculpture. In Information Processing 71: Proceedings of the IFIP Congress 1971; North-Holland Publishing Co.: Amsterdam, The Netherlands, 1972; Volume 2, pp. 1460–1465. [Google Scholar]
  4. Xiong, T.; Wang, N. Exploring dual pathways for traditional pattern innovation: Shape grammar and diffusion models. npj Herit. Sci. 2025, 13, 639. [Google Scholar] [CrossRef] [Scilit]
  5. Lee, J.D.; See, K.A. Trust in Automation: Designing for Appropriate Reliance. Hum. Factors 2004, 46, 50–80. [Google Scholar] [CrossRef] [Scilit]
  6. Ren, K.; Lam, J.F.I. Knowledge graph-driven digital preservation of intangible cultural heritage: A cross-cultural comparative study of Chinese and Western implementation paradigms. Humanit. Soc. Sci. Commun. 2026, 13, 147. [Google Scholar] [CrossRef] [Scilit]
  7. Oppenlaender, J.; Linder, R.; Silvennoinen, J. Prompting AI Art: An Investigation into the Creative Skill of Prompt Engineering. Int. J. Hum. Comput. Interact. 2025, 41, 10207–10229. [Google Scholar] [CrossRef] [Scilit]
  8. Lin, H.; Jiang, X.; Deng, X.; Bian, Z.; Fang, C.; Zhu, Y. Comparing AIGC and traditional idea generation methods: Evaluating their impact on creativity in the product design ideation phase. Think. Ski. Creat. 2024, 54, 101649. [Google Scholar] [CrossRef] [Scilit]
  9. Torricelli, M.; Martino, M.; Baronchelli, A.; Aiello, L.M. The Role of Interface Design on Prompt-mediated Creativity in Generative AI. In Proceedings of the ACM Web Science Conference, Stuttgart, Germany, 21–24 May 2024; Association for Computing Machinery: New York, NY, USA, 2024; pp. 235–240. [Google Scholar] [CrossRef] [Scilit]
  10. Hou, J.; Wang, L.; Wang, G.; Wang, H.J.; Yang, S. The Double-Edged Roles of Generative AI in the Creative Process: Experiments on Design Work. Inf. Syst. Res. 2025, ahead of printing. [Google Scholar] [CrossRef] [Scilit]
  11. Wang, M.; Liu, Y.; Liang, X.; Huang, Y.; Wang, D.; Yang, X.; Shen, S.; Feng, S.; Zhang, X.; Guan, C.; et al. Minstrel: Structural Prompt Generation with Multi-Agents Coordination for Non-AI Experts. arXiv 2024, arXiv:2409.13449. [Google Scholar] [CrossRef] [Scilit]
  12. Long, D.X.; Yen, D.N.; Luu, A.T.; Kawaguchi, K.; Kan, M.Y.; Chen, N.F. Multi-expert Prompting Improves Reliability, Safety and Usefulness of Large Language Models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing; Association for Computational Linguistics: Miami, FL, USA, 2024; pp. 20370–20401. [Google Scholar] [CrossRef] [Scilit]
  13. El Assadi, A. AI vs. human creativity: The impact of text-to-image and text-to-video ads on customer engagement. Electron. Commer. Res. 2026. [Google Scholar] [CrossRef] [Scilit]
  14. Jan, M.T.; Al-Jassani, M.G.; Nadar, M.; Vunnava, E.M.; Chakrapani, V.; Ullah, H.; Khan, A.; Abbas, S.A.; Furht, B. Text-to-video generators: A comprehensive survey. J. Big Data 2025, 12, 253. [Google Scholar] [CrossRef] [Scilit]
  15. Sangamuang, S.; Ariya, P.; Intawong, K.; Khanchai, S.; Puritat, K. Integrating generative AI and the metaverse for cultural heritage: A case study on the preservation of Lamphun Brocade Fabric. Humanit. Soc. Sci. Commun. 2025, 12, 1974. [Google Scholar] [CrossRef] [Scilit]
  16. Huang, L.; Yu, W.; Ma, W.; Zhong, W.; Feng, Z.; Wang, H.; Chen, Q.; Peng, W.; Feng, X.; Qin, B.; et al. A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions. ACM Trans. Inf. Syst. 2025, 43, 1–55. [Google Scholar] [CrossRef] [Scilit]
  17. Cross, N. Expertise in design: An overview. Des. Stud. 2004, 25, 427–441. [Google Scholar] [CrossRef] [Scilit]
  18. Liu, V.; Chilton, L.B. Design Guidelines for Prompt Engineering Text-to-Image Generative Models. In Proceedings of the CHI Conference on Human Factors in Computing Systems; Association for Computing Machinery: New York, NY, USA, 2022; pp. 1–23. [Google Scholar] [CrossRef] [Scilit]
  19. Boussioux, L.; Lane, J.N.; Zhang, M.; Jacimovic, V.; Lakhani, K.R. The Crowdless Future? Generative AI and Creative Problem-Solving. Organ. Sci. 2024, 35, 1589–1607. [Google Scholar] [CrossRef] [Scilit]
  20. Chong, L.; Lo, I.-P.; Rayan, J.; Dow, S.; Ahmed, F.; Lykourentzou, I. Prompting for products: Investigating design space exploration strategies for text-to-image generative models. Des. Sci. 2025, 11, e2. [Google Scholar] [CrossRef] [Scilit]
  21. Brade, S.; Wang, B.; Sousa, M.; Oore, S.; Grossman, T. Promptify: Text-to-Image Generation through Interactive Prompt Exploration with Large Language Models. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology; Association for Computing Machinery: New York, NY, USA, 2023; pp. 1–14. [Google Scholar] [CrossRef] [Scilit]
  22. Mahdavi Goloujeh, A.; Sullivan, A.; Magerko, B. Is It AI or Is It Me? Understanding Users’ Prompt Journey with Text-to-Image Generative AI Tools. In Proceedings of the CHI Conference on Human Factors in Computing Systems; Association for Computing Machinery: New York, NY, USA, 2024; pp. 1–13. [Google Scholar] [CrossRef] [Scilit]
  23. Masson, D.; Malacria, S.; Casiez, G.; Vogel, D. DirectGPT: A Direct Manipulation Interface to Interact with Large Language Models. In Proceedings of the CHI Conference on Human Factors in Computing Systems; Association for Computing Machinery: New York, NY, USA, 2024; pp. 1–16. [Google Scholar] [CrossRef] [Scilit]
  24. Drosos, I.; Williams, J.; Sarkar, A.; Wilson, N.; Rintel, S.; Panda, P. Dynamic Prompt Middleware: Contextual Prompt Refinement Controls for Comprehension Tasks. In Proceedings of the 4th Annual Symposium on Human-Computer Interaction for Work; Association for Computing Machinery: New York, NY, USA, 2025; pp. 1–23. [Google Scholar] [CrossRef] [Scilit]
  25. Chung, J.J.Y.; Adar, E. PromptPaint: Steering Text-to-Image Generation Through Paint Medium-like Interactions. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology; Association for Computing Machinery: New York, NY, USA, 2023; pp. 1–17. [Google Scholar] [CrossRef] [Scilit]
  26. Kim, S.; Ko, T.; Kwon, Y.; Lee, K. Designing interfaces for text-to-image prompt engineering using stable diffusion models: A human-AI interaction approach. In Proceedings of the IASDR 2023: Life-Changing Design, Milan, Italy, 9–13 October 2023. [Google Scholar] [CrossRef] [Scilit]
  27. Reynolds, L.; McDonell, K. Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm. arXiv 2021, arXiv:2102.07350. [Google Scholar] [CrossRef] [Scilit]
  28. Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. Training language models to follow instructions with human feedback. Adv. Neural Inf. Process. Syst. 2022, 35, 27730–27744. [Google Scholar] [CrossRef] [Scilit]
  29. Inan, M.; Sicilia, A.; Xie, A.; Vaduguru, S.; Fried, D.; Alikhani, M. Identifying and interactively refining ambiguous user goals for data visualization code generation. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing; Association for Computational Linguistics: Stroudsburg, PA, USA, 2025; pp. 25246–25263. Available online: https://aclanthology.org/2025.emnlp-main.1283 (accessed on 17 September 2026).
  30. Zhang, Z.; Ge, L.; Li, H.; Zhu, W.; Zhang, C.; Ye, Y. MAPRO: Recasting multi-agent prompt optimization as maximum a posteriori inference. In Findings of the Association for Computational Linguistics: EACL 2026; Association for Computational Linguistics: Rabat, Morocco, 2026; pp. 4458–4480. [Google Scholar] [CrossRef] [Scilit]
  31. Heseltine, M.; Clemm Von Hohenberg, B. Large language models as a substitute for human experts in annotating political text. Res. Politics 2024, 11, 20531680241236239. [Google Scholar] [CrossRef] [Scilit]
  32. Elias, S.; Alshammari, B.S.; Alfraidi, K.N.; Karam, K.M. Rethinking literary creativity in the digital age: A comparative study of human versus AI playwriting. Humanit. Soc. Sci. Commun. 2025, 12, 689. [Google Scholar] [CrossRef] [Scilit]
  33. Krupp, L.; Bley, J.; Gobbi, I.; Geng, A.; Müller, S.; Suh, S.; Moghiseh, A.; Medina, A.C.; Bartsch, V.; Widera, A.; et al. LLM-generated tips rival expert-created tips in helping students answer quantum-computing questions. EPJ Quantum Technol. 2025, 12, 33. [Google Scholar] [CrossRef] [Scilit]
  34. Wischnewski, M.; Krämer, N.; Müller, E. Measuring and Understanding Trust Calibrations for Automated Systems: A Survey of the State-Of-The-Art and Future Directions. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems; Association for Computing Machinery: New York, NY, USA, 2023; pp. 1–16. [Google Scholar] [CrossRef] [Scilit]
  35. Leder, H.; Belke, B.; Oeberst, A.; Augustin, D. A model of aesthetic appreciation and aesthetic judgments. Br. J. Psychol. 2004, 95, 489–508. [Google Scholar] [CrossRef] [Scilit]
  36. Leder, H.; Nadal, M. Ten years of a model of aesthetic appreciation and aesthetic judgments: The aesthetic episode—Developments and challenges in empirical aesthetics. Br. J. Psychol. 2014, 105, 443–464. [Google Scholar] [CrossRef] [Scilit]
  37. Bhattacherjee, A. Understanding Information Systems Continuance: An Expectation-Confirmation Model. MIS Q. 2001, 25, 351–370. [Google Scholar] [CrossRef] [Scilit]
  38. Oliver, R.L. A cognitive model of the antecedents and consequences of satisfaction decisions. J. Mark. Res. 1980, 17, 460–469. [Google Scholar] [CrossRef] [Scilit]
  39. Liu, S.; Zhang, Y.; Li, W.; Lin, Z.; Jia, J. Video-P2P: Video Editing with Cross-Attention Control. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2024; pp. 8599–8608. [Google Scholar] [CrossRef] [Scilit]
  40. UNESCO. Report of the Independent Expert Group on Artificial Intelligence and Culture. Available online: https://www.unesco.org/en/articles/new-expert-report-explores-how-ai-transforming-culture (accessed on 13 December 2025).
  41. Rapp, A.; Di Lodovico, C.; Torrielli, F.; Di Caro, L. How do people experience the images created by generative artificial intelligence? An exploration of people’s perceptions, appraisals, and emotions related to a Gen-AI text-to-image model and its creations. Int. J. Hum.-Comput. Stud. 2025, 193, 103375. [Google Scholar] [CrossRef] [Scilit]
  42. Palinkas, L.A.; Horwitz, S.M.; Green, C.A.; Wisdom, J.P.; Duan, N.; Hoagwood, K. Purposeful Sampling for Qualitative Data Collection and Analysis in Mixed Method Implementation Research. Adm. Policy Ment. Health Ment. Health Serv. Res. 2015, 42, 533–544. [Google Scholar] [CrossRef] [Scilit]
  43. Patton, M.Q. Qualitative Research & Evaluation Methods: Integrating Theory and Practice, 4th ed.; SAGE Publications: Thousand Oaks, CA, USA, 2015. [Google Scholar]
  44. Likert, R. A technique for the measurement of attitudes. Arch. Psychol. 1932, 22, 1–55. [Google Scholar]
  45. Kuaishou Technology. Kling AI. Available online: https://app.klingai.com/cn/ (accessed on 13 December 2025).
  46. Guo, D.; Yang, D.; Zhang, H.; Song, J.; Wang, P.; Zhu, Q.; Xu, R.; Zhang, R.; Ma, S.; Bi, X. DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning. arXiv 2025, arXiv:2501.12948. [Google Scholar] [CrossRef] [Scilit]
  47. Simon, H.A. Rational choice and the structure of the environment. Psychol. Rev. 1956, 63, 129–138. [Google Scholar] [CrossRef] [Scilit]
  48. Hart, S.G.; Staveland, L.E. Development of NASA-TLX (Task Load Index): Results of empirical and theoretical research. In Human Mental Workload; Hancock, P.A., Meshkati, N., Eds.; North-Holland Publishing Co.: Amsterdam, The Netherlands, 1988; pp. 139–183. [Google Scholar]
  49. Rheinberg, F.; Vollmeyer, R.; Engeser, S. Die Erfassung des Flow-Erlebens. In Diagnostik von Motivation und Selbstkonzept; Stiensmeier-Pelster, J., Rheinberg, F., Eds.; Hogrefe Verlag: Gottingen, Germany, 2003; pp. 261–279. [Google Scholar]
  50. Laugwitz, B.; Held, T.; Schrepp, M. Construction and evaluation of a User Experience Questionnaire. In HCI and Usability for Education and Work; Holzinger, A., Ed.; Springer: Berlin/Heidelberg, Germany, 2008; pp. 63–76. [Google Scholar] [CrossRef] [Scilit]
  51. Osgood, C.E.; Suci, G.J.; Tannenbaum, P.H. The Measurement of Meaning; University of Illinois Press: Urbana, IL, USA, 1957. [Google Scholar]
  52. Schrepp, M. User Experience Questionnaire Handbook. Available online: https://www.ueq-online.org/Material/Handbook.pdf (accessed on 5 December 2025).
  53. Changsha Ranxing Information Technology. Wenjuanxing. Available online: https://www.wjx.cn/ (accessed on 5 December 2025).
  54. SPSSAU. Available online: https://spssau.com/ (accessed on 1 December 2025).
  55. Faul, F.; Erdfelder, E.; Lang, A.-G.; Buchner, A. G*Power 3: A flexible statistical power analysis program for the social, behavioral, and biomedical sciences. Behav. Res. Methods 2007, 39, 175–191. [Google Scholar] [CrossRef] [Scilit]
  56. Faul, F.; Erdfelder, E.; Buchner, A.; Lang, A.-G. Statistical power analyses using G*Power 3.1: Tests for correlation and regression analyses. Behav. Res. Methods 2009, 41, 1149–1160. [Google Scholar] [CrossRef] [Scilit]
  57. Csikszentmihalyi, M. Flow: The Psychology of Optimal Experience; Harper & Row: New York, NY, USA, 1990. [Google Scholar]
  58. Green, M.C.; Brock, T.C. The role of transportation in the persuasiveness of public narratives. J. Personal. Soc. Psychol. 2000, 79, 701–721. [Google Scholar] [CrossRef]
  59. Jennett, C.; Cox, A.L.; Cairns, P.; Dhoparee, S.; Epps, A.; Tijs, T.; Walton, A. Measuring and defining the experience of immersion in games. Int. J. Hum. Comput. Stud. 2008, 66, 641–661. [Google Scholar] [CrossRef] [Scilit]
  60. Payne, J.W.; Bettman, J.R.; Luce, M.F. When Time Is Money: Decision Behavior under Opportunity-Cost Time Pressure. Organ. Behav. Hum. Decis. Process. 1996, 66, 131–152. [Google Scholar] [CrossRef] [Scilit]
  61. MacNeil, S.; Tran, A.; Kim, J.; Huang, Z.; Bernstein, S.; Mogil, D. Prompt Middleware: Mapping Prompts for Large Language Models to UI Affordances. arXiv 2023, arXiv:2307.01142. [Google Scholar] [CrossRef] [Scilit]
  62. Westphal-Fitch, G.; Huber, L.; Gómez, J.C.; Fitch, W.T. Production and perception rules underlying visual patterns: Effects of symmetry and hierarchy. Philos. Trans. R. Soc. B Biol. Sci. 2012, 367, 2007–2022. [Google Scholar] [CrossRef] [Scilit]
  63. Altaweel, M.; Khelifi, A.; Zafar, M.H. Using Generative AI for Reconstructing Cultural Artifacts: Examples Using Roman Coins. J. Comput. Appl. Archaeol. 2024, 7, 301–315. [Google Scholar] [CrossRef] [Scilit]
  64. Ming, Y.; Xia, X. Generative AI Technology for Safeguarding Intangible Cultural Heritage: A Systematic Review. In Proceedings of the 2025 2nd International Conference on Artificial Intelligence and Future Education; Association for Computing Machinery: New York, NY, USA, 2025; pp. 7–17. [Google Scholar] [CrossRef] [Scilit]
  65. Zhang, L.; Rao, A.; Agrawala, M. Adding Conditional Control to Text-to-Image Diffusion Models. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision; IEEE: Piscataway, NJ, USA, 2023; pp. 3836–3847. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Overview of the experimental procedure.
Figure 1. Overview of the experimental procedure.
Applsci 16 09429 g001
Figure 2. Recorded task-duration and iteration comparison between T2I and T2V tasks. (a) Mean total observed task duration for T2I and T2V tasks; (b) Mean observed duration per iteration for T2I and T2V tasks; (c) Relationship between mean number of iterations and mean total observed task duration. Error bars represent ±1 SD. Panels (b,c) are presented descriptively. Each point in Panel (c) represents one participant’s condition-specific mean.
Figure 2. Recorded task-duration and iteration comparison between T2I and T2V tasks. (a) Mean total observed task duration for T2I and T2V tasks; (b) Mean observed duration per iteration for T2I and T2V tasks; (c) Relationship between mean number of iterations and mean total observed task duration. Error bars represent ±1 SD. Panels (b,c) are presented descriptively. Each point in Panel (c) represents one participant’s condition-specific mean.
Applsci 16 09429 g002
Figure 3. Subjective experience measures. (a) NASA-TLX-based workload-item mean; (b) overall flow-related score; (c) process satisfaction; (d) outcome satisfaction by terminal prompting strategy. Error bars represent ±1 SD. Panels (ac) show participant-level condition means, whereas Panel (d) shows unadjusted task-level descriptive means across all 144 tasks (S1, n = 9; S2, n = 75; S3, n = 60).
Figure 3. Subjective experience measures. (a) NASA-TLX-based workload-item mean; (b) overall flow-related score; (c) process satisfaction; (d) outcome satisfaction by terminal prompting strategy. Error bars represent ±1 SD. Panels (ac) show participant-level condition means, whereas Panel (d) shows unadjusted task-level descriptive means across all 144 tasks (S1, n = 9; S2, n = 75; S3, n = 60).
Applsci 16 09429 g003
Figure 4. Subjective experience by participant background. (a) NASA-TLX-based workload-item mean; (b) overall flow-related score; (c) process satisfaction. Results are shown for the design-background group (n = 18) and non-design-background group (n = 18). Error bars represent ±1 SD.
Figure 4. Subjective experience by participant background. (a) NASA-TLX-based workload-item mean; (b) overall flow-related score; (c) process satisfaction. Results are shown for the design-background group (n = 18) and non-design-background group (n = 18). Error bars represent ±1 SD.
Applsci 16 09429 g004
Table 1. T2I and T2V theme pools used in the formal experiment.
Table 1. T2I and T2V theme pools used in the formal experiment.
T2I ThemeT2V Theme
1.
Generate a honeysuckle motif in a blue-and-white porcelain style.
1.
A close-up shot of a lotus flower slowly blooming in ink-wash style.
2.
Design a symmetrical peony motif in an embroidery style.
2.
A video of peony petals gently swaying in the wind.
3.
Generate the side profile of a nine-coloured deer in the style of Dunhuang murals.
3.
A slow-motion shot of a bamboo grove swaying in the wind, with falling bamboo leaves.
4.
Design a Vermilion Bird motif in the style of Han dynasty roof tiles.
4.
A dynamic video of a nine-coloured deer running through a forest.
5.
Create a seated qilin motif in a traditional painted style.
5.
A shot of a phoenix flying across the sky with its magnificent tail streaming behind.
6.
Generate a close-up of a Chinese dragon’s head in a paper-cut style.
6.
A koi swimming in water, with ripples spreading across the surface.
7.
Create a static pose of a flying apsaras in the “rebounding pipa” posture.
7.
A video of a Chinese dragon soaring through clouds and mist.
8.
Design a Peking opera facial-mask motif in a cute cartoon-like or trendy stylized form.
8.
A flying apsaras dancing slowly in the air, with ribbons fluttering.
9.
Generate an exquisite auspicious-cloud motif with a jade-carving texture.
9.
A looping video of auspicious-cloud patterns slowly flowing and unfolding in the sky.
10.
Create a set of flame motifs inspired by grotto art.
10.
A dynamic video of flame patterns showing burning and flickering effects.
Table 2. Prompt strategies at the iteration and task levels.
Table 2. Prompt strategies at the iteration and task levels.
Level/ConditionS1, n (%)S2, n (%)S3, n (%)
Iteration level (N = 255)
All iterations28 (11.0%)129 (50.6%)98 (38.4%)
T2I iterations (n = 133)15 (11.3%)69 (51.9%)49 (36.8%)
T2V iterations (n = 122)13 (10.7%)60 (49.2%)49 (40.2%)
Task level—strategy used in the first iteration (N = 144)
All tasks17 (11.8%)94 (65.3%)33 (22.9%)
Task level—tasks using each strategy at least once (N = 144)
All tasks18 (12.5%)105 (72.9%)64 (44.4%)
T2I tasks (n = 72)11 (15.3%)57 (79.2%)30 (41.7%)
T2V tasks (n = 72)7 (9.7%)48 (66.7%)34 (47.2%)
Note. Iteration-level frequencies are descriptive because iterations were nested within tasks and participants. The first-iteration row indicates the initial strategy used in each task. The remaining task-level rows indicate whether each strategy appeared at least once; their percentages may therefore exceed 100%.
Table 3. Combination patterns of prompt strategies at the task level.
Table 3. Combination patterns of prompt strategies at the task level.
Combination Typen Tasks% of All Tasks
Single-strategy tasks (n = 105, 72.9%)
Only S164.2%
Only S26746.5%
Only S33222.2%
Mixed-strategy tasks (n = 39, 27.1%)
S1 + S274.9%
S1 + S310.7%
S2 + S32718.8%
S1 + S2 + S342.8%
Note. Only S1/S2/S3 indicate tasks that consistently used a single prompt strategy across all iterations. S1 + S2, S2 + S3, etc., signify tasks that combine multiple strategies within the same generative process.
Table 4. Recorded task measures by terminal prompting strategy.
Table 4. Recorded task measures by terminal prompting strategy.
Terminal StrategynTotal Observed Task Duration, M (SD), minIteration Count, M (SD)
S196.29 (2.53)2.56 (1.13)
S2752.70 (2.06)1.36 (0.71)
S3605.58 (2.97)2.17 (1.09)
Note. M and SD are reported in minutes for total observed task duration and in counts for iterations. Values are unadjusted task-level descriptive statistics rather than model-adjusted estimates. Terminal-strategy counts by task condition were T2I: S1 (n = 4), S2 (n = 39), and S3 (n = 29); and T2V: S1 (n = 5), S2 (n = 36), and S3 (n = 31).
Table 5. Exploratory participant-background comparisons within each task condition.
Table 5. Exploratory participant-background comparisons within each task condition.
MeasureTaskGroup E, M (SD) Group L, M (SD) Δ [95% CI]davWelch’s t (df)pHolm p
Subjective workloadT2I2.78 (0.60)2.47 (0.83)0.31 [−0.18, 0.80]0.421.270.2140.231
T2V3.22 (0.85)2.75 (0.88)0.47 [−0.12, 1.06]0.541.620.1160.231
Overall flow-related scoreT2I4.84 (0.92)5.54 (0.60)−0.69 [−1.22, −0.17]−0.89−2.680.0110.046
T2V4.97 (0.86)5.69 (0.59)−0.72 [−1.22, −0.22]−0.98−2.930.0060.036
Process satisfactionT2I4.67 (1.33)6.19 (0.57)−1.53 [−2.22, −0.83]−1.49−4.48<0.001<0.001
T2V4.97 (1.30)6.08 (0.58)−1.11 [−1.79, −0.43]−1.11−3.320.0020.015
Outcome satisfactionT2I4.56 (1.42)6.00 (0.94)−1.44 [−2.26, −0.63]−1.20−3.590.0010.009
T2V4.83 (1.18)5.64 (1.12)−0.81 [−1.58, −0.03]−0.70−2.100.0430.129
Perceived understandingT2I4.17 (1.36)5.61 (1.02)−1.44 [−2.26, −0.63]−1.20−3.600.0010.009
T2V4.19 (1.28)5.36 (1.25)−1.17 [−2.02, −0.31]−0.92−2.770.0090.046
Note. Group E = design-background group (n = 18); Group L = non-design-background group (n = 18). Δ = Group E − Group L. Effect sizes are reported as dav, calculated using the square root of the average of the two group variances. Holm correction was applied across the 10 RQ3 comparisons. Higher workload scores indicate greater subjective workload, with Q4 reverse-scored as 8 minus the original response.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Huang, J.; Wang, W.; Liu, Y.; Song, F. User Experiences with Prompt-Supported Text-to-Image and Text-to-Video Task Conditions for Visual Creation in a Traditional Chinese Cultural Context: An Exploratory Study. Appl. Sci. 2026, 16, 9429. https://doi.org/10.3390/app16199429

AMA Style

Huang J, Wang W, Liu Y, Song F. User Experiences with Prompt-Supported Text-to-Image and Text-to-Video Task Conditions for Visual Creation in a Traditional Chinese Cultural Context: An Exploratory Study. Applied Sciences. 2026; 16(19):9429. https://doi.org/10.3390/app16199429

Chicago/Turabian Style

Huang, Jinming, Weihao Wang, Yan Liu, and Fanghao Song. 2026. "User Experiences with Prompt-Supported Text-to-Image and Text-to-Video Task Conditions for Visual Creation in a Traditional Chinese Cultural Context: An Exploratory Study" Applied Sciences 16, no. 19: 9429. https://doi.org/10.3390/app16199429

APA Style

Huang, J., Wang, W., Liu, Y., & Song, F. (2026). User Experiences with Prompt-Supported Text-to-Image and Text-to-Video Task Conditions for Visual Creation in a Traditional Chinese Cultural Context: An Exploratory Study. Applied Sciences, 16(19), 9429. https://doi.org/10.3390/app16199429

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Article metric data becomes available approximately 24 hours after publication online.
Back to TopTop