Abstract
This study investigated whether humans and generative Large Language Models (LLMs) exhibit similar performance in divergent ideation but diverge in convergent selection. To address the critical oversight in current AI creativity research, which predominantly focuses on generative output, this study introduces the original conceptual framework of ‘Selection Alignment’ and a ‘novel dual-phase experimental protocol.’ This research transcends traditional generation-centric evaluations to establish a new paradigm for assessing the evaluative stage of creativity. A controlled experiment involved 240 design professionals (120 idea generators, 120 independent selectors) and two LLM agents (GPT-4o, Gemini 1.5 Pro). Participants and LLMs responded to identical divergent prompts, including 10 Alternative Uses Task-style prompts and 10 design problems. Both humans and LLMs generated candidate idea pools, then performed convergent selection by choosing the top five items per prompt. Idea generation was evaluated based on Fluency, Flexibility, and Semantic Breadth. Selection outcomes were compared using top-5 overlap rates derived from semantic clustering. The results indicated near-parity in generation metrics, showing no statistically significant differences between human and AI outputs. However, a substantial divergence was observed in convergent selection: the mean human–AI top-5 overlap was 19.2% for Model-A and 22.4% for Model-B, both significantly below permutation-based chance levels (null mean overlap ≈ 35%). AI selections were strongly predicted by embedding- and probability-based metrics, while human choices were better predicted by context- and experience-based criteria, highlighting a fundamental mechanistic divide. This suggests that convergent selection amplifies human–AI divergence, carrying significant implications for designing co-creative interfaces that integrate human experience into AI’s selection mechanisms.
1. Introduction
Creative cognition is widely understood as a two-phase process: (1) divergent ideation, where candidate ideas are generated through associative exploration, and (2) convergent selection, where these ideas are synthesized, evaluated, and refined into final solutions [1]. Recent advancements in Large Language Models (LLMs) have demonstrated impressive capabilities in divergent ideation, often matching or even surpassing human performance on standardized tests of fluency and diversity. This has spurred considerable debate regarding the nature of computational creativity and the unique contributions of human intelligence.
However, while much attention has been placed on AI’s generative prowess, a critical question for design theory and practice is the extent to which this demonstrated parity in ideation extends to the convergent selection phase. This phase is heavily influenced by human intuition, domain-specific tacit knowledge, and contextual understanding.
Existing methods predominantly focus on evaluating AI’s generative capabilities, often overlooking the nuanced and context-dependent nature of convergent selection, a critical phase where practical feasibility, human values, and tacit knowledge are paramount. Current research provides ample evidence of AI’s proficiency in divergent ideation [2]—matching or even surpassing human performance in generating a wide array of ideas. Yet, this body of work falls short in addressing whether this generative parity translates into the evaluative rigor required for selecting truly impactful and human-aligned solutions.
Specifically, the limitations of existing methods include: scalability challenges in qualitative assessment, a lack of interpretability in AI’s selection mechanisms [3], implicit assumptions of convergent alignment, and insufficient consideration of contextual and experiential factors crucial for real-world design problems. These limitations highlight a significant research gap in understanding the full spectrum of human–AI creative collaboration.
To bridge this gap and address these structural limitations, the primary objective of this study is to introduce methodological innovations in three key areas, ensuring a rigorous evaluation of the selection phase:
- First, a novel dual-phase experimental protocol for ‘Selection Alignment’: We move beyond traditional generation-centric evaluations by systematically analyzing Human–AI divergence specifically in the convergent selection phase.
- Second, quantitative modeling and cross-prediction analysis: We implement a framework that statistically demonstrates the fundamental mechanistic divide between human and AI selection processes, rather than merely observing disparate outcomes.
- Third, the Construction of an ‘Idealization Vector’: Grounded in multi-level, multi-dimensional human expert judgments, this vector explicitly anchors AI’s internal evaluation criteria to rigorous human domain expertise.
By systematically addressing these objectives, this study provides a critical framework for the design of effective mixed-initiative co-creative systems. We posit a focused hypothesis: while humans and contemporary generative AIs may exhibit comparable association-based generation, they will systematically diverge during convergent selection. This divergence is anticipated because human selection is profoundly shaped by contextual factors, experiential biases, and pragmatic constraints [4], whereas AI’s decision-making processes are predominantly statistical and pattern-based. The structural relationship of this human–AI dual creative process is conceptualized in the Squid Model (SQM) shown in Figure 1.
Figure 1.
Squid Model (SQM)—a conceptual model of the human–AI dual creative process. This model highlights the importance of selection sorting in bridging the gap between AI’s statistical predictions and human contextual judgments.
2. Background and Rationale
The dual-process theory of creativity distinguishes between divergent thinking, which emphasizes the quantity and diversity of ideas, and convergent thinking, which focuses on evaluating, refining, and selecting the most promising ideas. Metrics such as Fluency (number of ideas), Flexibility (number of distinct categories), and Semantic Breadth (diversity of ideas in a semantic space) are commonly used to quantify divergent output [5]. Modern generative models [6], particularly LLMs [7], have shown remarkable performance on these measures, frequently equaling or exceeding average human capabilities.
However, the internal selection mechanisms of generative AI systems fundamentally differ from human cognitive processes. AI models typically rank and select candidate outputs based on factors such as token-level log probabilities, embedding proximities, and model confidence scores. In stark contrast, human evaluative criteria are multifaceted, incorporating deep contextual understanding, prior domain expertise, implicit biases, ethical considerations, and pragmatic constraints related to implementation, user needs, and market viability. These human-specific criteria often involve an intuitive synthesis that transcends purely statistical correlations.
Prior research in computational creativity has predominantly focused on the generative capabilities of AI, often leading to an overemphasis on divergent ideation benchmarks. While studies are actively exploring the ability of LLMs to evaluate creative outputs, there remains a significant gap in understanding whether the observed human–AI parity in generation translates to the critical selection and convergent stage. This framework consists of two stages: divergent generation and convergent selection. Demonstrating generation parity within this framework is critical because it isolates the ‘selection gap’ as a distinct cognitive divergence, proving that superior generative output does not guarantee human-aligned evaluative logic.
If selection systematically diverges, then relying solely on generation-based evaluations would provide an incomplete and potentially misleading assessment of human–AI co-creative alignment. Therefore, ‘good’ ideas are not simply quantitatively superior, but are directly tied to the selection stage criteria that human designers prioritize, such as contextual relevance, practical feasibility, human intuition, domain-specific tacit knowledge, and contextual understanding. This divergence stems from AI’s inherent lack of emotional depth and personal experience, which are crucial for evaluating subjective, contextual, and business-relevant design criteria.
This would necessitate a re-evaluation of how we design mixed-initiative systems, leading to specific design principles to foster alignment. These practical guidelines are grounded in theoretical foundations concerning human–AI interaction and co-creativity frameworks, such as the interaction framework for studying co-creative AI and the three-level framework for effective collaboration between human and AI in human-centered AI-co-creation. These frameworks provide a structured understanding of interaction dynamics, including communication, turn-taking, and the balance of agency between humans and AI, which are crucial for developing co-creative interfaces that effectively integrate human experience into AI’s selection mechanisms:
- Transparency in AI’s Ranking Signals to allow humans to interpret and challenge the statistical basis of AI selection.
- Mechanisms for Explicit Human Contextual Control to allow users to inject domain-specific tacit knowledge and pragmatic constraints into the selection process.
- Interface Affordances that Reconcile Divergence to bridge the gap between AI’s data-driven evaluations and humans’ experience-driven choices.
This study directly addresses this gap by offering a mechanistic comparison that dissects human–AI behavior at both generation and selection phases, and by modeling the underlying mechanisms of selection for both agent types.
3. Research Hypotheses
To systematically investigate the potential divergence between human and AI creativity, we formulated three hierarchical hypotheses that progress from generative capability (H1) to selection outcomes (H2) and finally to the underlying decision-making mechanisms (H3). This structure is designed to deconstruct the creative process into logical stages, ensuring that any observed divergence in selection is not merely a byproduct of generation differences but is rooted in distinct evaluative architectures.
H1 (Generation Parity):
Under matched prompts, humans and AIs exhibit statistically equivalent distributions across generation metrics, specifically fluency, flexibility, and semantic breadth.
H2 (Selection Divergence):
Human and AI agents show a statistically significant divergence in selection outcomes, manifested as a low overlap rate in their respective prioritized idea lists.
H3 (Distinct Mechanisms):
The divergence in selection outcomes is driven by distinct evaluative logics, where AI selection is positively correlated with statistical pattern metrics, while human selection is correlated with contextual and experiential metrics.
4. Methodology
This research utilized a preregistered, two-phase experimental protocol to compare generation and selection behaviors between humans and AI, incorporating several rigorous methodological controls to ensure robustness. The methodology emphasizes matched candidate pools, reproducible preprocessing, and a pre-specified analysis pipeline to ensure robustness and transparency. To clarify the methodological distinctions of this study, Table 1 summarizes our key innovations in contrast to conventional baseline techniques, with pointers to the relevant subsections.
Table 1.
Distinction between conventional approaches and the methodological innovations of this study.
4.1. Experimental Design Overview
The purpose of this section is to outline the dual-phase experimental framework, which isolates convergent selection from divergent generation to identify the specific stage where human–AI divergence occurs. The experiment establishes a controlled environment that allows for a direct, comparative analysis of human and AI performance across both creative phases.
The protocol comprises two sequential phases:
- Phase A (Divergent Generation): Both humans and LLMs produce candidate ideas in response to identical prompts.
- Phase B (Convergent Selection): Pooled, deduplicated sets of candidate ideas are presented to independent human selectors and to LLM selection procedures to identify the top five items per prompt.
This design facilitates the quantification of generation metrics (Fluency, Flexibility, Semantic Breadth), selection outcomes (top-5 lists, overlap rates), and an in-depth analysis of the underlying selection mechanisms (AI internal scores versus human-rated criteria). A schematic of this experimental design is provided in Figure 2.
Figure 2.
Pipeline flowchart: The arrows indicate the linear procedural flow from ideation to final analysis. Capital letters denote the specific experimental modules: (A) Generation Phase, where initial ideas are produced by both Humans and LLMs; (B) Pooling, the aggregation of all generated ideas into a unified set; (C) Preprocessing and Embedding, involving normalization, vector embedding, and other pre-analysis steps; (D) Deduplication, using Hierarchical Agglomerative Clustering (HAC) with a 0.15 cut distance to remove near-duplicates; (E) Idea Selection and Quantitative Analysis, where independent human/LLM ranking occurs and the mechanistic divide is modeled; and (F1, F2) Final Metrics, representing generation performance and selection/overlap analysis results.
To ensure empirical rigor, our experimental design included three key technical safeguards:
- Parameter Sensitivity (Deduplication): We employed Hierarchical Agglomerative Clustering (HAC) with a precisely defined cut distance threshold (, corresponding to a cosine similarity of 0.85). This threshold was rigorously justified through sensitivity analyses, confirming the stability of our flexibility metric across a range of 0.12–0.18.
- Robustness to Noise (Double-Blind): In the selection phase, we implemented a double-blind procedure where human selectors were unaware of the provenance of any idea (human vs. AI). This minimized potential cognitive bias, ensuring that human judgments were based solely on the merit of the ideas.
- Behavior under Edge Cases (Non-parametric Inference): To evaluate selection divergence reliably, we used a permutation-based chance-level overlap calculation. For each prompt, 10,000 random reassignments were conducted to estimate a null mean overlap, providing a robust baseline even in cases where idea distributions were highly skewed.
Sensitivity analyses and cross-prediction tests are integrated to rigorously examine human–AI divergence.
- STEP 1: Generation and Preprocessing
- Generation (Humans, LLMs): Initial phase where ideas are produced.
- Pooling: Aggregation of all generated ideas.
- Preprocessing and Embedding: Normalization, embedding into vectors, and other pre-analysis steps.
- Deduplication (HAC, ): Hierarchical Agglomerative Clustering with a cut distance of 0.15 to remove near-duplicates.
- STEP 2: Idea Selection
- Selection (Human selectors; LLM ranking): Independent selection of top ideas by humans and LLMs.
- STEP 3: Quantitative Analysis
- Analyses (F1-Generation Metrics Analysis, F2-Selection/Overlap Analysis): Final stage for quantitative analysis of results.
4.2. Experimental Subjects: Humans and AI Agents
This subsection details the selection criteria and expertise of the human participants and the specific configurations of the LLM agents, ensuring that the human baseline reflects expert-level intuition and the AI baseline represents state-of-the-art generative capabilities. To begin, we ensured the objectivity of our participants by recruiting N = 240 design professionals (120 idea generators, 120 independent selectors). The recruitment criteria and group justifications are as follows:
- Expertise Criteria: The objective of setting high expertise standards is to ensure the human baseline captures nuanced professional judgment rather than mere novice intuition. All participants were required to have a minimum of two years of full-time professional experience or be advanced design graduate-level researchers (Master’s or Ph.D. candidates) with equivalent professional backgrounds. This deliberate choice grounded our investigations within a field where contextual understanding and practical applicability are paramount, directly addressing the nuanced aspects of human judgment that differ from AI’s statistical processes.
- Statistical Power and Sample Size Justification: This justification aims to demonstrate that our study is sufficiently powered to detect practically significant effects. The sample size of N = 120 for both groups was determined through a robust statistical power analysis. Adhering to conventional power levels of 0.80 to 0.90 and an alpha of 0.05 [8], we anticipated a medium effect size (Cohen’s d of 0.5) for design judgment tasks. This N = 120 per group provides enhanced precision to detect even small yet practically significant effects.
- Human Generators and Selectors: To maintain independent variables, we recruited 120 generators and 120 separate selectors. The generators completed randomized prompts during Phase A, while the selectors (who did not participate in generation) performed the Phase B selections.
- Ethical Considerations: The study was conducted in accordance with the Declaration of Helsinki. The project was evaluated for ethical compliance by the School of Art and Design at Zhejiang A&F University and deemed exempt from formal IRB review based on minimal risk and participant anonymity. The informed consent form has been added to the Supplementary Information to ensure ethical transparency.
AI Agents and Baseline Rationale
To increase the generalizability of our results, we selected two high-performance modern LLMs: GPT-4o (Model-A) and Gemini 1.5 Pro (Model-B). These models were chosen as representative state-of-the-art (SOTA) baselines because they currently define the upper bound of generative parity with human creative cognition. Furthermore, these models provide stable API support, allowing for the extraction of internal ranking signals (e.g., log-probabilities), which is essential for our mechanistic analysis. A comparative overview of the AI agents and human participants is provided in Table 2.
Table 2.
Participants and AI agents.
4.3. Materials—Prompt Corpus
A corpus of 20 distinct prompts was curated to evaluate the generative and evaluative capabilities of both humans and AI agents. This set consists of 10 Alternative Uses Task (AUT)-style prompts and 10 short design-problem prompts. This dual-category approach provides the foundational data for our “Selection Alignment” analysis, ensuring that the prompts challenge both humans and AI across various constraint-driven problem-solving scenarios.
The AUT prompts include common objects such as a brick, a paperclip, and a newspaper, while the design prompts address contemporary challenges such as low-cost faucet filters for renters and modular phone accessories for health sensing. To ensure full transparency and facilitate research replication, the complete list of all 20 prompts, including their specific instructions and task constraints, is detailed in Appendix A (Table A1).
4.4. Procedure—Generation Phase
The purpose of this section is to provide a step-by-step account of the experimental flow to ensure the transparency and reproducibility of the data collection and preprocessing stages. It details the initial phase of the procedure, corresponding to STEP 1 (Generation and Preprocessing) in Figure 2, encompassing the workflow from initial ideation to the final refined pool for selection.
4.4.1. AI Generation Procedure and Parameter Selection (Phase A)
The objective of this sub-stage is to establish a robust AI generation baseline that mirrors human creative constraints while maximizing output variance. To ensure a rigorous comparison, each LLM was prompted with instructions mirroring those given to human participants. As part of the generative stage, the models were sampled 10 times per prompt using a range of decoding parameters to optimize for creative diversity, as detailed in Table 3.
Table 3.
AI parameters and rationale for selection.
4.4.2. Pooling and Preprocessing (Phases B–D)
This sub-stage aims to unify disparate data sources into a refined, high-quality dataset suitable for comparative analysis. Following the generation phase, the outputs from human participants and both LLM agents were processed through the following systematic sequence:
- Aggregation (Phase B): All generated outputs were aggregated into a single candidate pool for each prompt.
- Vector Embedding (Phase C): The pool underwent preprocessing and vector embedding to facilitate mathematical analysis of the idea space.
- Deduplication (Phase D): To ensure a high-quality dataset, we performed deduplication using Hierarchical Agglomerative Clustering (HAC) with a cut distance (). This process removed near-duplicate entries, ensuring the integrity of the selection phase.
This refined, unified pool was then presented identically to all selectors—both human and AI—establishing a common baseline that allows for a fair and direct comparison of selection behaviors across all agents. For a comprehensive visual overview of this operational sequence, please refer to STEP 1 in Figure 2.
4.5. Text Preprocessing, Embedding, and Deduplication Pipeline
The objective of this subsection is to describe the technical pipeline used to normalize and deduplicate the generated idea pools, establishing a common, unbiased baseline for the subsequent selection phase.
- Normalization: A deterministic text normalization procedure will be applied, including lowercasing, removal of non-essential punctuation (while preserving content-critical punctuation), whitespace normalization, and canonicalization of common abbreviations
- Embedding model: Each candidate idea will be encoded into a fixed-length dense vector using a pre-trained sentence encoder from the SBERT family [9]. All embeddings are generated using the all-MiniLM-L6-v2 model [10]. The embedding generation process for a dataset of approximately 10,000 ideas took an average of 3 h on a single NVIDIA RTX 4080 SUPER GPU.
- Similarity and distance: A pairwise cosine similarity matrix (S) will be computed for all embedded ideas, from which a distance matrix (D = 1 − S) will be derived for clustering [11].
- Clustering and medoid selection: Hierarchical Agglomerative Clustering with average linkage will be employed to group near-duplicate ideas [12], utilizing Cosine Distance to map the semantic proximity between embeddings [13]. The resultant clusters serve two distinct purposes: (1) measuring Flexibility by quantifying the variety of semantic categories, and (2) supporting Semantic Breadth measurement by providing the foundation for calculating mean pairwise distance among unique item embeddings. A cut distance threshold (, equivalent to a cosine similarity ≥ 0.85) will define cluster boundaries, a value chosen to distinguish between nuanced variations and distinct conceptual shifts. The clustering and medoid selection process for 10,000 ideas is typically completed within 30 min on a 12-core CPU. The number of unique clusters () will be used as the direct measure of Flexibility for both human and AI agents. Within each cluster, the medoid (the item minimizing its mean distance to all other items in the cluster) will be selected as the cluster representative to maintain the interpretive integrity of the original data. As a sensitivity check, alternative clustering methods like DBSCAN [14] (tuned with eps0.15–0.25, min_samples = 1) will be explored to ensure the robustness of the findings against different density-based grouping assumptions.
- Threshold justification and sensitivity: The 0.80–0.90 cosine similarity range is a standard practice for short-text deduplication. We have preset 0.85 for balance and will report the robustness of our findings across thresholds of 0.80, 0.85, and 0.90.
4.6. Selection Phase—Humans and LLMs
This section details the convergent process corresponding to STEP 2 (Idea Selection) in Figure 2, where both human selectors and LLM agents identify the optimal ideas from the pooled candidate set (Phase E). This phase enables a mechanistic comparison of their underlying selection criteria. To ensure the psychological independence of the items in this pool and avoid selection bias from redundant content, we employed Hierarchical Agglomerative Clustering (HAC) for deduplication. A fixed cut distance of (cosine distance) was applied based on a sensitivity analysis conducted over a range of 0.10 to 0.20. This specific threshold was chosen because it effectively collapsed lexical redundancies (e.g., minor paraphrasing) while preserving distinct semantic concepts, thereby maintaining the granularity required for a valid comparison between human and AI selection logic. Empirical observation during the sensitivity analysis confirmed that served as the optimal elbow point to maximize diversity without losing unique conceptual kernels.
4.6.1. Human Selection
- Selection Participant Rationale: Rather than prioritizing a large sample size, which can introduce ‘average-preference noise’ in high-level creative evaluation, this study focused on ‘expert depth.’ The human selectors consisted of senior design professionals with an average of 15.4 years of experience, ensuring that the baseline for ‘expert-referenced’ selection was rooted in stabilized professional intuition—serving as a human-centric benchmark for evaluating selection alignment rather than an absolute measure of objective truth.
- Selection Procedure: Each human selector independently reviewed the deduplicated candidate pool for their assigned prompts and identified their top-5 preferred items. To mitigate social conformity and individual cognitive bias, this selection was conducted in a double-blind environment where selectors were unaware of the ideas’ origins. To maintain the integrity of the double-blind protocol, selectors were not informed of the specific origin (human vs. AI) of any individual idea during the selection and rating process. While the qualitative inquiry phase retrospectively asked participants to consider human-specific insights, this open-ended question was only presented after the primary selection, and Likert-scale ratings were completed. This ensures that their evaluation of specific ideas was based on intrinsic merit and professional judgment rather than a reactive bias for or against AI-generated content.
- Criterion Elicitation: Following selection, they will complete a brief questionnaire, rating the influence of various selection criteria on a 1–5 Likert scale: Context Relevance, Prior Experience Influence, Feasibility, Novelty, and Intuitive Coherence. Additionally, they will provide an open-text rationale for their top-ranked choice.
- Ensuring Robust Judgment: To ensure robust human judgment, all ideas will be rated by at least two independent human selectors. Furthermore, to establish the ‘Human-Consensus Baseline’ (), which serves as a comparative reference for AI alignment, a strict consensus protocol was applied: an idea was only included in the consensus baseline if it was mutually selected by a high threshold of experts (at least 70% agreement). This approach acknowledges that while individual expert judgment may vary, a shared professional standard provides a robust point of comparison for analyzing the degree of ‘Selection Alignment’ in co-creative systems. This ensures that the resulting criteria represent a collective professional standard that transcends individual subjective variance. Inter-rater reliability (IRR) will be rigorously assessed using Intraclass Correlation Coefficients (ICCs) for Likert scale ratings [15] and Cohen’s Kappa for categorical judgments [16] (e.g., in open-text rationale coding).
4.6.2. LLM Selection
The LLM agents rank the candidate items using computed internal metrics and select the top five accordingly.
Calculation of Metrics: For each candidate item in the pool, the LLM will calculate the following:
- Item-to-Ideal Cosine Similarity: The cosine similarity between the item’s embedding vector and a “human-centric reference” vector (formerly termed ‘idealization’ vector).
- Mean Token Log-Probability: This measures the internal likelihood of an idea given the LLM’s language model.
- Ranking Score: Composite scores combine these statistical signals. The LLM’s top-5 items are then determined by the highest composite scores.
Reference Vector Construction: The reference vector is crucial for evaluating the suitability of ideas from a human-aligned perspective. It is constructed by combining the initial prompt text with a set of exemplar responses, functioning as an internal criterion point that embodies consensus-based professional standards. Rather than relying on a broad but potentially diluted sample, this study prioritized ‘expert depth’ by utilizing a panel of 10 senior design professionals (average 15.4 years of experience). This ensures the vector is anchored in stabilized professional intuition—providing a benchmark for human-centric design values—rather than mere statistical popularity.
In this study, the reference vector is constructed based on a ‘pool of high-quality potential ideas’ selected by ‘a panel of at least 10 design experts’ applying ‘multi-level, multi-dimensional evaluation criteria’ such as ‘contextual relevance, feasibility, and domain appropriateness’, and not merely a statistical average or a simple keyword-based ‘generation of ideal words’ [17]. This distinct approach ensures that AI’s internal evaluation criteria are explicitly compared against rigorous human domain expertise, providing a benchmark for analyzing selection alignment.
Exemplar Responses: These responses are meticulously pre-selected through a rigorous multi-stage evaluation process by an independent panel of N = 10 design experts. Experts independently evaluate a comprehensive set of high-quality candidate ideas based on predefined criteria (e.g., context relevance, feasibility, domain appropriateness). To ensure the robust quality of the human-consensus baseline without over-reliance on post hoc statistical corrections, a strict consensus protocol was applied: an idea was only included in the exemplar pool if it was mutually selected by a high threshold of the panel (at least 70% agreement). This stringent filtering ensures that the reference vector represents a concentrated ‘professional consensus’ that transcends individual subjective variance. Only responses within the top 25% of the score distribution are included in the initial selection pool. The final set of exemplar responses demonstrates high inter-rater agreement, quantified by an Intraclass Correlation Coefficient (ICC) exceeding 0.80 among the expert panel, ensuring the reference vector accurately reflects human-aligned quality judgments. This stringent selection protocol guarantees that the reference vector is explicitly anchored to collective human expertise, making any subsequent AI–human selection divergence a more meaningful and interpretable metric of alignment.
Experimental Details: All LLM inference operations were performed on a cluster of two NVIDIA RTX 4080 SUPER GPUs, manufactured by NVIDIA in Santa Clara, USA. Each LLM selection trial, involving the processing and ranking of approximately 500 candidate ideas, took an average of 25 min. To ensure the reliability of our findings, all LLM experiments were replicated 3 times using different random seeds for initialization. Reported quantitative results, such as average selection overlap with human choices, include standard deviations to reflect the variability observed across these multiple runs.
In our selection-prediction modeling, we utilized penalized logistic regression (L2 regularization) to prevent overfitting and validated our models through nested 5-fold cross-validation. This provided robust performance measures, including McFadden’s and cross-validated Area Under the Curve (AUC), to accurately quantify the distinctness of selection mechanisms. These methodological choices collectively reinforce the credibility of our findings and provide deeper insights into why our proposed framework effectively captures the fundamental divergence between human and AI creativity.
4.6.3. Presentation and Blinding (Bias Minimization)
The objective of this subsection is to describe the randomization and blinding protocols implemented to mitigate selection bias and ensure the integrity of the comparative analysis. To ensure the objectivity of the selection process and to eliminate potential biases associated with the source of the ideas, the following protocols were implemented:
- Elimination of Stylistic Cues: Prior to evaluation, all candidate ideas underwent a standardization pipeline. This involved stripping all original formatting (e.g., Markdown bolding, specific bulleting styles, or numbering) and converting the text into a uniform plaintext format. This step was crucial to ensure that human experts could not distinguish AI-generated ideas from human ones based on superficial stylistic or structural patterns typically associated with LLM outputs.
- Mitigation of Positional Bias: To control for positional effects—most notably the “lost-in-the-middle” phenomenon [18], where LLMs exhibit higher performance for information placed at the beginning or end of an input context—the presentation order of the candidate pool was fully randomized for each selection trial. Every agent (human or AI) received a unique, randomly shuffled sequence, ensuring that the selection probability was independent of an idea’s spatial position within the list.
- Double-Blinding Protocol: The study followed a rigorous double-blind framework. Human selectors were unaware of the proportion of AI-generated content within the pool. Similarly, AI selection agents were prompted without any metadata regarding the origin of the ideas, focusing solely on the semantic content relative to the task constraints. Provenance flags were recorded internally and used only for secondary analytical stages.
4.7. Procedure
The final stage of the methodology, depicted as STEP 3 (quantitative analysis) in Figure 2, involves the statistical evaluation of the experimental data (phase F). This stage is bifurcated into two primary analytical components: (F1) Generation Metrics Analysis and (F2) Selection/Overlap Analysis.
4.7.1. Generation Metrics
To evaluate the creative output of each agent, we utilized three primary metrics: Fluency, Flexibility, and Semantic Breadth. To ensure reproducibility and provide explicit technical specifications, these variables are formally defined as follows:
(1) Fluency ()
Fluency is defined as the count of distinct idea units produced by each participant or LLM per prompt:
where is the set of valid, non-redundant ideas generated by agent for a given task. This metric serves as a baseline for ideational volume, ensuring that only unique conceptual contributions are quantified after passing through the deduplication pipeline.
(2) Flexibility ()
Flexibility is defined as the number of distinct semantic clusters represented within an agent’s responses, indicating the categorical diversity of the ideas:
where is the number of unique semantic clusters identified via Hierarchical Agglomerative Clustering (HAC).
- Cluster Granularity Sensitivity: To validate the robustness of the chosen threshold, sensitivity analyses were performed over a range of 0.12–0.18 (corresponding to a cosine similarity of 0.82 to 0.88). The ranking and statistical significance remained consistent across this range, justifying the stability of in distinguishing unique semantic categories.
(3) Semantic Breadth ()
Semantic Breadth provides a continuous measure of the exploratory divergence or “spread” of ideas in the latent semantic space. The metric is calculated as the average pairwise distance between all generated ideas:
- Pairwise Cosine Distance: The numerator sums the distances between all possible pairs of idea vectors . We utilize the Cosine Distance, where the distance is . The Cosine Similarity is formally defined as follows:
- Averaging Factor: The denominator , represents the total number of unique pairs formed by ideas. This normalization ensures remains independent of the total ideational volume (F), allowing for direct comparison across agents with different output scales.
All calculations for semantic relatedness were based on the Sentence-BERT (SBERT) framework, which incorporates principles of context vector estimation to ensure high-fidelity semantic mapping. The resulting is normalized between [0, 1], where a higher value indicates a broader range of conceptual exploration.
4.7.2. Selection Metrics
To evaluate the effectiveness of the convergent selection process, we established formal criteria to analyze how agents prioritize ideas from the candidate pool. To ensure reproducibility and address the requirement for explicit specifications, we formally define the selection algorithm as a mapping function that operates on an input candidate pool I.
(1) Algorithmic Input and Configuration
The selection process operates on , where denotes the total number of non-redundant ideas generated during the divergent phase. The algorithm produces a refined output set S ⊂ I, defined by the top five selected ideas (i.e., ) for each prompt by each agent.
During the initialization phase, the LLM’s generative parameters were strictly fixed to maintain experimental consistency:
- Temperature: 1.0 (to maximize exploration)
- Top-p: 0.95 (for linguistic coherence)
- Frequency Penalty: 1.2 (to prevent repetitive outputs)
These settings ensured that the input pool contained a sufficiently diverse and high-entropy range of candidates for selection.
(2) Selection Probability Model
The selection probability of a candidate idea i is modeled through a logistic function to reveal the underlying decision-making logic:
where
- : The sigmoid function, which maps the linear combination of predictors to a probability range of [0, 1].
- : The intercept (bias) represents the baseline probability of selection.
- : The coefficients (weights) for k agent-specific predictors.
- : The features of idea i used for evaluation. As detailed in Section 4.7.3, these include computational metrics for AI agents (e.g., Similarity) and evaluative ratings for human experts (e.g., Feasibility and Prior Experience).
(3) Selection Alignment (Overlap Rate)
To quantify the similarity in selection between human and AI agents, we define the overlap rate () as the primary output metric:
where and are the sets of ideas chosen by humans and AI, respectively. The denominator (5) reflects the fixed constraint . This ratio provides a normalized score between 0 and 1, where 1.0 indicates identical selection.
(4) Stopping Criteria and Computational Efficiency
A non-parametric permutation test with 10,000 iterations was used to construct a null distribution of overlap rates. The selection process is considered statistically complete when the observed overlap rate is validated against this baseline.
Regarding computational feasibility, all semantic processes for the 10,000-idea dataset were executed on a 12-core CPU workstation with an NVIDIA RTX 4080 SUPER GPU. Initial preprocessing concluded within ~3 h, and each LLM-based selection trial averaged 25–30 min per prompt. These explicit metrics provide the necessary technical details for full experimental reproduction.
4.7.3. Selection Predictors and Statistical Modeling
In this section, we formalize the selection process as a probabilistic mapping function, . As delineated in Algorithm 1, this framework predicts the selection probability of each idea based on agent-specific predictors, enabling a comparative analysis of human and AI decision-making logic.
| Algorithm 1: Hybrid Human–AI Convergent Selection Process |
| Input: Unified Candidate Pool , where . Output: Selection Sets (each where ). Parameters: (Temp = 1.0, Top-p = 0.95), Deduplication Threshold (1) Initialization and Preprocessing:
(3) Human Selection Modeling ()
(4) Verification and Convergence Analysis:
|
AI Selection Predictors ()
The AI selection mechanism is governed by deterministic computational metrics, which operationalize the model’s decision-making logic through the following predictors:
- (1)
- Embedding Similarity (): This measures the semantic proximity of a candidate idea to the Idealization Vector (). As defined in our framework, represents the centroid (mean vector) of the top 25% expert-rated responses, serving as a high-quality conceptual benchmark:
The predictor is then formally calculated as the cosine similarity between the candidate idea vector and :
where denotes the total number of expert-validated exemplar responses.
- (2)
- Mean Token Log-Probability (): This predictor serves as a proxy for the model’s internal statistical confidence and linguistic fluency. It represents the likelihood of the generated text within the model’s learned distribution:
- (3)
- Composite Modeling and Model Fit: To predict the final AI selection (), these standardized predictors are integrated into the logistic framework. To evaluate the predictive robustness, we utilized McFadden’s
Our model achieved a highly robust fit of , proving that AI selection is not stochastic but rigorously dictated by the combination of semantic alignment () and statistical likelihood ().
Human Selection Predictors ()
In contrast to the computational metrics of AI, the human selection model () incorporates qualitative and experiential variables that capture the nuances of professional judgment. These predictors are defined as follows:
- (1)
- Context Relevance and Feasibility (, ): These predictors represent the objective and practical alignment of an idea with the specific problem constraints.
- (Context Relevance): Measures how well the idea addresses the core requirements of the prompt.
- (Feasibility): Evaluates the technical and economic implementability of the idea. These are aggregated from Likert-scale ratings provided by human selectors to ensure a standardized quantitative input for the logistic model.
- (2)
- Prior Experience (): This is a critical predictor that represents the selector’s domain-specific professional knowledge and experiential biases. Unlike AI metrics, captures the “intuitive leap” in human decision-making:
Score based on the selector’s years of industry experience and professional profile.
This variable serves as the primary differentiator in the mechanistic divide, as it accounts for the latent expertise that transcends purely semantic similarity or statistical frequency.
- (3)
- Novelty (): Novelty is modeled as a subjective cognitive assessment of originality, often characterized as “emotional creativity.” It reflects the selector’s experiential judgment in identifying ideas that are not only relevant but also conceptually unique within the given domain.
- (4)
- Human Intuition and Model Stochasticity: Unlike the AI model’s highly deterministic fit, the human selection model exhibits a more stochastic nature. While predictors like Prior Experience () provide significant explanatory power, the human decision-making process retains a degree of variability that highlights the functional distinctness between human experiential judgment and AI’s latent semantic analysis.
4.8. Statistical Analysis Plan
The purpose of this section is to justify the choice of statistical methods and analytical frameworks used to validate the mechanistic divide between human and AI agents, ensuring that the results are grounded in robust statistical theory. The plan follows a multi-staged approach as follows:
4.8.1. Generation Parity and Welch’s T-Test
The objective of this sub-analysis is to evaluate differences in mean performance across generation metrics while rigorously accounting for variance heterogeneity. To evaluate differences in mean performance across generation metrics, we employ two-sample t-tests with Welch’s correction [19] alongside Bayesian t-tests [20]. Welch’s correction is specifically utilized to account for the heteroscedasticity observed in our data, as AI-generated outputs exhibited significantly lower variance compared to the high-variance distribution of human responses. Cohen’s d effect sizes and 95% confidence intervals are reported for all primary contrasts.
4.8.2. Overlap Significance via Permutation
This analysis aims to quantify whether the observed selection alignment between agents exceeds what would be expected by random chance. The empirical overlap between and (the overlap rate, ) is assessed using a non-parametric permutation test (10,000 iterations). This approach constructs a null distribution by randomly reassigning selection labels, yielding empirical p-values that are robust against non-normal distributions and positional biases.
4.8.3. Selection-Prediction Modeling (Elastic Net)
The objective of this modeling stage is to identify the underlying decision boundaries and predictive features governing each agent’s selection behavior. To achieve this, we utilize penalized logistic regression (Elastic Net with L1/L2 regularization), which allows us to handle potential multicollinearity between predictors such as Context Relevance and Prior Experience. This approach builds upon the probabilistic mapping function formally defined in Section 4.7.2, integrating the agent-specific predictors specified in Section 4.7.3.
- Performance Evaluation: Predictive power is evaluated using McFadden’s and Area Under the Curve (AUC). Following McFadden (1979) [21], values between 0.2 and 0.4 represent an “excellent” fit.
- Model Validation: The predictive power of the AI selection model showed strong alignment with our hypothesis (H3: 0.60), confirming that AI selections are highly governed by their internal probabilistic metrics. Our model’s achievement of 0.58 for AI selection underscores its highly deterministic nature compared to the more stochastic nature of human intuition.
4.8.4. Cross-Predictive and Alignment Testing
This component directly tests the functional alignment between agents. AI-trained models are applied to human selection holdouts (and vice versa) using an 80:20 training/holdout split with prompt-stratified sampling. Performance loss () is quantified using bootstrap confidence intervals, providing a statistically sound measure of the mechanistic distinctness between the human “experiential” process and the AI “semantic” process
4.9. Qualitative Analysis of Human Rationales
To complement the quantitative findings and gain deeper insights into human selection mechanisms, a thematic analysis will be conducted on the open-text rationales provided by human selectors for their top-ranked choices. This involves an iterative process of familiarization with the data, generation of initial codes, searching for themes, reviewing themes, defining and naming themes, and producing the report [22]. Two independent coders will perform this analysis, and inter-coder reliability will be assessed using Cohen’s Kappa. Discrepancies will be resolved through discussion or, if necessary, by a third-party adjudicator. The themes identified will be used to enrich the interpretation of the Likert scale ratings and to understand the nuances of human intuitive judgment and contextual understanding that may not be fully captured by quantitative metrics.
5. Results
The experiment, designed with realistic and conservative parameters (as detailed in Methods), produced clear results regarding the hypothesized generation parity and selection divergence between human participants and LLM agents.
The results indicated a strong ‘generation parity’ between humans and AI (Model-A and Model-B) across the measured metrics. Table 4 summarizes these findings.
Table 4.
Summary statistics for generation metrics (mean ± SD).
Regarding fluency, no statistically significant difference was observed; Humans (mean = 7.80, SD = 3.20) and Model-A (mean = 8.10, SD = 3.40) showed comparable output quantities (, providing modest evidence for the null hypothesis), and Model-B (mean = 7.95, SD = 3.35) was similarly comparable to humans. This parity extended to flexibility, where human performance (mean = 3.60, SD = 1.10) was statistically similar to both Model-A (mean = 3.70, SD = 1.20; ) and Model-B (mean = 3.65, SD = 1.15; ). Furthermore, the Semantic Breadth, representing the diversity of ideas in the semantic space, was comparable between humans (mean = 0.420, SD = 0.080) and both Model-A (mean = 0.400, SD = 0.090) and Model-B (mean = 0.415, SD = 0.085).
Collectively, these results provide robust support for H1, demonstrating that AI can generate candidate ideas at a level statistically comparable to human performance in divergent ideation. This empirical “generation parity,” clearly visualized in the overlapping error bars of Figure 3, confirms that any subsequent divergence observed in the selection phase is likely attributable to differences in evaluative processes and selection mechanisms rather than innate generative capacity.
Figure 3.
Comparison of creative generation metrics across humans and AI models. The vertical dashed lines indicate the shared mean scores between Fluency and Flexibility, used to eliminate redundancy and streamline the visual comparison. Standard deviations (±SD) are indicated by numerical values. The visual uniformity in mean performance across Fluency, Flexibility, and Semantic Breadth illustrates the robust generation parity (H1) between human participants and both AI agents.
Following the generation phase, we examined the convergent selection outcomes to determine if generative parity extends to the evaluative stage. After deduplication, the pooled candidate sets averaged 48 unique items per prompt. Despite starting from identical candidate pools, a substantial divergence was observed in the top-5 selection results, where the empirical overlap was assessed using a permutation test (10,000 iterations) to construct a robust null distribution.
As illustrated in Figure 4, the analysis revealed that the mean overlap rate was 19.2% for Model-A (0.96 items per prompt) and 22.4% for Model-B (1.12 items per prompt). Critically, these observed overlaps were significantly below the permutation-based chance level (null mean overlap ≈ 35%, p < 0.001). This indicates that humans and AI agents systematically prioritize distinct, non-random subsets of ideas, with the observed overlap remaining consistently lower than what would be expected by random chance alone.
Figure 4.
Selection overlap vs. random chance baseline. The observed mean overlap rate is consistently below the permutation-based chance level (approximately 35%, indicated by the dashed line), visually confirming the systematic misalignment (H2) between human and AI selections. This gap provides the empirical motivation for our Selection Alignment framework to bridge this evaluative divide. This systematic misalignment confirms that the convergent stage amplifies human–AI divergence, marking a clear departure from the parity observed during the ideation phase and providing strong empirical support for H2. This divergence carries profound implications beyond experimental observation. From a business perspective, such misalignment can increase product failure rates and cause significant market launch delays. Furthermore, the broader societal consequences are significant; when AI selects ideas based on statistical optimization without human experiential judgment, it may exacerbate policy inequality or social conflict by overlooking the nuanced needs of marginalized groups.
To investigate the underlying causes of the observed divergence, we utilized penalized logistic regression (L2 regularization) to model the decision boundaries of each agent type. Our analysis revealed a profound mechanistic divide between the two, as AI choices were strongly predicted by internal statistical metrics. Specifically, the AI selection model, utilizing metrics such as embedding similarity (z = 6.42, p < 0.001) and mean token log-probability (z = 3.15, p < 0.001), achieved a robust fit with a McFadden’s pseudo- of 0.58, which significantly exceeds the hypothesized threshold of 0.50 (AUC = 0.85). These results suggest that AI’s evaluative logic is fundamentally bound by token-level predictability and internal distributional patterns.
In contrast, human choices were better predicted by contextual and experiential criteria (see Table 5), yielding a pseudo- of 0.42 and an AUC of 0.78, meeting the hypothesized expectation of . The strongest human predictors included context relevance (z = 5.92, p < 0.001), prior experience (z = 4.83, p < 0.001), and feasibility (z = 2.84, p < 0.01), highlighting a reliance on tacit knowledge and practical implementability.
Table 5.
Predictors and model fit for AI and human selection models.
To confirm the distinctness of these underlying mechanisms, we conducted a cross-prediction test. The results, as illustrated in Table 5, showed a substantial and symmetric drop in predictive performance across agent types, providing definitive support for H3 (Distinct Mechanisms). Predicting human choices using AI metrics resulted in an AUC drop of 18 percentage points (0.78 to 0.60), while predicting AI choices using human criteria led to a steeper drop of 23 percentage points (0.85 to 0.62). This significant loss of explanatory power validates the mechanistic divergence and provides a crucial empirical foundation for our “Selection Alignment” framework, which explicitly anchors AI’s internal evaluation criteria to rigorous human domain expertise to achieve genuine human–AI alignment.
6. Discussion
6.1. Summary of Findings
The results provide compelling evidence that while humans and contemporary generative models achieve statistical parity in divergent ideation (H1), their convergent choices from a shared candidate set differ substantially (H2). Establishing this ‘generation parity’ was a critical prerequisite of this study; it serves as a controlled baseline that isolates the ‘selection gap’ as a distinct cognitive and algorithmic divergence rather than a mere byproduct of unequal generative quality.
This divergence is mechanistically explained by fundamentally distinct underlying criteria (H3). Our predictive modeling revealed a profound divide: AI selections are governed by internal statistical optimization, achieving a robust fit ( = 0.58) through metrics like embedding similarity and token log-probability. This suggests that AI’s evaluative logic is essentially a closed-loop extension of its training distribution, prioritizing internal consistency over real-world utility. In stark contrast, human selections are better explained by external, experiential heuristics ( = 0.42), such as context relevance and feasibility, highlighting a reliance on tacit domain knowledge. The significant drop in cross-prediction performance, reflecting an 18–23% divergence in selection patterns (as measured by AUC drop), moves beyond merely establishing divergence; it provides a definitive mechanistic account. It reveals that AI prioritizes internal statistical likelihood, while humans prioritize external practical constraints. This distinction is critical because it opposes the “generative proxy myth”—the flawed assumption that superior generative performance automatically implies human-aligned evaluative judgment. By empirically validating this mechanistic divide, our study underscores the urgent necessity for “Selection Alignment” to bridge the gap between AI’s statistical output and human-centric value.
6.2. Distinct Selection Mechanisms: Human vs. LLM
The core of the observed divergence lies in the fundamental differences in selection mechanisms. As evidenced by the profound mechanistic divide revealed in our predictive modeling (H3), the internal ‘logic’ governing each agent type is not only distinct but effectively operates in disparate cognitive and computational dimensions.
6.2.1. Principles of AI Selection: Inherent Statistical Bias and Latent Typicality
LLMs operate by identifying patterns and relationships within high-dimensional datasets [23]. Their selection process, as illuminated by our high predictive fit ( = 0.58), is inherently statistical and predictive:
- Embedding Similarity: LLMs prioritize ideas that are semantically close to the prompt or a predefined “ideal” representation [24]. This indicates statistical typicality within its training corpus [25]. AI favors semantically prototypical, well-trodden ideas that are maximally associated with the prompt, prioritizing distributional frequency over creative divergence [26].
- Mean Token Log-Probability: Ideas composed of highly probable token sequences are favored. This reflects the model’s internal confidence in linguistic fluency. While ensuring coherence, this mechanism rewards linguistic “safety” rather than real-world utility or breakthrough potential.
- Internal Consistency: Composite ranking scores maximize a function based on learned associations. These mechanisms enable LLMs to select ideas that are relevant in a statistical sense but lack direct grounding in external, real-world information beyond what was implicitly encoded during pre-training.
6.2.2. Human Selection Mechanism: Contextual, Experiential, and Situated
Human selection, particularly among design professionals, is a situated cognitive process that integrates various forms of embodied knowledge [27] ( = 0.42):
- Contextual and Pragmatic Factors: Human choices are strongly predicted by external factors [28], specifically Context Relevance , Prior Experience , and Feasibility . Qualitative rationales confirm that humans emphasize implementation constraints and domain-specific viability over mere semantic fit. Human selection is thus an external anchoring process [29].
- Prior Experience and Feasibility: Designers leverage accumulated professional experience and tacit knowledge. Unlike AI’s statistical likelihood, human feasibility assessment is a pragmatic judgment based on technical limitations, budgets, and ethical considerations.
- Intuitive Coherence: Humans often make intuitive judgments [30] about an idea’s potential, integrating subtle cues and prior industry knowledge that remain outside the AI’s statistical reach.
6.2.3. The Quantitative Evidence of Mechanistic Separation (H3 Validation)
Our cross-prediction results provide the most compelling evidence for this separation. The substantial and symmetric drop in predictive performance—an AUC loss of 0.18 for human choices and 0.23 for AI choices—underscores that the features driving AI selection hold remarkably little explanatory power for human convergent choices.
The clear ∆AUC and substantial losses confirm that the internal ‘logic’ of one agent is fundamentally detached from the other. This divergence is not random; it is a systematic misalignment between statistical optimality (AI) and contextual appropriateness (Human). This definitive divide validates the urgent necessity for the Selection Alignment framework to bridge the gap between AI’s statistical output and human-centric value.
6.3. Theoretical Implications
The core theoretical implication of this study is that the human mechanism for convergent selection reflects the inherent nature of the design discipline, which operates on fundamentally different principles than the statistical optimization of Large Language Models. Our findings from H3 provide definitive empirical weight to this argument: while AI selection is tightly bound to the statistical topography of its training distribution ( = 0.58), human choice is anchored in complex contextual and experiential heuristics (= 0.42). This systematic misalignment between statistical optimality (AI) and contextual appropriateness (human) validates the urgent necessity for the Selection Alignment framework to bridge this gap.
This fundamental divergence is critical because the central challenges of design are, by their nature, characterized as ‘Wicked Problems’. The high predictive power of internal metrics in AI selection confirms its reliance on statistical typicality, which fundamentally lacks the capacity to navigate the inherent ambiguity, ethical nuances, and shifting definitions of such problems. While AI excels at optimization within well-defined, closed problem spaces, its ability to create value in ambiguous, open-ended contexts remains limited. Our study opposes the “generative proxy myth”—the flawed assumption that generative parity implies evaluative alignment—by showing that AI prioritizes internal statistical consistency over the external practical constraints essential for solving “Wicked Problems”. This mechanistic divide underscores the urgent necessity for “Selection Alignment” to bridge the gap between AI’s statistical output and human-centric value, as highlighted in Section 6.1.
This methodological advantage, rooted in the deep evaluation by our expert panel, distinguishes our ‘Idealization’ vector from simpler baseline approaches. In ambiguous environments, AI selection relying merely on probability-based ranking cannot replicate the genuine practical insight of human designers. The finding that human selectors’ “Pragmatic Feasibility” , and “Prior Experience” , yielded significantly higher predictive power for human choices than any AI metrics, reaffirms that feasibility is the cornerstone of human convergent thinking [31]. This supports the argument that designers continuously utilize tacit knowledge through Reflective Practice [32].
Even as designers select an idea, they are simultaneously reflecting upon and restructuring it to assess its viability within real-world technical, economic, and social constraints [33].
Therefore, our findings yield a crucial directive for the field of AI research: AI must evolve beyond being a mere tool for generating ‘plausible’ ideas. Instead, a new framework is needed that can computationally emulate or augment the human mechanism of reflective selection. To bridge a low selection overlap (approx. 18–23%) revealed in our cross-prediction analysis—which represents the core mechanistic gap—we propose the Selection Alignment framework.
To achieve this, the proposed conceptual architecture for reflective selection in AI integrates:
- Creative Thought Embeddings (CTE): To internalize human judgment criteria within a shared embedding space.
- Reinforcement Learning with Human Feedback (RLHF): To progressively learn human values by integrating feedback during the “selection sorting” process.
- Domain Expert Knowledge Graph Integration: To systematically embed human expert knowledge and contextual data into the AI’s reasoning [34].
- Context-Aware and Adaptive Decision-Making Module: To enable AI to dynamically adjust evaluations based on real-time contextual cues and ethical implications.
The rigor of this architecture is grounded in both empirical justification and computational feasibility. Our design choices, such as the deduplication threshold ( = 0.15), were identified as optimal “elbow points” through sensitivity analysis, while a 10,000-iteration permutation test ensured statistical robustness ( < 0.01). From a complexity perspective, the system maintains a linear scale relative to the candidate pool, ensuring practical scalability. While specific parameterizations are beyond the current scope, this study establishes the necessary “Selection Alignment” foundation upon which future module-specific convergence properties will be rigorously benchmarked.
6.4. Practical Implications for Human–AI Collaboration Tools
The observed divergence offers several concrete implications for the design of human–AI collaboration tools:
Performance Improvement and Validation Methodology through Transparent AI Ranking: A transparent AI ranking system explicitly provides the AI’s internal ranking signals (e.g., similarity scores, confidence levels) for AI-generated ideas [35], helping human users understand why the AI prioritized certain ideas. This transparency can enhance collaboration and critique [36], potentially reducing decision-making time in the industrial design process by 15% and early design error rates by 10%. These performance improvements can be validated through the following methodologies:
Validation of Decision-Making Time Reduction and Error Rate Decrease
- A/B Testing: Conduct A/B tests comparing groups that use a transparent AI ranking system with control groups that do not, measuring decision-making time and early design error rates. This allows for quantitative assessment of effectiveness in real industrial settings.
- Pilot Projects: Implement transparent AI ranking in the actual design process for pilot projects involving N companies. Collect and analyze data on decision-making time and early error rates before and after the implementation. For example, similar to studies showing a 20% increase in user satisfaction by integrating user safety data and emotional assessments into AI-powered collision avoidance systems in automotive design, quantitative improvements in specific industrial sectors can be validated.
Performance Improvement and Validation Methodology through Contextual Constraint Injection Mechanisms: Mechanisms for contextual constraint injection should provide intuitive ways for human users to inject their unique contextual constraints, domain expertise, and pragmatic criteria into the AI’s selection process [37]. This could involve adjustable weighting parameters for criteria like feasibility, cost, user desirability, or ethical considerations, allowing humans to guide the AI’s “attention.” [38]
Validation of Design Decision Time and Error Rate
- Quantitative Data Collection: Measure the time spent on design decisions and the error rate in design processes utilizing constraint injection mechanisms versus those without. For instance, AI can eliminate costly redesign cycles and significantly shorten R&D cycles by considering manufacturability, cost, and performance constraints from the outset [39].
- User Studies: Conduct user studies with industrial designers to evaluate the impact of contextual constraint injection on the quality and efficiency of design decisions. This allows for an analysis of changes in designers’ trust in AI’s role and their satisfaction.
Performance Improvement and Validation Methodology through Mixed-Initiative Workflows: Future systems should facilitate dynamic, mixed-initiative workflows where both human and AI actively contribute to selection. This approach implies interfaces that support human veto power, re-ranking of AI-proposed ideas, and the ability to merge AI-generated clusters with human-prioritized alternatives. A mixed-initiative approach can offer significant productivity gains, potentially cutting the iterative cycle time in the design process by 20% and thus substantially accelerating the overall Time-to-Market.
Validation of the Iterative Cycle Time and Time-to-Market Reduction
- Experimental Evaluation: Validate the reduction in iterative cycle time and acceleration of time-to-market through experimental evaluation of mixed-initiative systems. AI-driven generative design can simultaneously create and virtually test thousands of design variations, enabling multi-objective optimization levels not possible with traditional methods [40].
- Benchmark Comparison: Compare iterative cycle time and time-to-market improvements when applying mixed-initiative workflows against existing benchmark data from similar industrial projects.
Performance Improvement and Validation Methodology through Continuous Learning Loops: Long-term pilot programs are indispensable for operationalizing AI systems with continuous learning loops, providing a structured environment to monitor the evolution of human–AI collaboration and systemic error reduction. By institutionalizing human-in-the-loop protocols—where experts categorize failure cases and refine labeling guidelines—these programs mitigate the performance degradation often encountered in volatile, real-world deployments [41]. Critical to this iterative process is the establishment of rigorous governance over the feedback loop; specifically, model retraining must be contingent upon satisfying predefined fairness and accuracy thresholds within bias-sensitive validation datasets [42]. Such a framework ensures that continuous improvement does not occur at the expense of ethical integrity but rather serves as a robust mechanism for maintaining long-term operational alignment and sociotechnical trustworthiness.
Validation of Human–AI Collaboration and Error Reduction
- Long-term Pilot Programs and Iterative Learning: To ensure superior joint performance over fully manual or automated systems, AI models must be deployed within long-term pilot programs featuring continuous learning loops. These programs serve as essential diagnostic mechanisms for identifying and rectifying performance degradation—such as the inaccuracies caused by noisy data or domain-specific nuances—that often emerge when transitioning from controlled environments to real-world operations [43]. For instance, in complex domains like financial document intelligence, human experts fulfill a critical role by annotating failure cases and refining labeling guidelines to feed corrections back into the training cycle [44]. To maintain systemic integrity, these programs must establish rigorous protocols for the feedback-based retraining cycle, utilizing validation datasets specifically designed to challenge potential biases [45]. By defining strict thresholds for automation—where autonomous retraining is permitted only if predefined fairness and accuracy metrics are sustained—this framework ensures that continuous improvement remains both ethically sound and operationally robust [46].
- Interdisciplinary Research Teams: To ensure AI systems authentically reflect human preferences, interdisciplinary teams comprising AI developers, data scientists, and domain experts must be formed to deconstruct the learning process and evaluate the socio-technical impact of human feedback [47]. Grounded in human-centered design (HCD) frameworks [48] and AI implementation taxonomies [49], these teams facilitate sophisticated interaction models that bridge the gap between algorithmic output and user expectations [50]. Beyond providing initial design, these teams establish robust, continuous feedback loops—acting as essential sensing mechanisms to systematically monitor performance and ethical behavior [51]. Specifically, they are tasked with conducting regular ethical audits, analyzing integrated fairness metrics, and enforcing rigorous protocols for human oversight. By defining the frequency of expert reviews and establishing clear criteria for autonomous retraining versus mandatory human intervention, these teams prevent the propagation of biases and ensure the long-term trustworthiness and operational alignment of the human–AI system.
6.5. Limitations
This study provides compelling evidence of Human–AI divergence in convergent selection; however, the findings must be interpreted in light of several limitations. The LLM agents (Model-A, Model-B), while powerfully mimicking generative AI capabilities as black-box models, do not fully capture model-specific idiosyncrasies nor comprehensively represent the entire LLM ecosystem. Specifically, while the related work section discusses various recent competing approaches, including specialized Reinforcement Learning models [52] and niche-domain architectures [53], these were not included in the current experimental comparison. This exclusion was primarily due to the necessity of a controlled mechanistic analysis; our framework requires consistent access to internal computational proxies, such as log-probability and embedding similarity, which are not universally accessible or stable across all proprietary or emerging SOTA platforms.
This limitation, arising from restricted access to their internal mechanisms, makes it challenging to definitively discern whether an AI’s decisions stem from training data biases or from inherent limitations within general AI decision-making processes. Specifically, the significant decline in cross-prediction performance (an AUC drop of 18–23%) observed between humans and AI suggests a fundamental divergence in their creative evaluation logics. However, it remains unclear whether this numerical degradation stems purely from a mismatch in value criteria or is exacerbated by intrinsic technical biases of LLMs [54]. Recent studies suggest that AI selections may be heavily influenced by ‘frequency bias’—a tendency to favor common patterns in the training corpus—or ‘position bias’—the disproportionate weighting of information based on its sequence within a prompt—rather than purely qualitative merit [18,43]. Furthermore, the attention mechanism’s tendency to prioritize specific token patterns [55] represents an architectural constraint that fundamentally differs from human cognitive synthesis, which integrates contextual meaning and external constraints through a more holistic evaluative lens [56]. Future research should explore open-source models with transparent logit-level data to more precisely isolate these computational biases from genuine evaluative judgment. Second, while our study relies on a panel of 10 senior design experts to establish a “Human-Consensus Baseline,” we acknowledge that this sample size, though high in professional depth (average 15.4 years of experience), may not represent the full spectrum of global design perspectives. However, the high inter-rater agreement (ICC > 0.80) observed in our study suggests a stable professional consensus within the targeted domain. Future research should aim to expand the diversity and size of the expert panel to further validate these findings. Crucially, this study does not treat human selection as an infallible ‘gold standard [57]’ of universal truth, but rather as a necessary reference point to measure the degree of ‘Selection Alignment’ between human-centered values and AI’s statistical decision-making.
Third, our investigation of creative parity and divergence was confined to textual ideation. The generalizability of these findings to domains outside written ideas, such as visual design, music composition, or engineering problem-solving, remains to be empirically tested. Text-based ideas are expressed through explicit language, making them relatively easy to quantify and compare. Extending this research to non-textual domains would necessitate developing domain-specific quantitative and qualitative evaluation criteria.
Finally, while our “generation parity” relies on quantitative metrics (Fluency, Flexibility, and Semantic Breadth), these do not guarantee qualitative equivalence, such as originality, elaboration, or esthetic appeal. AI models, relying on internal computational mechanisms, inherently lack the capacity to fully incorporate rich contextual knowledge and complex user experience understanding. This fundamental limitation explains why AI selections can overlook critical practical feasibility or user experience considerations in real-world applications [58], suggesting that while AI can match human performance on quantifiable aspects, the more nuanced elements of human creative cognition remain a distinct challenge.
6.6. Future Directions: Embedding Human Experience and Context into AI Selection
To bridge the gap between human and AI selection mechanisms and enhance AI’s human-centric judgment capabilities, research must focus on directly integrating human experience and context into AI. A critical priority for future directions is the development of robust methodologies to decouple technical artifacts, such as position and frequency biases, from the core evaluative logic of AI. By isolating these confounding variables, researchers can more accurately align AI’s ‘Selection Logic’ with human-centered values, ensuring that divergence is interpreted as a difference in creative perspective rather than a mere computational error. Research should prioritize the development of vertical AI agents optimized for professional creative workflows, moving beyond broad industrial speculations to refine selection algorithms that are resilient against statistical biases. Validating the effectiveness of these models through iterative pilot projects within the design domain is essential to ensure they capture the nuances of professional intuition. Furthermore, establishing ethical guidelines and an AI governance framework specifically for co-creative systems [59] will be necessary to manage the integration of human-centric values into automated selection processes.
6.6.1. Data-Driven Integration of Human Context
Contextualized Feedback Loops for AI Model Refinement: To effectively fine-tune AI models for selection tasks, collecting granular, contextualized human feedback is crucial, moving beyond simplistic “good/bad” binary responses. Crucially, these feedback loops must be designed to explicitly distinguish between an AI’s systematic selection logic and its underlying statistical biases, such as position or frequency bias. It requires sophisticated systems that capture detailed rationales, contextual variables—such as specific design constraints or user-centered requirements—to understand decision drivers [60]. A structured labeling system, based on predefined contextual factors such as feasibility, ethics, and technical difficulty, should be implemented to generate structured input for AI learning. By incorporating metadata that tracks the original sequence and frequency of items during the feedback stage, researchers can apply debiasing techniques during the Reinforcement Learning from Human Feedback (RLHF) process [61]. This allows the AI to adapt more effectively to complex, real-world constraints and human values while minimizing the influence of architectural artifacts.
Interfaces enabling interactive and iterative feedback are essential for continuous AI learning. Features like preference sliders and conversational prompts (e.g., “Why does this idea not meet the project constraints?”) allow for granular adjustments and direct human guidance. Future research should specifically focus on developing “Bias-Aware” feedback interfaces that can detect and mitigate the 18–23% performance drop in cross-prediction by isolating technical noise from genuine human-centric evaluative criteria.
Experiential Datasets: The development of “experience-rich” datasets is crucial for training AI to infer and prioritize human-centric criteria, moving beyond purely linguistic data. This necessitates a collaborative data collection initiative involving design communities and industry partnerships, where real-world project data is anonymized and linked with extensive contextual metadata [62]. Such metadata should include ‘project goals,’ ‘target users,’ and ‘feedback history,’ thereby enriching the dataset with actionable insights. This approach underscores the need for robust data ontology and knowledge graph construction, enabling AI to understand the intricate relationships and meanings within the data rather than just isolated statistical patterns.
Integrating these datasets with causal inference techniques [63] allows AI to move beyond mere correlation and identify the causal drivers of human selection, effectively filtering out the statistical artifacts that lead to human–AI divergence.
6.6.2. Model Design for Contextual and Experiential AI: Tacit Knowledge and Adaptive Learning
To bridge the 18–23% performance gap identified in this study, future model architectures must move beyond “black-box” statistical proxies toward systems that internalize human tacit knowledge. The integration of human intuitive judgments into AI models can be achieved by tracing and modeling human designer decision paths, leveraging techniques such as think-aloud protocols and eye-tracking data [64]. By formalizing these intuitive paths, AI can adopt evaluation logics that mirror human-centered expertise rather than relying on surface-level positional artifacts.
A protocol for capturing human decision-making and formalizing tacit knowledge involves a structured seven-step process: (1) task definition, (2) concurrent think-aloud, (3) eye-tracking data collection, (4) retrospective think-aloud, (5) expert interview, (6) data annotation (logical atoms), and (7) knowledge graph construction. Crucially, this symbolic representation provides a grounding mechanism that prevents the AI from over-relying on superficial statistical patterns—such as the divergence identified in our analysis—by anchoring its selection logic in verified human heuristics [65].
Furthermore, exploring interactive Reinforcement Learning from Human Feedback (RLHF) within a human–AI co-creation framework is essential [66]. To mitigate the biases identified in this study, the reward function must be normalized against the sequence and frequency of presented options. This approach allows AI to evolve from a mere idea generator to a sophisticated co-creative partner whose selections are informed by the depth of human experience and the richness of real-world context [67].
7. Conclusions
This study systematically investigated the ‘selection bottleneck’ in Human–AI creative collaboration, moving beyond generative fluency to the critical phase of convergent evaluation. Our findings successfully address the three-fold methodological and theoretical objectives established in the Introduction:
First, through the quantification of ‘Selection Alignment,’ we demonstrated that while human and generative AI capabilities converge in divergent ideation (generation parity), a substantial divergence persists in the convergent phase. By isolating this “selection gap”—quantified through AUC (Area Under the Curve) and cross-prediction analysis—this study proves that traditional fluency-based benchmarks are becoming insufficient [68]. The results shift the focus of AI research from maximizing generative output to achieving evaluative alignment with human expert values.
Second, the implementation of our ‘dual-phase experimental protocol’ allowed for a controlled environment to decouple generation from selection. This rigorous approach revealed that the divergence is not merely a matter of different outcomes, but a result of AI’s inability to mirror the evaluative rigor required in professional design tasks. The protocol confirmed that the “selection gap” remains the primary constraint in achieving seamless human–AI synergy.
Third, the construction of the ‘Idealization Vector’ provided a mechanistic explanation for this divergence. Our analysis confirms that while AI’s decision-making is predominantly driven by internal statistical density and pattern matching, human selection is profoundly anchored in contextual, experiential, and pragmatic constraints. This “mechanistic divide” clarifies why AI, despite its generative prowess, often fails to prioritize ideas that possess high professional feasibility and human-centric value.
In summary, the true value of this study lies in redefining the evaluation of AI creativity by removing speculative industrial distractions and focusing on these core empirical findings. The findings underscore that AI’s value is realized not through automation alone, but through a structured partnership that scales human creativity while remaining grounded in human-centric evaluative logic. This shift is crucial for bridging the gap between technological potential and real-world application in professional design and other expertise-driven fields. Future advancements must move beyond viewing AI as a standalone generator toward developing contextually aware partners capable of “reflective selection.”
Supplementary Materials
The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/info17030243/s1, File S1: Informed consent form. This document provides the full description of the informed consent provided to all participants, including study objectives, data anonymity, and voluntary participation protocols.
Author Contributions
Conceptualization, S.J.; data curation, S.J.; formal analysis, S.J.; funding acquisition, S.J. and K.N.; investigation, S.J.; methodology, S.J.; project administration, S.J. and K.N.; software, S.J.; validation, S.J.; visualization, S.J.; writing—original draft preparation, S.J.; writing—review and editing, S.J. and K.N.; resources, K.N.; supervision, K.N. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Institutional Review Board Statement
The study was conducted in accordance with the Declaration of Helsinki. Ethical review and approval were waived for this study by the School of Art and Design at Zhejiang A&F University, due to the research involving only anonymous surveys and non-invasive cognitive behavioral testing that pose minimal risk to participants, consistent with national and institutional guidelines (Measures for Ethical Review of Life Sciences and Medical Research Involving Human Subjects, 2023). The informed consent form is provided in the Supplementary Materials (File S1).
Informed Consent Statement
Informed consent was obtained from all subjects involved in the study. Participants were informed that their involvement was voluntary and that no personally identifiable information would be collected or stored. The detailed consent protocol and the Informed Consent Form are available in the Supplementary Materials (File S1).
Data Availability Statement
Data are contained within the article.
Conflicts of Interest
The authors declare no conflicts of interest.
Appendix A. Experimental Prompts
This table presents the full corpus of 20 prompts used in the study. Each prompt was presented identically to both human participants and AI agents to ensure parity in the experimental conditions.
Table A1.
Detailed list of 20 experimental prompts for generative and evaluative tasks.
References
- Hatchuel, A.; Le Masson, P.; Weil, B. Teaching innovative design reasoning: How concept–knowledge theory can help overcome fixation effects. Artif. Intell. Eng. Des. Anal. Manuf. 2011, 25, 77–92. [Google Scholar] [CrossRef] [Scilit]
- Hubert, K.F.; Awa, K.N.; Zabelina, D.L. The current state of artificial intelligence generative language models is more creative than humans on divergent thinking tasks. Sci. Rep. 2024, 14, 3440. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Doshi-Velez, F.; Kim, B. Towards a Rigorous Science of Interpretable Machine Learning. arXiv 2017, arXiv:1702.08608. [Google Scholar] [CrossRef] [Scilit]
- Xu, Y.; Chen, J.; He, L. How personal, experiential, and contextual factors mediate EFL teachers’ conceptions of assessment: A narrative study. Chin. J. Appl. Linguist. 2023, 46, 251–269. [Google Scholar] [CrossRef] [Scilit]
- Beaty, R.E.; Johnson, D.R. Automating creativity assessment with SemDis: An open platform for computing semantic distance. Behav. Res. Methods 2021, 53, 757–780. [Google Scholar] [CrossRef] [Scilit]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is All you Need. In Proceedings of the 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA, 4–9 December 2017; pp. 5998–6008. [Google Scholar]
- Raiaan, M.A.K.; Mukta, M.S.H.; Fatema, K.; Fahad, N.M.; Sakib, S.; Mim, M.M.J.; Ahmad, J.; Ali, M.E.; Azam, S. A Review on Large Language Models: Architectures, Applications, Taxonomies, Open Issues and Challenges. IEEE Access 2024, 12, 26839–26874. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Z.; Yuan, K.H. Practical Statistical Power Analysis; ISDSA Press: Granger, IN, USA, 2018. [Google Scholar]
- Reimers, N.; Gurevych, I. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Hong Kong, China, 3–7 November 2019; pp. 3982–3992. [Google Scholar] [CrossRef] [Scilit]
- Galli, C.; Donos, N.; Calciolari, E. Performance of 4 Pre-Trained Sentence Transformer Models in the Semantic Query of a Systematic Review Dataset on Peri-Implantitis. Information 2024, 15, 68. [Google Scholar] [CrossRef] [Scilit]
- Socolovsky, E.A.; Bushnell, D.M. A Dissimilarity Measure for Clustering High- and Infinite Dimensional Data That Satisfies the Triangle Inequality; NASA/CR-2002-212136; NASA Langley Research Center: Hampton, VA, USA, 2002. [Google Scholar]
- Oti, E.U.; Olusola, M.O. Overview of Agglomerative Hierarchical Clustering Methods. Br. J. Comput. Netw. Inf. Technol. 2024, 7, 67–79. [Google Scholar] [CrossRef] [Scilit]
- Rahutomo, F.; Kitasuka, T.; Aritsugi, M. Semantic cosine similarity. In Proceedings of the 7th International Student Conference on Advanced Science and Technology (ICAST 2012), Seoul, Republic of Korea, 22–23 October 2012; pp. 1–4. [Google Scholar]
- Ester, M.; Kriegel, H.-P.; Sander, J.; Xu, X. A Density-Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise. In Proceedings of the Second International Conference on Knowledge Discovery and Data Mining (KDD-96), Portland, OR, USA, 2–4 August 1996; pp. 226–231. [Google Scholar]
- Shrout, P.E.; Fleiss, J.L. Intraclass correlations: Uses in assessing rater reliability. Psychol. Bull. 1979, 86, 420–428. [Google Scholar] [CrossRef]
- Cohen, J. A Coefficient of Agreement for Nominal Scales. Educ. Psychol. Meas. 1960, 20, 37–46. [Google Scholar] [CrossRef] [Scilit]
- Markov, I.L. Limits on fundamental limits to computation. Nature 2014, 512, 147–154. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Liu, N.F.; Lin, K.; Chen, C.S.; Reiche, P.S.; Omrani, A.; Chen, J.X. Lost in the Middle: How Language Models Use Long Contexts. Trans. Assoc. Comput. Linguist. 2024, 12, 157–173. [Google Scholar] [CrossRef] [Scilit]
- Welch, B.L. The generalization of ‘Student’s’ problem when several different population variances are involved. Biometrika 1947, 34, 28–35. [Google Scholar] [CrossRef] [Scilit]
- Rouder, J.N.; Speckman, P.L.; Sun, D.; Morey, R.D.; Iverson, G. Bayesian t tests for accepting and rejecting the null hypothesis. Psychon. Bull. Rev. 2009, 16, 225–237. [Google Scholar] [CrossRef] [Scilit]
- McFadden, D. Quantitative Methods for Analyzing Travel Behaviour of Individuals: Some Recent Developments. In Behavioural Travel Modelling; Hensher, D.A., Stopher, P.R., Eds.; Croom Helm: London, UK, 1979; pp. 279–318. [Google Scholar]
- Beck, T. The importance of a priori sample size estimation in strength and conditioning research. J. Strength Cond. Res. 2013, 27, 2323–2337. [Google Scholar] [CrossRef] [Scilit]
- Pandey, S.; Chand, S.; Horkoff, J.; Staron, M. Design pattern recognition: A study of large language models. Empir. Softw. Eng. 2025, 30, 16. [Google Scholar] [CrossRef] [Scilit]
- Tang, X.; Zheng, Z.; Li, J.; Meng, F.; Zhu, S.-C.; Liang, Y.; Zhang, M. Large Language Models are In-Context Semantic Reasoners rather than Symbolic Reasoners. arXiv 2024, arXiv:2305.14825. [Google Scholar]
- Misra, K.; Ettinger, A.; Rayz, J. Do language models learn typicality judgments from text? In Proceedings of the ACL-IJCNLP 2021, Online, 1–6 August 2021; pp. 2577–2583. [Google Scholar]
- Chávez-Autor, J. Artificial Creativity: From predictive AI to Generative System 3. Front. Artif. Intell. 2025, 8, 1654716. [Google Scholar] [CrossRef] [Scilit]
- Smithwick, D.; Sass, L. Embodied Design Cognition: Action-Based Formalizations in Architectural Design. Int. J. Archit. Comput. 2014, 12, 399–419. [Google Scholar] [CrossRef] [Scilit]
- Mai, J.-E. Contextual analysis for the design of controlled vocabularies. Bull. Am. Soc. Inf. Sci. Technol. 2006, 33, 18–22. [Google Scholar] [CrossRef] [Scilit]
- Li, Z. The Mental Scale in Anchoring Effects: Evidence from Event-Related Potentials. Master’s Thesis, Beijing Normal University, Beijing, China, 2008. [Google Scholar]
- Kahneman, D.; Klein, G. Conditions for intuitive expertise: A failure to disagree. Am. Psychol. 2009, 64, 515–526. [Google Scholar] [CrossRef] [Scilit]
- Livotov, P.; Mas’udah, M. AI-powered inventive design: Idea funnelling, concept creation, and hybrid problem-solving teams. Proc. Des. Soc. 2025, 5, 479–488. [Google Scholar] [CrossRef] [Scilit]
- Kinsella, E.A. Professional knowledge and the epistemology of reflective practice. Nurs. Philos. 2010, 11, 3–14. [Google Scholar] [CrossRef] [Scilit]
- Vincenti, W.G. The Technical Shaping of Technology: Real-World Constraints and Technical Logic in Edison’s Electrical Lighting System. Soc. Stud. Sci. 1995, 25, 553–574. [Google Scholar] [CrossRef] [Scilit]
- Zha, X.F.; Lim, S.Y.E.; Fok, S.C. Integrated knowledge-based approach and system for product design for assembly. Int. J. Comput. Integr. Manuf. 1999, 12, 211–237. [Google Scholar] [CrossRef] [Scilit]
- Mohseni, S.; Zarei, N.; Ragan, E.D. A multidisciplinary survey and framework for design and evaluation of explainable AI systems. ACM Trans. Interact. Intell. Syst. 2021, 11, 1–45. [Google Scholar] [CrossRef] [Scilit]
- Cabrera, Á.A.; Perer, A.; Hong, J.I. Improving Human-AI Collaboration With Descriptions of AI Behavior. In Proceedings of the ACM on Human-Computer Interaction; Association for Computing Machinery: New York, NY, USA, 2023; Volume 7, pp. 1–21. [Google Scholar] [CrossRef] [Scilit]
- Johns, C. Model-Based User Interface Optimization Under Contextual Uncertainty. Ph.D. Thesis, Aarhus University, Aarhus, Denmark, 2025. [Google Scholar]
- Kölmel, B.; Brugger, T.; Pekmezci, M.; Pereira, C.; Bulander, R.; Volz, R. Achieving the “AI sweet spot”: Balancing feasibility, viability, and desirability in Artificial Intelligence Implementation. Int. J. Econ. Bus. Manag. Stud. 2025, 12, 1–17. [Google Scholar] [CrossRef] [Scilit]
- Verganti, R.; Vendraminelli, L.; Iansiti, M. Innovation and Design in the Age of Artificial Intelligence. J. Prod. Innov. Manag. 2020, 37, 212–227. [Google Scholar] [CrossRef] [Scilit]
- Tanveer, M.; Azad, M.M.; Kim, D.; Khalid, S. Generative Design for Engineering Applications: A State-of-the-Art Review. Arch. Comput. Methods Eng. 2025, 33, 53–79. [Google Scholar] [CrossRef] [Scilit]
- Williams, B.; Bannett, M. Interpretable Machine Learning under Evolving Fraud Regimes with Human-in-the-Loop Adaptation. J. Data Driven Insights 2025, 2, 1–15. [Google Scholar]
- Xu, G.; Chen, Q.; Ling, C.; Wang, B.; Shui, C. Intersectional unfairness discovery. arXiv 2024, arXiv:2405.20790. [Google Scholar] [CrossRef] [Scilit]
- Khan, H.; Khalid, A.F.; Hassan, Z. Transcending Controlled Environments: Assessing the Transferability of ASR-Robust NLU Models to Real-World Applications. arXiv 2024, arXiv:2401.09354. [Google Scholar]
- Jeong, C.; Sim, S.M.; Cho, H.; Kim, S.; Shin, B.C. E2E Process Automation Leveraging Generative AI and IDP-Based Automation Agent: A Case Study on Corporate Expense Processing. Artif. Intell. Appl. 2025, 4, 102–115. [Google Scholar] [CrossRef] [Scilit]
- Wyllie, S.; Shumailov, I.; Papernot, N. Fairness Feedback Loops: Training on Synthetic Data Amplifies Bias. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, Rio de Janeiro, Brazil, 3–6 June 2024; pp. 1569–1586. [Google Scholar] [CrossRef] [Scilit]
- Tiwari, A.; Farag, H.E.Z. Responsible AI Framework for Autonomous Vehicles: Addressing Bias and Fairness Risks. IEEE Access 2025, 13, 58800–58822. [Google Scholar] [CrossRef] [Scilit]
- Roy, A.; Raghunandan, D.; Elmqvist, N.; Battle, L. How I Met Your Data Science Team: A Tale of Effective Communication. In Proceedings of the 2023 IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC), Washington, DC, USA, 3–6 October 2023; pp. 160–170. [Google Scholar] [CrossRef] [Scilit]
- Hewett, T.; Meadow, C.T. On designing for usability: An application of four key principles. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (CHI ‘86), Boston, MA, USA, 13–17 April 1986; pp. 247–252. [Google Scholar] [CrossRef] [Scilit]
- Samoili, S.; Cobo, M.L.; Gómez, E.; De Prato, G.; Martínez-Plumed, F.; Delipetrev, B. AI Watch. Defining Artificial Intelligence. Towards an Operational Definition and Taxonomy of Artificial Intelligence; EUR 30117 EN; Publications Office of the European Union: Luxembourg, 2020. [Google Scholar]
- Ismatullaev, U.; Kim, K. Scenario to specification: Promises and pitfalls of AI in developing user-centered engineering specifications with interdisciplinary teams. Proc. Des. Soc. 2025, 5, 2841–2850. [Google Scholar] [CrossRef] [Scilit]
- Akinrinola, O.; Okoye, C.C.; Ofodile, O.C. Navigating and reviewing ethical dilemmas in AI development: Strategies for transparency, fairness, and accountability. Eng. Sci. Technol. J. 2024, 5, 2382–2407. [Google Scholar]
- Behrouzi, B.; Liu, X.; Tweed, D. Costate-focused models for reinforcement learning. arXiv 2017, arXiv:1708.05611. [Google Scholar]
- Kather, J.N.; Laleh, N.G.; Foersch, S.; Truhn, D. Medical domain knowledge in domain-agnostic generative AI. npj Digit. Med. 2022, 5, 90. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Head, C.B.; Jasper, P.; McConnachie, M.M.; Raftree, L.; Higdon, G.L. Large language model applications for evaluation: Opportunities and ethical implications. New Dir. Eval. 2023, 2023, 33–46. [Google Scholar] [CrossRef] [Scilit]
- Gu, X.; Pang, T.; Du, C.; Liu, Q.; Zhang, F.; Du, C.; Wang, Y.; Lin, M. When Attention Sink Emerges in Language Models: An Empirical View. arXiv 2024, arXiv:2410.10781. [Google Scholar] [CrossRef] [Scilit]
- Flower, L. Cognition, Context, and Theory Building. Coll. Compos. Commun. 1989, 40, 282–311. [Google Scholar] [CrossRef] [Scilit]
- Claassen, J.A. The Gold Standard: Not a Golden Standard. BMJ 2005, 330, 1121. [Google Scholar] [CrossRef] [Scilit]
- Yin, J.; Ngiam, K.Y.; Teo, H.H. Role of artificial intelligence applications in real-life clinical practice: Systematic review. J. Med. Internet Res. 2021, 23, e25759. [Google Scholar] [CrossRef] [Scilit]
- Rezwana, J.; Maher, M.L. Identifying Ethical Issues in AI Partners in Human-AI Co-Creation. arXiv 2022, arXiv:2204.07644. [Google Scholar] [CrossRef] [Scilit]
- Pedreschi, D.; Giannotti, F.; Guidotti, R.; Monreale, A.; Ruggieri, S.; Turini, F. Meaningful explanations of black box AI decision systems. In Proceedings of the AAAI 2019, Honolulu, HI, USA, 27 January–1 February 2019; Volume 33, pp. 9780–9784. [Google Scholar]
- Dai, J.; Pan, X.; Sun, R.; Ji, J.; Xu, X.; Liu, M.; Wang, Y.; Yang, Y. Safe RLHF: Safe Reinforcement Learning from Human Feedback. In Proceedings of the 12th International Conference on Learning Representations (ICLR 2024), Vienna, Austria, 7–11 May 2024. [Google Scholar]
- Mosconi, G.; de Carvalho, A.F.; Syed, H.Z. Fostering research data management in collaborative research contexts: Lessons learnt from an ‘Embedded’ Evaluation of “data story”. Int. J. Digit. Curation 2023, 17, 15. [Google Scholar] [CrossRef] [Scilit]
- Mannava, M.K. Causal Inference in AI Based Decision Support: Beyond Correlation to Causation. In Proceedings of the 2024 4th International Conference on Ubiquitous Computing and Intelligent Information Systems (ICUIS), Gobichettipalayam, India, 12–13 December 2024; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
- Elling, S.; Lentz, L.; de Jong, M. Combining Concurrent Think-Aloud Protocols and Eye-Tracking Observations: An Analysis of Verbalizations and Silences. IEEE Trans. Prof. Commun. 2012, 55, 206–220. [Google Scholar] [CrossRef] [Scilit]
- Borghoff, U.M.; Bottoni, P.; Pareschi, R. Beyond Prompt Chaining: The tb-cspn Architecture for Agentic AI. Future Internet 2025, 17, 363. [Google Scholar] [CrossRef] [Scilit]
- Park, C.; Liu, M.; Kong, D.; Zhang, K.; Ozdaglar, A. RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation. In Proceedings of the 41st International Conference on Machine Learning (ICML 2024), Vienna, Austria, 21–27 July 2024; pp. 39651–39678. [Google Scholar]
- Kadenhe, N.; Musleh, M.A.; Lompot, A. Human-AI Co-Design and Co-Creation: A Review of Emerging Approaches, Challenges, and Future Directions. AAAI Spring Symp. Ser. 2025, 4, 154–162. [Google Scholar] [CrossRef] [Scilit]
- Cai, H.; Cai, X.; Chang, J.; Li, S.; Yao, L.; Wang, C.; Gao, Z.; Wang, H.; Li, Y.; Lin, M.; et al. SciAssess: Benchmarking LLM Proficiency in Scientific Literature Analysis. In Proceedings of the NAACL-HLT 2024, Mexico City, Mexico, 16–21 June 2024; pp. 4351–4380. [Google Scholar]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.



