Evaluation of Large Language Models as Tools, Models, and Partners in Creative Thinking Research: A Selective Narrative Review with the GCA Framework
Abstract
1. Introduction
2. Literature Selection Strategy and Methods
3. Foundations of Creative Thinking and the Rise of LLMs
3.1. Theoretical Frameworks of Creative Cognition
3.2. Neurocognitive Architecture of Human Creativity
3.3. Transformer Architectures as Associative Engines
4. LLMs as Creative Idea Generators
4.1. Benchmarking LLM Creativity on Standard and Alternative Tasks
4.2. Prompt Engineering and the Controllability of Creative Output
4.3. The Novelty–Typicality Trade-Off and Decoding Strategies
4.4. Evidence Against Attributing Creativity to LLMs
4.5. Contradictory Findings and Boundary Conditions
5. LLMs as Cognitive Models of Creative Thought
5.1. Marr’s Levels Applied to LLM Creativity
5.2. Empirical Evidence for and Against LLMs as Cognitive Models
5.3. Simulation Versus Instantiation: A Methodological Framework
6. LLMs as Automated Creativity Assessment Tools
6.1. Semantic Distance as a Computational Proxy for Originality
6.2. Automated Scoring Systems for Divergent Thinking
6.3. Cross-Cultural Validity and the Cross-Linguistic Reliability Gap
7. Human–AI Co-Creativity and Collaborative Ideation
7.1. Mechanisms of Effective Human–AI Creative Collaboration
7.2. Anchoring Effects and the Homogenisation Risk
8. Methodological Challenges, Ethical Concerns, and Open Questions
8.1. Training-Data Contamination and Benchmark Validity
8.2. Ethics of AI Creativity: Authorship, Bias, and Environmental Costs
8.3. Future Directions: Intentionality, Evaluation, and Ecological Validity
8.4. Limitations of the Present Review
9. Conclusions
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| Abbreviation | Full Form |
| AI | Artificial Intelligence |
| AUT | Alternative Uses Task |
| CAN | Creative Adversarial Network |
| CAT | Consensual Assessment Technique |
| COFI | Co-Creative Framework for Interaction |
| DMN | Default Mode Network |
| ECN | Executive Control Network |
| FIQ | Figural Interpretation Quest |
| GCA | Generation–Capability–Assessment |
| ICC | Intraclass Correlation Coefficient |
| LLM | Large Language Model |
| OCSAI | Organisciak’s Computational Scoring of Alternate Uses and Ideation |
| PRISMA | Preferred Reporting Items for Systematic Reviews and Meta-Analyses |
| TTCT | Torrance Tests of Creative Thinking |
| DAT | Divergent Association Task |
| fNIRS | Functional Near-Infrared Spectroscopy |
| LSA | Latent Semantic Analysis |
References
- Amabile, T. M. (1982). Social psychology of creativity: A consensual assessment technique. Journal of Personality and Social Psychology, 43(5), 997–1013. [Google Scholar] [CrossRef]
- Antony, V. N., & Huang, C.-M. (2024). ID.8: Co-creating visual stories with generative AI. ACM Transactions on Interactive Intelligent Systems, 14(3), 1–29. [Google Scholar] [CrossRef] [Scilit]
- Aru, J., Drüke, M., Pikamäe, J., & Larkum, M. E. (2023). Mental navigation and the neural mechanisms of insight. Trends in Neurosciences, 46(2), 100–109. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Baer, J. (2012). Domain specificity and the limits of creativity theory. The Journal of Creative Behavior, 46(1), 16–29. [Google Scholar] [CrossRef] [Scilit]
- Bearman, M., Tai, J., Dawson, P., Boud, D., & Ajjawi, R. (2024). Developing evaluative judgement for a time of generative artificial intelligence. Assessment & Evaluation in Higher Education, 49(6), 893–905. [Google Scholar] [CrossRef] [Scilit]
- Beaty, R. E., & Johnson, D. R. (2021). Automating creativity assessment with SemDis: An open platform for computing semantic distance. Behavior Research Methods, 53(2), 757–780. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Beaty, R. E., Johnson, D. R., Zeitlen, D. C., & Forthmann, B. (2022). Semantic distance and the alternate uses task: Recommendations for reliable automated assessment of originality. Creativity Research Journal, 34(3), 245–260. [Google Scholar] [CrossRef] [Scilit]
- Beaty, R. E., & Kenett, Y. N. (2023). Associative thinking at the core of creativity. Trends in Cognitive Sciences, 27(7), 671–683. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Beaty, R. E., Kenett, Y. N., Christensen, A. P., Rosenberg, M. D., Benedek, M., Chen, Q., Fink, A., Qiu, J., Kwapil, T. R., Kane, M. J., & Silvia, P. J. (2018). Robust prediction of individual creative ability from brain functional connectivity. Proceedings of the National Academy of Sciences of the United States of America, 115(5), 1087–1092. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Bellemare-Pepin, A., Lespinasse, F., Thölke, P., Harel, Y., Mathewson, K., Olson, J. A., Bengio, Y., & Jerbi, K. (2026). Divergent creativity in humans and large language models. Scientific Reports, 16, 1279. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency (pp. 610–623). Association for Computing Machinery. [Google Scholar] [CrossRef] [Scilit]
- Benedek, M., & Fink, A. (2019). Toward a neurocognitive framework of creative cognition: The role of memory, attention, and cognitive control. Current Opinion in Behavioral Sciences, 27, 116–122. [Google Scholar] [CrossRef] [Scilit]
- Boden, M. A. (2004). The creative mind: Myths and mechanisms (2nd ed.). Routledge. [Google Scholar]
- Brandt, A. K. (2025). Amplifying the anomaly: How humans choose unproven options and large language models avoid them. Creativity Research Journal, 37(4), 582–603. [Google Scholar] [CrossRef] [Scilit]
- Breithaupt, F., Otenen, E., Wright, D. R., Kruschke, J. K., Li, Y., & Tan, Y. (2024). Humans create more novelty than ChatGPT when asked to retell a story. Scientific Reports, 14(1), 875. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Brickman, J., Gupta, M., & Oltmanns, J. R. (2025). Large language models for psychological assessment: A comprehensive overview. Advances in Methods and Practices in Psychological Science, 8(3), 25152459251343582. [Google Scholar] [CrossRef] [Scilit]
- Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., … Amodei, D. (2020). Language models are few-shot learners. In Advances in neural information processing systems (Vol. 33, pp. 1877–1901). Curran Associates, Inc. [Google Scholar]
- Cheng, X., & Zhang, L. (2025). Inspiration booster or creative fixation? The dual mechanisms of LLMs in shaping individual creativity in tasks of different complexity. Humanities and Social Sciences Communications, 12, 1563. [Google Scholar] [CrossRef] [Scilit]
- Cheng, Y., & Huang, X. (2026). Beyond cognitive support: Social interaction and affective experience in AI-assisted creative thinking among university students. Journal of Intelligence, 14(7), 151. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chrysikou, E. G. (2019). Creativity in and out of (cognitive) control. Current Opinion in Behavioral Sciences, 27, 94–99. [Google Scholar] [CrossRef] [Scilit]
- Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Lawrence Erlbaum Associates. [Google Scholar]
- Cropley, A. (2006). In praise of convergent thinking. Creativity Research Journal, 18(3), 391–404. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Demszky, D., Yang, D., Yeager, D. S., Bryan, C. J., Clapper, M., Chandhok, S., Eichstaedt, J. C., Hecht, C., Jamieson, J., Johnson, M., Jones, M., Krettek-Cobb, D., Lai, L., Jones-Mitchell, N., Ong, D. C., Dweck, C. S., Gross, J. J., & Pennebaker, J. W. (2023). Using large language models in psychology. Nature Reviews Psychology, 2, 255–266. [Google Scholar] [CrossRef] [Scilit]
- Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the Association for Computational Linguistics: Human language technologies (pp. 4171–4186). Association for Computational Linguistics. [Google Scholar] [CrossRef] [Scilit]
- Doshi, R. M., & Hauser, O. P. (2024). Generative AI enhances individual creativity but reduces the collective diversity of novel content. Science Advances, 10(28), eadn5290. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Elgammal, A., Liu, B., Elhoseiny, M., & Mazzone, M. (2017). CAN: Creative adversarial networks, generating “art” by learning about styles and deviating from style norms. In Proceedings of the eighth international conference on computational creativity (pp. 96–103). Association for Computational Creativity. [Google Scholar]
- Fazi, M. B. (2018). Can a machine think (anything new)? Automation beyond simulation. AI & Society, 34(4), 813–824. [Google Scholar] [CrossRef] [Scilit]
- Finke, R. A., Ward, T. B., & Smith, S. M. (1992). Creative cognition: Theory, research, and applications. MIT Press. [Google Scholar]
- Glăveanu, V. P. (2014). Distributed creativity: Thinking outside the box of the creative individual. Springer. [Google Scholar] [CrossRef] [Scilit]
- Goecke, B., DiStefano, P. V., Aschauer, W., Haim, K., Beaty, R., & Forthmann, B. (2024). Automated scoring of scientific creativity in German. The Journal of Creative Behavior, 58(3), 321–327. [Google Scholar] [CrossRef] [Scilit]
- Grassini, S., & Koivisto, M. (2025). Artificial creativity? Evaluating AI against human performance in creative interpretation of visual stimuli. International Journal of Human–Computer Interaction, 41(7), 4037–4048. [Google Scholar] [CrossRef] [Scilit]
- Green, A. E., Beaty, R. E., Kenett, Y. N., & Kaufman, J. C. (2024). The process definition of creativity. Creativity Research Journal, 36(3), 544–572. [Google Scholar] [CrossRef] [Scilit]
- Guilford, J. P. (1950). Creativity. American Psychologist, 5(9), 444–454. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Guo, Y., Lin, S., Williams, Z. J., Grantham, T. C., Guo, J., Cole Clark, L. Q., & Zou, W. (2024). Creative potential and creative self-belief: Measurement invariance in cross-cultural contexts. The Journal of Creative Behavior, 58(2), 209–226. [Google Scholar] [CrossRef] [Scilit]
- Han, S. J., Ransom, K., Perfors, A., & Kemp, C. (2024). Inductive reasoning in humans and large language models. Cognitive Systems Research, 83, 101155. [Google Scholar] [CrossRef] [Scilit]
- He, L., Kenett, Y. N., Zhuang, K., Liu, C., Zeng, R., Yan, T., Huo, T., & Qiu, J. (2021). The relation between semantic memory structure, associative abilities, and verbal and figural creativity. Thinking & Reasoning, 27(2), 268–293. [Google Scholar] [CrossRef] [Scilit]
- Hubert, K. F., Awa, K. N., & Zabelina, D. L. (2024). The current state of artificial intelligence generative language models is more creative than humans on divergent thinking tasks. Scientific Reports, 14, 3440. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kaufman, J. C., & Baer, J. (2004). Sure I’m creative—But not in mathematics! Self-reported creativity in diverse domains. Empirical Studies of the Arts, 22(2), 143–155. [Google Scholar] [CrossRef] [Scilit]
- Kaufman, J. C., & Beghetto, R. A. (2009). Beyond big and little: The Four C model of creativity. Review of General Psychology, 13(1), 1–12. [Google Scholar] [CrossRef] [Scilit]
- Kenett, Y. N., Anaki, D., & Faust, M. (2014). Investigating the structure of semantic networks in low and high creative persons. Frontiers in Human Neuroscience, 8, 407. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Koivisto, M., & Grassini, S. (2023). Best humans still outperform artificial intelligence in a creative divergent thinking task. Scientific Reports, 13, 13601. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kostikova, A., Wang, Z., Bajri, D., Pütz, O., Paaßen, B., & Eger, S. (2026). LLLMs: A data-driven survey of evolving research on limitations of large language models. ACM Computing Surveys, 58(11), 282. [Google Scholar] [CrossRef] [Scilit]
- Kozbelt, A., Beghetto, R. A., & Runco, M. A. (2010). Theories of creativity. In J. C. Kaufman, & R. J. Sternberg (Eds.), The Cambridge handbook of creativity (pp. 20–47). Cambridge University Press. [Google Scholar]
- Lazovsky, G. S., Raz, T., & Kenett, Y. N. (2025). The art of creative inquiry—From question asking to prompt engineering. The Journal of Creative Behavior, 59(1), e671. [Google Scholar] [CrossRef] [Scilit]
- Lee, B. C., & Chung, J. (2024). An empirical investigation of the impact of ChatGPT on creativity. Nature Human Behaviour, 8(10), 1906–1914. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lin, Z. (2026). A validity-guided workflow for robust LLM research in psychology. Behavior Research Methods, 58, 216. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Liu, C., Ren, Z., Zhuang, K., He, L., Yan, T., Zeng, R., & Qiu, J. (2021). Semantic association ability mediates the relationship between brain structure and human creativity. Neuropsychologia, 151, 107722. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lloyd-Cox, J., Chen, Q., & Beaty, R. E. (2022). The time course of creativity: Multivariate classification of default and executive network contributions to creative cognition over time. Cortex, 156, 90–105. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lou, J., & Sun, Y. (2026). Anchoring bias in large language models: An experimental study. Journal of Computational Social Science, 9, 11. [Google Scholar] [CrossRef] [Scilit]
- Magni, F., Park, J., & Chao, M. M. (2024). Humans as creativity gatekeepers: Are we biased against AI creativity? Journal of Business and Psychology, 39(3), 643–656. [Google Scholar] [CrossRef] [Scilit]
- Mahowald, K., Ivanova, A. A., Blank, I. A., Kanwisher, N., Tenenbaum, J. B., & Fedorenko, E. (2024). Dissociating language and thought in large language models. Trends in Cognitive Sciences, 28(6), 517–540. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Marr, D. (1982). Vision: A computational investigation into the human representation and processing of visual information. W. H. Freeman. [Google Scholar]
- McCoy, R. T., Smolensky, P., Linzen, T., Gao, J., & Celikyilmaz, A. (2023). How much do language models copy from their training data? Evaluating linguistic novelty in text generation using RAVEN. Transactions of the Association for Computational Linguistics, 11, 652–670. [Google Scholar] [CrossRef] [Scilit]
- Mednick, S. (1962). The associative basis of the creative process. Psychological Review, 69(3), 220–232. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Nusbaum, E. C., & Silvia, P. J. (2011). Are intelligence and creativity really so different?: Fluid intelligence, executive processes, and strategy use in divergent thinking. Intelligence, 39(1), 36–45. [Google Scholar] [CrossRef] [Scilit]
- Organisciak, P., Acar, S., Dumas, D., & Berthiaume, K. (2023). Beyond semantic distance: Automated scoring of divergent thinking greatly improves with large language models. Thinking Skills and Creativity, 49, 101356. [Google Scholar] [CrossRef] [Scilit]
- Orwig, W., Diez, I., Vannini, P., Beaty, R., & Sepulcre, J. (2021). Creative connections: Computational semantic distance captures individual creativity and resting-state functional connectivity. Journal of Cognitive Neuroscience, 33(3), 499–509. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Orwig, W., Edenbaum, E. R., Greene, J. D., & Schacter, D. L. (2024). The language of creativity: Evidence from humans and large language models. The Journal of Creative Behavior, 58(1), 128–136. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Orwig, W., Luchini, S. A., Beaty, R. E., & Schacter, D. L. (2025). A “sweet spot” for creative ideation: Non-linear associations between semantic distance and creativity. The Journal of Creative Behavior, 59(3), e70041. [Google Scholar] [CrossRef] [Scilit]
- Page, M. J., McKenzie, J. E., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., Shamseer, L., Tetzlaff, J. M., Akl, E. A., Brennan, S. E., Chou, R., Glanville, J., Grimshaw, J. M., Hróbjartsson, A., Lalu, M. M., Li, T., Loder, E. W., Mayo-Wilson, E., McDonald, S., … Moher, D. (2021). The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ, 372, n71. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Park, J., & Choo, S. (2024). Generative AI prompt engineering for educators. Journal of Special Education Technology, 40(3), 411–417. [Google Scholar] [CrossRef] [Scilit]
- Perchtold-Stefan, C. M., Papousek, I., Rominger, C., Schertler, M., Weiss, E. M., & Fink, A. (2020). Humor comprehension and creative cognition: Shared and distinct neurocognitive mechanisms as indicated by EEG alpha activity. NeuroImage, 213, 116695. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Rezwana, J., & Maher, M. L. (2023). Designing creative AI partners with COFI: A framework for modeling interaction in human-AI co-creative systems. ACM Transactions on Computer-Human Interaction, 30(5), 67. [Google Scholar] [CrossRef] [Scilit]
- Rhodes, M. (1961). An analysis of creativity. The Phi Delta Kappan, 42(7), 305–310. [Google Scholar]
- Ruan, K., Wang, X., Hong, J., Wang, P., Liu, Y., & Sun, H. (2026). Evaluating LLMs’ divergent thinking capabilities for scientific idea generation with minimal context. Nature Communications, 17, 3625. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Runco, M. A., & Jaeger, G. J. (2012). The standard definition of creativity. Creativity Research Journal, 24(1), 92–96. [Google Scholar] [CrossRef] [Scilit]
- Saretzki, J., & Benedek, M. (2026). Investigating the validity evidence of automated scoring methods for divergent thinking assessments. Creativity Research Journal, 38(1), 1–17. [Google Scholar] [CrossRef] [Scilit]
- Scott, G. M., Leritz, E. C., & Mumford, M. D. (2004). The effectiveness of creativity training: A quantitative review. Creativity Research Journal, 16(4), 361–388. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Silvia, P. J., Winterstein, B. P., Willse, J. T., Barona, C. M., Cram, J. T., Hess, K. I., Martinez, J. L., & Richard, C. A. (2008). Assessing creativity with divergent thinking tasks: Exploring the reliability and validity of new subjective scoring methods. Psychology of Aesthetics, Creativity, and the Arts, 2(2), 68–85. [Google Scholar] [CrossRef] [Scilit]
- Stanko-Kaczmarek, M., Dera, L., & Koscielska, H. (2025). “Between the lines”: Perceptions of poetry with authorship attributed to artificial intelligence or humans—A comparative analysis. The Journal of Creative Behavior, 59(3), e1513. [Google Scholar] [CrossRef] [Scilit]
- Stevenson, C., Smal, I., Baas, M., Grasman, R., & van der Maas, H. (2022). Putting GPT-3’s creativity to the (alternative uses) test. In Proceedings of the 13th international conference on computational creativity (pp. 164–168). Association for Computational Creativity. [Google Scholar]
- Tang, M., Hofreiter, S., Werner, C. H., Zielińska, A., & Karwowski, M. (2025). “Who” is the best creative thinking partner? An experimental investigation of human–human, human–internet, and human–AI co-creation. The Journal of Creative Behavior, 59(3), e1519. [Google Scholar] [CrossRef] [Scilit]
- Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2017). Attention is all you need. In Advances in neural information processing systems (Vol. 30, pp. 5998–6008). Curran Associates, Inc. [Google Scholar]
- Vinchon, F., Gironnay, V., & Lubart, T. (2024). GenAI creativity in narrative tasks: Exploring new forms of creativity. Journal of Intelligence, 12(12), 125. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, P., Zhang, X., Wei, L., Childs, P., Wang, S. J., Guo, Y., & Kleinsmann, M. (2026). Human–AI co-ideation via combinational generative model. Journal of Engineering Design, 37(2), 458–494. [Google Scholar] [CrossRef] [Scilit]
- Wang, X., Chen, Q., Zhuang, K., Zhang, J., Cortes, R. A., Holzman, D. D., Fan, L., Liu, C., Sun, J., Li, X., Li, Y., Feng, Q., Chen, H., Feng, T., Lei, X., He, Q., Green, A. E., & Qiu, J. (2024). Semantic associative abilities and executive control functions predict novelty and appropriateness of idea generation. Communications Biology, 7(1), 703. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wingström, R., Hautala, J., & Lundman, R. (2024). Redefining creativity in the era of AI? Perspectives of computer scientists and new media artists. Creativity Research Journal, 36(2), 177–193. [Google Scholar] [CrossRef] [Scilit]
- Xu, C., Sun, Y., & Zhou, H. (2025). Artificial aesthetics and ethical ambiguity: Exploring business ethics in the context of AI-driven creativity. Journal of Business Ethics, 199, 671–692. [Google Scholar] [CrossRef] [Scilit]
- Yang, T., Zhang, Q., Sun, Z., & Hou, Y. (2023). Automatic assessment of divergent thinking in Chinese language with TransDis: A transformer-based language model approach. Behavior Research Methods, 56(6), 5798–5819. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhao, C., Habule, M., & Zhang, W. (2025). Large language models (LLMs) as research subjects: Status, opportunities, and challenges. New Ideas in Psychology, 79, 101167. [Google Scholar] [CrossRef] [Scilit]
- Zhou, Z., Qiao, X., Zhang, W., Tong, S., & Hao, N. (2026). The effects of AI viewpoint divergence on group creative performance and its cognitive-neural mechanisms. Acta Psychologica Sinica. Advance online publication. [Google Scholar]



| Criterion | Type | Justification |
|---|---|---|
| LLMs as primary tool, subject, or object | Inclusion | Ensures relevance to LLM–creativity nexus |
| Addresses ≥1 GCA dimension | Inclusion | Aligns with the review’s evaluative framework |
| Published in English, 2018–2026 | Inclusion | Captures the transformer era and ensures accessibility |
| Non-LLM AI systems only | Exclusion | Outside the scope of LLM-specific evaluation |
| Metaphorical creativity | Exclusion | Lacks operationalised creativity measurement |
| Editorials, commentaries, book reviews | Exclusion | No original empirical or theoretical contribution |
| Study | Task | Measure | g | 95% CI | Note |
|---|---|---|---|---|---|
| Hubert et al. (2024) | AUT | Originality | 2.61 | [2.30, 2.91] | Independent; post hoc fluency matching |
| Hubert et al. (2024) | AUT | Elaboration | 4.07 | [3.68, 4.46] | Same sample |
| Hubert et al. (2024) | Consequences | Originality | 1.47 | [1.22, 1.72] | Same sample |
| Grassini and Koivisto (2025) | FIQ | Flexibility | 0.74 | [0.56, 0.92] | Independent |
| Grassini and Koivisto (2025) | FIQ | Subjective creativity | −1.31 | [−1.50, −1.11] | Same sample; human advantage |
| Koivisto and Grassini (2023) | AUT | Originality | 0.70 | [0.33, 1.06] | Independent; prospective fluency matching |
| Bellemare-Pepin et al. (2026) | DAT | Originality vs. population average | 0.52 | [0.39, 0.64] | Independent; 503 model responses vs. 500 drawn from 100,000 humans |
| Bellemare-Pepin et al. (2026) | DAT | Originality vs. top 50% of humans | −0.88 | [−1.01, −0.75] | Same sample; human advantage |
| Study Design | Applicable GCA Axis/Axes | Criteria to Apply (Heuristic Benchmarks) | Required Reporting Items |
|---|---|---|---|
| Generation-focused comparison (LLM vs. human) | Generation | Fluency; originality (d ≥ 0.50, heuristic); flexibility; elaboration | Full prompt; model name and version; temperature; number of responses per prompt; fluency-control procedure (prospective vs. post hoc); average-level and peak-level comparisons; task and scoring method; contamination-risk classification (Table S3) |
| Assessment-tool validation | Assessment | Inter-rater reliability (ICC ≥ 0.80, heuristic); semantic validity (r ≥ 0.70, heuristic); cultural fairness | ICC form, unit, and confidence interval; blinding of human raters; embedding-model language and target language; group-level vs. individual-level use claims; language-specific calibration evidence; construct-equivalence evidence for cross-cultural use |
| Cognitive-modelling study | Capability | Architectural fidelity; mechanistic plausibility; experimental utility (Marr’s three levels) | Level at which correspondence is claimed (computational, algorithmic, implementation); simulation vs. instantiation stance; falsifiable predictions tested |
| Co-creativity study | Generation + interaction quality | Ideation efficiency; idea diversity; participant satisfaction; diversity preservation | Interaction, contribution, and adaptation models; semantic spread of human ideas before and after AI exposure; task complexity; divergent vs. convergent outcomes |
| Training-data contamination screening (all designs) | All axes | Graded exposure risk (low/moderate/high) | Model type (open vs. proprietary); corpus documentation status; stimulus provenance (post-cutoff, dynamic, or classic); holdout or exclusion-list procedures |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Huang, K.; Liu, C.; Yang, J. Evaluation of Large Language Models as Tools, Models, and Partners in Creative Thinking Research: A Selective Narrative Review with the GCA Framework. J. Intell. 2026, 14, 218. https://doi.org/10.3390/jintelligence14090218
Huang K, Liu C, Yang J. Evaluation of Large Language Models as Tools, Models, and Partners in Creative Thinking Research: A Selective Narrative Review with the GCA Framework. Journal of Intelligence. 2026; 14(9):218. https://doi.org/10.3390/jintelligence14090218
Chicago/Turabian StyleHuang, Kexin, Chunlei Liu, and Jiaqin Yang. 2026. "Evaluation of Large Language Models as Tools, Models, and Partners in Creative Thinking Research: A Selective Narrative Review with the GCA Framework" Journal of Intelligence 14, no. 9: 218. https://doi.org/10.3390/jintelligence14090218
APA StyleHuang, K., Liu, C., & Yang, J. (2026). Evaluation of Large Language Models as Tools, Models, and Partners in Creative Thinking Research: A Selective Narrative Review with the GCA Framework. Journal of Intelligence, 14(9), 218. https://doi.org/10.3390/jintelligence14090218

