Advanced Studies in Human-Centred AI

A Special Issue of Behavioral Sciences (ISSN 2076-328X) belonging to the section "Cognition".

Deadline for manuscript submissions: closed (31 October 2025) | Viewed by 15385

Editor


E-Mail Website
Guest Editor
School of Psychology, Plymouth University, Drake Circus, Plymouth PL4 4AG, UK
Interests: categorization; attention; artificial intelligence
Special Issues, Collections and Topics in MDPI journals

Special Issue Information

Dear Colleagues,

Human-centred AI puts human behaviour and experience at the heart of artificial intelligence research. For example, it is sometimes claimed that artificial neural networks (ANNs) now perform at human levels in a variety of tasks—to what extent is this claim substantiated by the evidence? Does that performance, human-level or otherwise, extend to replicating (or perhaps amplifying) well-documented biases in human decision-making? Can ANNs effectively and safely be used to support the work of highly trained professionals—for example, radiologists, therapists, legal advisors, or researchers? Can we effectively adapt the skills and techniques of behavioural research, previously applied to humans and other animals, to better understand the ‘psychology’ of complex black-box ANNs? To what extent can our understanding of how humans explain their decisions inform explainable AI? What makes an AI system seem trustworthy, and is that trust well placed? Can people spontaneously distinguish real photographs and videos from deepfakes—and, if not, can they be trained to do so? Can work on goal-setting and reinforcement learning in humans inform agentic behaviour and AI alignment? If the technical issues of AI alignment are indeed solvable, to what values should they be aligned? We welcome original papers on these and related topics in human-centred AI. The papers may be theoretical, empirical, or both. They may report new findings, or synthetically review the existing literature.

Prof. Dr. Andy J. Wills
Guest Editor

Manuscript Submission Information

Manuscripts should be submitted online at www.mdpi.com by registering and logging in to this website. Once you are registered, click here to go to the submission form. Manuscripts can be submitted until the deadline. All submissions that pass pre-check are peer-reviewed. Accepted papers will be published continuously in the journal (as soon as accepted) and will be listed together on the special issue website. Research articles, review articles as well as short communications are invited. For planned papers, a title and short abstract (about 250 words) can be sent to the Editorial Office for assessment.

Submitted manuscripts should not have been published previously, nor be under consideration for publication elsewhere (except conference proceedings papers). All manuscripts are thoroughly refereed through a single-anonymized peer-review process. A guide for authors and other relevant information for submission of manuscripts is available on the Instructions for Authors page. Behavioral Sciences is an international peer-reviewed open access monthly journal published by MDPI.

Please visit the Instructions for Authors page before submitting a manuscript. The Article Processing Charge (APC) for publication in this open access journal is 2400 CHF (Swiss Francs). Submitted papers should be well formatted and use good English. Authors may use MDPI's English editing service prior to publication or during author revisions.

Keywords

  • human-centred AI
  • neural networks
  • AGI
  • bias
  • decision-making
  • therapy
  • healthcare
  • experimental psychology
  • explainable AI
  • trust and trustworthiness
  • deepfake detection
  • AI alignment

Benefits of Publishing in a Special Issue

  • Ease of navigation: Grouping papers by topic helps scholars navigate broad scope journals more efficiently.
  • Greater discoverability: Special Issues support the reach and impact of scientific research. Articles in Special Issues are more discoverable and cited more frequently.
  • Expansion of research network: Special Issues facilitate connections among authors, fostering scientific collaborations.
  • External promotion: Articles in Special Issues are often promoted through the journal's social media, increasing their visibility.
  • Reprint: MDPI Books provides the opportunity to republish successful Special Issues in book format, both online and in print.

Further information on MDPI's Special Issue policies can be found here.

Published Papers (9 papers)

Order results
Result details
Select all
Export citation of selected articles as:

Research

Jump to: Review

39 pages, 8250 KB  
Article
Discerning Quantity: Numerosity in Two Embodied Machine Learning Agents
by Niall Donnelly and Edward Keedwell
Behav. Sci. 2026, 16(5), 813; https://doi.org/10.3390/bs16050813 - 19 May 2026
Viewed by 320
Abstract
As artificial intelligence systems continue to overcome evermore challenging tasks, researchers have suggested that the time is ripe to begin evaluating these systems along more psychologically inspired lines. This study seeks to build upon these recommendations by evaluating two machine learning models, A-Learning [...] Read more.
As artificial intelligence systems continue to overcome evermore challenging tasks, researchers have suggested that the time is ripe to begin evaluating these systems along more psychologically inspired lines. This study seeks to build upon these recommendations by evaluating two machine learning models, A-Learning and Proximal Policy Optimisation, for the cognitive capability known as numerosity. In our experiment, these two models were embodied in a three-dimensional virtual environment, known as Animal-AI, and tested in a psychologically inspired numerosity experiment. In contrast to previous research, A-Learning failed to reliably express numerosity capabilities, as did Proximal Policy Optimisation. Both models displayed a tendency to overfit to the first policy that provided rewarding feedback. These results suggest that predicting the cognitive capabilities of machine learning models once embodied is non-trivial, and confounding factors such as environmental properties and perceptual processes complicate the expression of numerosity capabilities. Building on these findings, it is suggested that future researchers pay greater attention to the influence of environmental factors and perceptual mechanisms on the machine learning models they are developing, especially if such models are to be embodied in a virtual- or real-world environment. Full article
(This article belongs to the Special Issue Advanced Studies in Human-Centred AI)
Show Figures

Figure 1

24 pages, 13819 KB  
Article
What Does ‘Human-Centred AI’ Mean?
by Olivia Guest
Behav. Sci. 2026, 16(4), 583; https://doi.org/10.3390/bs16040583 - 13 Apr 2026
Cited by 3 | Viewed by 2578
Abstract
While it seems sensible that human-centred artificial intelligence (AI) means centring “human behaviour and experience,” it cannot be any other way. AI, I argue, is usefully seen as a relationship between technology and humans where it appears that artefacts can perform, to a [...] Read more.
While it seems sensible that human-centred artificial intelligence (AI) means centring “human behaviour and experience,” it cannot be any other way. AI, I argue, is usefully seen as a relationship between technology and humans where it appears that artefacts can perform, to a greater or lesser extent, human cognitive labour. This is evinced using examples that juxtapose technology with cognition, inter alia: abacus versus mental arithmetic; alarm clock versus knocker-upper; camera versus vision; and sweatshop versus tailor. Using novel definitions and analyses, sociotechnical relationships can be seen as varying types of: displacement (harmful), enhancement (beneficial), and/or replacement (neutral) of human cognitive labour. Ultimately, all AI implicates human cognition; no matter what. Obfuscation of cognition in the AI context—from clocks to artificial neural networks—results in distortion, in slowing critical engagement, perverting cognitive science, and indeed in limiting our ability to truly centre humans and humanity in the engineering of AI systems. To even begin to de-fetishise AI, we must look the human-in-the-loop in the eyes. Full article
(This article belongs to the Special Issue Advanced Studies in Human-Centred AI)
Show Figures

Figure 1

22 pages, 681 KB  
Article
Legal Decision Biases in GPT: A Comparison with Human Judgment
by Toscane F. Bessis, Andy J. Wills, Bartosz W. Wojciechowski, Lee C. White and Emmanuel M. Pothos
Behav. Sci. 2026, 16(3), 437; https://doi.org/10.3390/bs16030437 - 17 Mar 2026
Viewed by 1174
Abstract
Legal decision-making is expected to meet high standards of consistency and rationality, yet human judgments in this domain are known to be influenced by procedural factors such as evidence order and intermediate evaluations. Recent work has shown that even legal professionals, including judges, [...] Read more.
Legal decision-making is expected to meet high standards of consistency and rationality, yet human judgments in this domain are known to be influenced by procedural factors such as evidence order and intermediate evaluations. Recent work has shown that even legal professionals, including judges, are susceptible to such biases when assessing criminal cases. This raises a critical question: do large language models, which are increasingly proposed as decision-support tools in legal contexts, exhibit similar procedural biases—and if so, can these biases be mitigated? To address this question, we tested GPT-4o and GPT-5.2 using a controlled legal judgment task adapted from prior human research. The task involved simplified criminal cases in which we systematically manipulated (i) the order of incriminating and exonerating evidence and (ii) whether an intermediate guilt judgment was required before a final decision. Model responses were directly compared to human judgments from the original study. We additionally examined whether prompt engineering strategies, based on current best-practice recommendations, could reduce observed biases. GPT-4o exhibited robust order effects and a form of evaluation bias, although the latter differed in structure from the human pattern. GPT-5.2 showed similar but attenuated effects. Across both models, prompt engineering had limited and inconsistent impact, failing to reliably eliminate procedural sensitivity. These findings suggest that even advanced large language models remain vulnerable to normatively irrelevant procedural influences. More broadly, they advise caution in treating large language models as inherently rational or bias-resistant decision-support systems in high-stakes professional domains such as law. Full article
(This article belongs to the Special Issue Advanced Studies in Human-Centred AI)
Show Figures

Figure 1

17 pages, 399 KB  
Article
Beyond the Machine: An Integrative Framework of Anthropomorphism in AI
by Petru Lucian Curșeu and Ștefana Radu
Behav. Sci. 2026, 16(3), 358; https://doi.org/10.3390/bs16030358 - 3 Mar 2026
Cited by 3 | Viewed by 3255
Abstract
AI-enabled technology (AI) has a transformational role in our modern society because it is increasingly used as an interaction partner, making anthropomorphism (tendency to ascribe human features to non-human agents) a central mechanism shaping how people evaluate, accept or resist AI systems. Existing [...] Read more.
AI-enabled technology (AI) has a transformational role in our modern society because it is increasingly used as an interaction partner, making anthropomorphism (tendency to ascribe human features to non-human agents) a central mechanism shaping how people evaluate, accept or resist AI systems. Existing technology acceptance models and anthropomorphism frameworks, however, offer limited guidance on how human-like attributes of AI translate into perceptions of usefulness, perceived control, perceived opportunity or threats, particularly across different levels of AI autonomy. Building on the theory of planned behavior, the technology acceptance model and threat rigidity model, this paper develops a mid-range conceptual framework of AI anthropomorphism grounded in universal social perception dimensions of warmth and competence. We integrate fragmented research to derive three core propositions and four corollaries that specify how warmth and competence attributions shape evaluative cognitions in relation to AI. The framework further identifies AI autonomy as a boundary condition under which anthropomorphic cues may either facilitate acceptance or trigger perceptions of pseudo-empathy, cognitive superiority and identity threat. By offering a parsimonious, theoretically informed model, this paper clarifies when anthropomorphism fosters acceptance versus resistance in human–AI interaction and provides a structured agenda for future empirical research and AI design aimed at fostering synergies and resilience in human–AI ecosystems. Full article
(This article belongs to the Special Issue Advanced Studies in Human-Centred AI)
Show Figures

Figure 1

35 pages, 4355 KB  
Article
The Comparison of Human and Machine Performance in Object Recognition
by Gokcek Kul and Andy J. Wills
Behav. Sci. 2026, 16(1), 109; https://doi.org/10.3390/bs16010109 - 13 Jan 2026
Viewed by 1207
Abstract
Deep learning models have advanced rapidly, leading to claims that they now match or exceed human performance. However, such claims are often based on closed-set conditions with fixed labels, extensive supervised training, and do not considering differences between the two systems. Recent findings [...] Read more.
Deep learning models have advanced rapidly, leading to claims that they now match or exceed human performance. However, such claims are often based on closed-set conditions with fixed labels, extensive supervised training, and do not considering differences between the two systems. Recent findings also indicate that some models align more closely with human categorisation behaviour, whereas other studies argue that even highly accurate models diverge from human behaviour. Following principles from comparative psychology and imposing similar constraints on both systems, this study investigates whether these models can achieve human-level accuracy and human-like categorisation through three experiments using subsets of the ObjectNet dataset. Experiment 1 examined performance under varying presentation times and task complexities, showing that while recent models can match or exceed humans under conditions optimised for machines, they struggle to generalise to certain real-world categories without fine-tuning or task-specific zero-shot classification. Experiment 2 tested whether human performance remains stable when shifting from N-way categorisation to a free-naming task, while machine performance declines without fine-tuning; the results supported this prediction. Additional analyses separated detection from classification, showing that object isolation improved performance for both humans and machines. Experiment 3 investigated individual differences in human performance and whether models capture the qualitative ordinal relationships characterising human categorisation behaviour; only the multimodal CoCa model achieved this. These findings clarify the extent to which current models approximate human categorisation behaviour beyond mere accuracy and highlight the importance of incorporating principles from comparative psychology while considering individual differences. Full article
(This article belongs to the Special Issue Advanced Studies in Human-Centred AI)
Show Figures

Figure 1

37 pages, 5648 KB  
Article
Assessing the Relational Abilities of Large Language Models and Large Reasoning Models
by Matthias Raemaekers, Martin Finn and Jan De Houwer
Behav. Sci. 2026, 16(1), 45; https://doi.org/10.3390/bs16010045 - 25 Dec 2025
Viewed by 1181
Abstract
We assessed the relational abilities of two state-of-the-art large language models (LLMs) and two large reasoning models (LRMs) using a new battery of several thousand syllogistic problems, similar to those used in behavior-analytic tasks for relational abilities. To probe the models’ general (as [...] Read more.
We assessed the relational abilities of two state-of-the-art large language models (LLMs) and two large reasoning models (LRMs) using a new battery of several thousand syllogistic problems, similar to those used in behavior-analytic tasks for relational abilities. To probe the models’ general (as opposed to task- or domain-specific) abilities, the problems involved multiple relations (sameness, difference, comparison, hierarchy, analogy, temporal and deictic), specified between randomly selected nonwords and varied in terms of complexity (number of premises, inclusion of irrelevant premises) and format (valid or invalid conclusion prompted). We also tested transformations of stimulus function. Our results show that the models generally performed well in this new task battery. The models did show some variability across different relations and were to a limited extent affected by task variations. Model performance was, however, robust against the randomization of premise order in a replication study. Our research provides a new framework for testing a core aspect of intellectual (i.e., relational) abilities in artificial systems; we discuss the implications of this and future research directions. Full article
(This article belongs to the Special Issue Advanced Studies in Human-Centred AI)
Show Figures

Figure 1

27 pages, 4969 KB  
Article
LegalEye: Multimodal Court Deception Detection Across Multiple Languages
by Rommel Isaac A. Baldivas, Nivedha Sreenivasan, So Young Kang, Alexandra My-Linh Miller, Megan Chacko, Shreya Krishnan, Carmen Ayala, Esperanza Ayala and Dohyeong Kim
Behav. Sci. 2025, 15(12), 1707; https://doi.org/10.3390/bs15121707 - 9 Dec 2025
Cited by 1 | Viewed by 1738
Abstract
This study introduces LegalEye, a multimodal machine-learning model developed to detect deception in courtroom settings across three languages: English, Spanish, and Tagalog. The research investigates whether integrating audio, visual, and textual data can enhance deception detection accuracy and reduce bias in diverse legal [...] Read more.
This study introduces LegalEye, a multimodal machine-learning model developed to detect deception in courtroom settings across three languages: English, Spanish, and Tagalog. The research investigates whether integrating audio, visual, and textual data can enhance deception detection accuracy and reduce bias in diverse legal contexts. LegalEye uses neural networks and late fusion techniques to analyze multimodal courtroom testimony data. The dataset was carefully constructed with balanced representation across racial groups (White, Black, Hispanic, Asian) and genders, with attention to minimizing implicit bias. Performance was evaluated using accuracy and AUC across individual and combined modalities. The model achieved high deception detection rates—97% for English, 85% for Spanish, and 86% for Tagalog. Late fusion of modalities outperformed single-modality models, with visual features being most influential for English and Tagalog, while Spanish showed stronger audio and textual performance. The Tagalog audio model underperformed due to frequent code-switching. Dataset balancing helped mitigate demographic bias, though Asian representation remained limited. LegalEye shows strong potential for language-adaptive and culturally sensitive deception detection, offering a robust tool for pre-trial interviews and legal analysis. While not suited for real-time courtroom decisions, its objective insights can support legal counsel and promote fairer judicial outcomes. Future work should expand linguistic and demographic coverage. Full article
(This article belongs to the Special Issue Advanced Studies in Human-Centred AI)
Show Figures

Figure 1

19 pages, 964 KB  
Article
Human-Centred Perspectives on Artificial Intelligence in the Care of Older Adults: A Q Methodology Study of Caregivers’ Perceptions
by Seo Jung Shin, Kyoung Yeon Moon, Ji Yeong Kim, Youn-Gil Jeong and Song Yi Lee
Behav. Sci. 2025, 15(11), 1541; https://doi.org/10.3390/bs15111541 - 12 Nov 2025
Cited by 2 | Viewed by 1514
Abstract
This study used Q methodology to explore and categorise caregivers’ subjective perceptions of artificial intelligence (AI)-powered ‘virtual human’ (AVH) devices in caring for older adults. We derived 123 initial statements from literature and focus groups and narrowed them to 34 statements as the [...] Read more.
This study used Q methodology to explore and categorise caregivers’ subjective perceptions of artificial intelligence (AI)-powered ‘virtual human’ (AVH) devices in caring for older adults. We derived 123 initial statements from literature and focus groups and narrowed them to 34 statements as the final Q sample. Seventeen caregivers, nurses, and social workers completed the Q-sorting procedure. Using principal component analysis and Varimax rotation in Ken-Q, we identified three perception types: Active Acceptors, who emphasise the devices’ practical utility in patient communication; Improvement Seekers, who conditionally accept the technology while seeking greater accuracy and effectiveness; and Emotional Support Seekers, who view the device as a tool for emotional relief and psychological support. These findings suggest that technology acceptance in caregiving extends beyond functional utility. It also involves trust, affective experience, and interpersonal interaction. This study integrates multiple frameworks, including the Technology Acceptance Model (TAM), the Unified Theory of Acceptance and Use of Technology (UTAUT), Science and Technology Studies (STS), and Human–Machine Communication (HMC) theory, to provide a multifaceted understanding of caregivers’ acceptance of AI technology. The results offer valuable implications for designing user-centred AI care devices and enhanced emotional and communicative functions. Full article
(This article belongs to the Special Issue Advanced Studies in Human-Centred AI)
Show Figures

Figure 1

Review

Jump to: Research

18 pages, 3592 KB  
Review
Snake Oil or Panacea? How to Misuse AI in Scientific Inquiries of the Human Mind
by René Schlegelmilch and Lenard Dome
Behav. Sci. 2026, 16(2), 219; https://doi.org/10.3390/bs16020219 - 3 Feb 2026
Viewed by 820
Abstract
Large language models (LLMs) are increasingly used to predict human behavior from plain-text descriptions of experimental tasks that range from judging disease severity to consequential medical decisions. While these methods promise quick insights without complex psychological theories, we reveal a critical flaw: they [...] Read more.
Large language models (LLMs) are increasingly used to predict human behavior from plain-text descriptions of experimental tasks that range from judging disease severity to consequential medical decisions. While these methods promise quick insights without complex psychological theories, we reveal a critical flaw: they often latch onto accidental patterns in the data that seem predictive but collapse when faced with novel experimental conditions. Testing across multiple behavioral studies, we show these models can generate wildly inaccurate predictions, sometimes even reversing true relationships, when applied beyond their training context. Standard validation techniques miss this flaw, creating false confidence in their reliability. We introduce a simple diagnostic tool to spot these failures and urge researchers to prioritize theoretical grounding over statistical convenience. Without this, LLM-driven behavioral predictions risk being scientifically meaningless, despite impressive initial results. Full article
(This article belongs to the Special Issue Advanced Studies in Human-Centred AI)
Show Figures

Figure 1

Back to TopTop