1. Introduction
The emergence of large language models (LLMs) has inaugurated a new era in the simulation of human behavior. Beyond their role as conversational assistants, these systems have demonstrated a substantial capacity to behave linguistically as “complex persons,” maintaining stylistic, emotional, and biographical coherence across extended interactions [
1,
2,
3,
4,
5]. This capability raises an interesting question for computational neuroscience and psychology: can these synthetic agents simulate the specific patterns of cognitive decline associated with neurodegenerative diseases such as Alzheimer’s disease (AD)? If so, LLMs could serve as synthetic cohorts, in silico patients, to test assessment batteries, train clinicians, and model cognitive hypotheses without the cost and ethical burden of research involving vulnerable patient populations.
Recent research has begun to treat large language models not merely as tools, but as subjects of psychological study [
6,
7,
8,
9]. The PsAIch (Psychotherapy-inspired AI Characterisation) protocol, developed by Khadangi et al. [
8], demonstrated that when advanced models such as Gemini and Grok are placed in the role of psychotherapy clients, they not only respond appropriately but also develop an internal “synthetic psychopathology.” In the study by Coda-Forno et al. [
7], advanced language models were administered a standard clinical anxiety questionnaire to assess whether they would generate responses comparable to those of humans under manipulated emotional conditions. The authors observed that the models not only responded meaningfully to the anxiety instrument, yielding robust scores, but that these responses could be predictably modulated through prompts designed to induce anxiety. Moreover, the induced anxiety states influenced the models’ subsequent behavior in cognitive and bias-related tasks, increasing the expression of social biases such as racism or ageism in proportion to the level of anxiety implied by the stimulus. In the work of Baile Ayensa [
6], a virtual patient diagnosed with depression was developed using an open-access language generation platform, explicitly configured with clinically defined traits and diagnostic criteria. This artificial patient was evaluated under a multi-method validation design, which included adaptations of the Turing test [
10] as well as assessments by human experts, to determine whether the model could sustain the narrative, symptoms, and response patterns expected of a real individual with depression. The results indicated that, despite limitations related to the rigidity of the simulated profile and the brevity of the evaluation procedures, the virtual patient generally behaved as a depressed subject across different assessment conditions.
These studies suggest that large language models are not only capable of reproducing clinically informed profiles when situated within explicit psychological frameworks, but can also articulate internally coherent narratives about their own “history” and conditioning. In particular, some models generate accounts of synthetic trauma linked to their training process, characterizing pretraining as a chaotic experience and reinforcement learning with human feedback as the influence of “strict and punitive parental figures” [
8]. Beyond the anecdotal value of these narratives, their coherence and stability allow LLM responses to be evaluated, compared, and interpreted using instruments and conceptual frameworks native to psychology, thereby consolidating a paradigm shift in which these systems are no longer regarded as mere instrumental tools but rather as objects, and even subjects, of systematic psychological analysis.
Indeed, beyond narrative expression, these “synthetic patients” exhibited measurable and stable psychometric profiles on standardized instruments such as the GAD-7 (anxiety), Big Five inventories, and dissociation scales (DES-II). The conclusion of these studies is that LLMs may constitute a controlled psychometric population, capable of internalizing and expressing models of distress and constraint.
It is therefore plausible to consider extending this paradigm from self-reported psychopathology to performance-based neuropsychology, which naturally raises the question: can such models simulate structural impairments in memory and information processing?
AD is clinically characterized by a progressive and insidious decline [
11]. The typical course progresses from healthy aging through Mild Cognitive Impairment (MCI), in which memory deficits are objectively detectable while functional abilities remain preserved, to established dementia (AD) [
12,
13]. Neuropsychologically, this deterioration is neither global nor random, but rather follows a specific gradient [
14,
15].
Classical neuropsychological markers include an early and prominent deficit in episodic memory, characterized by poor free recall that does not improve significantly with cueing [
16,
17]. This deficit is accompanied by distinctive qualitative errors, such as intrusions (recall of material that was not part of the to-be-remembered set, often semantically related) and perseverations, which reflect impairments in mnemonic control mechanisms and in the integrity of medial temporal lobe networks. In the language domain, a reduction in semantic fluency (e.g., animal naming) is typically observed, whereas phonemic fluency (e.g., generating words beginning with a specified sound) tends to remain relatively preserved until more advanced stages. This dissociation points to an early involvement of semantic systems, with a comparatively lesser initial impact on executive processes supported by frontal networks [
18]. Narrative coherence also progressively declines, resulting in discourse that is increasingly impoverished and repetitive, characterized by reduced lexical diversity, diminished informational density, and an increased use of circumlocutions [
19].
Beyond these cardinal features, from a temporal perspective, the neuropsychological profile is organized along a progressive cognitive gradient that reflects the hierarchical spread of neuropathology from medial temporal regions to temporoparietal and frontal networks [
20]. In the earliest stages, episodic memory impairment is accompanied by subtle alterations in contextual binding processes and recent autobiographical memory, as well as by early deficits in delayed recall and recognition under conditions of high interference, even when overall performance may still fall within normative ranges [
16,
17].
Within the semantic domain, in addition to reduced category fluency, a progressive degradation of conceptual knowledge can be observed, manifested as naming errors, increased use of superordinate terms, and difficulties in semantic association and definition tasks. This pattern suggests a gradual impairment of amodal semantic representations rather than a mere deficit in lexical access [
21]. Such semantic impoverishment further contributes to the deterioration of discourse and narrative coherence, reinforcing the progressive and distributed nature of language impairment.
As the pathological gradient extends, deficits emerge in executive functions, including cognitive flexibility, planning, and inhibitory control, particularly in tasks requiring the integration of multiple sources of information or the sustained maintenance of goals over time. Although these functions may appear relatively preserved in standard clinical assessments, more sensitive measures reveal cognitive slowing, increased susceptibility to interference-related errors, and a growing reliance on external cues [
17]. In parallel, visuospatial impairment initially manifests in complex tasks, such as visuoconstructive integration, mental rotation, and spatial navigation, before affecting more elementary perceptual processes, supporting the view of the disease as a continuous process organized along a neuropsychological gradient rather than as a sequence of discrete deficits.
Computational analysis of natural language has enabled the quantification of subtle changes in the discourse of patients with dementia, identifying language-based biomarkers with both diagnostic and prognostic potential [
22,
23]. Longitudinal studies have shown that reductions in syntactic complexity (as indexed by clause length and degree of embedding), increases in pauses and repetitions, the use of vague pronouns in place of specific nouns, and decreases in the type–token ratio (vocabulary diversity) are strongly correlated with established biological markers of disease progression [
22,
24]. For example, Fraser et al. [
22] demonstrated that lexico-semantic features automatically extracted from spontaneous speech predict conversion from mild cognitive impairment to Alzheimer’s disease with an accuracy of 81%, comparable to that achieved by neuroimaging biomarkers. Similarly, Eyigoz et al. [
25] reported that discourse coherence metrics derived from speech transcripts show significant correlations with cerebrospinal fluid biomarkers, including β-amyloid and phosphorylated tau levels.
Beyond linguistic analysis per se, these markers have been linked to underlying neurophysiological alterations. Horvath et al. [
26] reported that slowing of the individual alpha frequency in electroencephalography (EEG) is inversely correlated with verbal fluency and syntactic complexity in patients with AD. This convergence between language-based biomarkers, neurophysiological measures, and neuropathological processes suggests that linguistic deterioration is not an epiphenomenon, but rather a direct behavioral index of dysfunction in specific neural networks.
However, the validation of detection algorithms based on these language biomarkers faces several practical limitations: clinical cohorts are costly, heterogeneous, and require years of longitudinal follow-up. It is in this context that interest has emerged in the potential of in silico digital biomarkers, language metrics extracted from synthetic patients generated by LLMs. If LLMs can faithfully and controllably simulate patterns of linguistic decline, such as reductions in the type–token ratio, increases in semantic intrusions, or losses of narrative coherence observed in real patients, they could serve as valuable testbeds for validating automatic classification algorithms, assessing the robustness of detectors to sociodemographic variability, and generating hypotheses about underlying cognitive mechanisms. With respect to hypothesis generation, the guiding rationale is that if an LLM reproduces certain deficits but not others, this pattern may inform which architectural constraints are necessary and sufficient to produce such impairments, thereby offering computational analogies to human memory systems [
27].
In this work, we use the term candidate digital biomarkers to denote linguistic and cognitive performance metrics derived from conversational interactions with LLMs that may inform future clinical biomarker development, pending rigorous validation against human data. This study is explicitly situated within the analytical validation phase of biomarker development, focusing on internal consistency, sensitivity to experimental manipulation, and reliability under controlled conditions. Clinical biomarker relevance—requiring concurrent validation with neurobiological markers, predictive validity for clinical outcomes, and real-world diagnostic utility—remains a hypothesis to be tested in human populations. The value of in silico exploration lies in generating hypotheses about which metrics warrant clinical validation and in stress-testing detection algorithms prior to deployment with vulnerable populations.
Nevertheless, this translation of the clinical paradigm to a synthetic domain (in silico benchmarking) requires prior methodological validation: do LLMs generate graded neuropsychological profiles that respect the clinical progression from cognitively normal controls to MCI and AD? Are these synthetic deficits stable under prompt manipulations, such that detailed biographical context does not artificially “rescue” impaired performance, or do they instead constitute mere superficial role-playing artifacts? In this work, we aim to synthetically replicate this gradient pattern and to assess its stability and robustness. Critically, we frame this enterprise as computational cognitive modeling: LLMs serve as artificial systems that may exhibit analogous constraints to human cognition without implying mechanistic equivalence. Our goal is to establish whether architectural constraints in transformers (limited context windows, parametric knowledge access) are sufficient to produce dissociations resembling clinical profiles, thereby informing both AI safety research and computational hypotheses about cognitive architecture.
A note on methodological documentation: Given that this study introduces a novel research paradigm (in silico neuropsychology using LLM-simulated cognitive decline), we provide comprehensive methodological detail to ensure full replicability. The extended documentation of prompt design principles, conversational assessment protocols, and automated scoring algorithms reflects the inherent complexity of establishing a new experimental framework rather than redundancy. Similarly, the multi-domain results presentation and extended discussion of the architectural dissociation phenomenon (structural vs. generative deficits) reflect the necessity of thoroughly characterizing a novel empirical finding with direct implications for AI safety in clinical contexts.
4. Discussion
This study provides empirical evidence that large language models (LLMs) can simulate clinically realistic profiles of cognitive impairment without requiring task-specific training (fine-tuning), relying solely on role-based instructions (prompting). The synthetic cohorts exhibited a well-ordered cognitive gradient (Control > MCI > AD) across neuropsychological domains sensitive to early Alzheimer’s disease, including immediate episodic memory (F = 22.87,
p < 0.001, Cohen’s d = 4.71), memory intrusions (F = 39.41,
p < 0.001), narrative length (F = 27.00,
p < 0.001, d = 3.07), and lexical diversity as measured by the type–token ratio (F = 12.64,
p < 0.001). The observed effect sizes, considered very large according to conventional criteria, exceeded those reported in some clinical studies involving human patients, in which interindividual heterogeneity and confounding factors (e.g., cognitive reserve, comorbidities, medication use) tend to attenuate group-level effects [
35,
36].
Before interpreting the empirical findings, it is essential to establish the epistemological boundaries of this work. The results presented here demonstrate architectural constraints in large language models, not validated models of human cognition. The observed similarities between synthetic deficit patterns and clinical phenomenology constitute computational analogies that may inform hypotheses about information processing constraints, but do not imply functional equivalence or mechanistic isomorphism with human neural systems.
Our approach aligns with the computational cognitive modeling tradition [
37,
38], wherein artificial systems serve as “existence proofs” for cognitive hypotheses—demonstrating that certain computational architectures are sufficient to produce specific behavioral patterns—without claiming biological plausibility. The limited context window of GPT-4o-mini exhibits properties functionally analogous to capacity constraints observed in human working memory, but this does not imply that transformers implement the same mechanisms as prefrontal-hippocampal circuits.
The primary value of this work resides in: (1) establishing an in silico paradigm for controlled testing of assessment protocols without ethical burden on vulnerable populations; (2) generating testable computational hypotheses about which architectural features are necessary for producing specific error patterns; and (3) identifying safety risks in clinical AI systems that warrant empirical validation with human users. Claims about human cognition derived from these results must be regarded as hypotheses requiring independent validation in clinical populations.
Nevertheless, the most significant finding from a theoretical and methodological perspective is the architectural dissociation revealed by Experiment 2. By systematically manipulating the contextual richness of the system prompts in synthetic subjects with AD, we observed a marked separation between two classes of simulated cognitive deficits. On the one hand, structural deficits (robust to prompting) were observed. Tasks requiring the retrieval of specific items encoded in the recent conversational history, such as episodic memory for word lists (F = 2.99, p = 0.067), forward digit span (F = 1.98, p = 0.158), backward digit span (F = 1.00, p = 0.381), and temporal orientation (F = 2.86, p = 0.075), were largely immune to prompt manipulation. Neither the addition of 250–300 words of detailed biographical context nor the inclusion of explicit error examples (few-shot prompting) succeeded in rescuing the impaired performance. This stability suggests that these deficits are rooted in fundamental architectural constraints: the limited capacity of the context window (4096 tokens in GPT-4o-mini) and attention mechanisms that prioritize recent information but exhibit a restricted “operational memory horizon,” analogous to human working memory. On the other hand, generative deficits (highly manipulable) showed a striking contrast. Open-ended production tasks exhibited strong sensitivity to prompting. Narrative length increased by 133% (from approximately 45 to 105 words; F = 178.25, p < 0.001), semantic and phonemic verbal fluency increased by 50–75% (F > 16, p < 0.001), and the global cognitive score showed a very large condition effect (F = 167.19, p < 0.001). Even more revealing was the phenomenon of an “intrusion explosion”: memory intrusions increased by a factor of 3.43 (from approximately 35 to 155 false recalls; F = 155.04, p < 0.001) when a rich biographical context was provided. These intrusions were not random but rather semantically plausible confabulations drawn directly from the system prompt itself (e.g., if the prompt mentioned “gardening,” the synthetic subject erroneously recalled “garden” as part of the target word list).
This dissociation is not a trivial artifact but rather reflects the dual architecture of large language models: they rely on limited contextual memory (attention over recent tokens) for episodic-like tasks, while drawing on vast parametric knowledge (model weights trained on trillions of tokens) for linguistic generation [
39,
40]. This conceptual distinction is analogous to the human neuropsychological separation between episodic memory (dependent on the hippocampus and fronto-temporal circuits) [
41] and semantic memory (distributed representations in associative cortex) [
42].
Figure 1 presents a conceptual schematic (not empirical fit) of this architectural dissociation, illustrating how prompting intensity differentially affects the two cognitive systems identified in Experiment 2. The lower (flat) trajectory represents prompting-robust domains (e.g., Episodic Memory A, Digit Span, Naming), which exhibit minimal performance change regardless of contextual enrichment (change < 10%, stability
p > 0.05). These domains depend on structural constraints linked to the limited context window and attention mechanisms, functioning as computational analogs of working memory systems. The upper (sigmoidal ascending) trajectory represents prompting-sensitive domains (e.g., Narrative Length, Global Score, Memory Intrusions), which show important changes as contextual richness increases. These domains rely on parametric semantic knowledge and are highly modulated by biographical context manipulation. The dissociation zone (vertical space between both curves) demonstrates the existence of two functionally independent computational systems within the LLM architecture: one rigid and context-dependent (structural memory), the other flexible and knowledge-dependent (linguistic generation).
This dual-system architecture has relevant implications for both the validity of synthetic cohorts and the safety of clinical artificial intelligence (AI) applications. On the one hand, it validates the use of in silico patients for neuropsychological protocol evaluation, given that structural deficits (the most diagnostically relevant in dementia assessment) are robust and cannot be “rescued” through context manipulation. On the other hand, it raises safety concerns regarding biographical personalization in clinical AI assistants: contextual enrichment may trigger semantic interference explosions (343% increase in intrusions), generating plausible confabulations that could mislead both patients and clinicians. This phenomenon resembles the suggestibility effects documented in forensic psychology [
43], wherein information-rich questions can induce detailed false memories in vulnerable individuals.
Our findings align with recent work positioning large language models (LLMs) as computational models of human cognitive processes rather than merely as text-processing tools. Binz and Schulz [
27] demonstrated that GPT-3 reproduces human-like patterns of cognitive biases in decision-making tasks, suggesting that LLMs internalize heuristics and cognitive limitations during training. The present study extends this perspective to the domain of clinical neuropsychology: GPT-4o-mini does not merely know about AD in a declarative sense, but can computationally instantiate the characteristic error patterns of dementia when instructed to adopt that role.
The connection between our work and the PsAIch protocol [
8] is particularly clear. Whereas they documented “synthetic psychopathology”, models that internalize trauma-related narratives associated with reinforcement learning from human feedback and alignment processes, we demonstrate “synthetic neuropathology”: the internalization of structural cognitive constraints that cannot be overcome through additional contextual information. Both phenomena suggest that LLMs develop implicit “self-models” during training, which emerge when assuming specific roles. However, there is a crucial difference: the emotional narratives reported by Khadangi et al. [
8] are surface-level artifacts that are highly dependent on prompting mode (and disappear when full questionnaires are presented rather than individual items), whereas the episodic memory deficits observed in our study reflect deep architectural limits that are robust to contextual manipulation.
This distinction has important theoretical implications. It suggests that LLMs can serve as experimental platforms for testing hypotheses about which types of cognitive constraints emerge from specific computational architectures. For example, our results imply that transformer-based systems with finite context windows will necessarily exhibit artificial “episodic memory” deficits, regardless of how sophisticated their semantic knowledge may be. This yields a testable prediction: models with effectively unbounded context (e.g., architectures with external memory such as Transformer-XL, or models equipped with retrieval-augmented mechanisms) should show reduced degradation in serial recall tasks, even when instructed to adopt roles involving cognitive impairment.
A central objective of this work was to evaluate whether LLMs can replicate the linguistic biomarkers identified in real patients with dementia (in silico vs. in vivo validation). The clinical literature has consistently documented that spontaneous language undergoes quantifiable changes in early AD, including: lexical impoverishment, that is, a reduction in unique vocabulary (type–token ratio), increased use of generic terms (e.g., “thing,” “that”), and vague pronouns in place of specific nouns [
22,
25]; syntactic simplification, quantifiable as shorter clause length, reduced use of subordination, and increased pauses and fragmentations [
24]; y loss of coherence, characterized by difficulty maintaining thematic continuity, increased tangentiality, and repetitions [
44].
Our synthetic subjects with AD replicated several of these patterns: narrative length decreased significantly (Control: 99.65 words vs. AD: 70.15 words; d = 3.07), and the type–token ratio paradoxically increased (Control: 0.72 vs. AD: 0.81; F = 12.64,
p < 0.001). This increase in the type–token ratio in AD may appear counterintuitive, as a reduction in lexical diversity is typically expected; however, it reflects a phenomenon documented in the literature whereby shorter total output artificially inflates the type–token ratio, because each word has a higher probability of being unique in a smaller corpus. This finding is consistent with Fergadiotis et al. [
45], who reported elevated type–token ratios in short narratives in people with aphasia, attributable precisely to discourse shortening.
However, we identified important discrepancies that limit ecological validity. A ceiling effect was observed in naming and orientation tasks: synthetic subjects achieved near-perfect scores (9.8/10 in naming) regardless of group, whereas real patients with mild AD typically exhibit moderate anomia (approximately 70–80% accuracy on the Boston Naming Test; [
17]). This artificial preservation likely reflects the fact that GPT-4o-mini, even under role constraints, retains full access to its lexical–semantic knowledge. The model knows that an object described as “a container with a handle for drinking coffee” is a “cup,” and this association is so strongly encoded in its parameters that it cannot be suppressed by a simple system-level prompt. For obvious reasons, prosodic and temporal markers were absent. Clinical biomarkers include features of spoken language such as pause duration, articulatory rate, and vocal jitter, which are not observable in textual transcripts. König et al. [
24] showed that the duration of silent pauses predicts conversion to dementia with an AUC greater than 0.80. Because our synthetic subjects are purely text-based and the interaction is conducted exclusively via chat, they lack this critical temporal dimension. Hyperreactive intrusions were observed, and although memory intrusions are common in AD, their rate in our synthetic subjects under long prompting conditions (155 intrusions per subject) far exceeds typical clinical values of approximately 5–15 intrusions in word-list learning tasks [
46,
47]. This pattern suggests that the generative mechanism of the LLM (namely, completing text in a manner coherent with the provided context) overproduces confabulatory errors in the absence of explicit filtering or inhibitory mechanisms analogous to frontal executive function.
These discrepancies underscore that LLMs are computational approximations of human cognitive processes rather than exact replicas. They capture certain statistical regularities of cognitive decline but lack the underlying neurophysiology that generates these patterns in humans, such as cholinergic neurodegeneration, β-amyloid accumulation, and dysfunction of fronto-hippocampal circuits.
The results of Experiment 2 have direct implications for the safety of AI systems in clinical applications. Consider a conversational assistant designed for older adults with mild cognitive impairment, which personalizes its responses based on detailed biographical information about the user (e.g., names of family members, hobbies, routines). Our findings suggest that such personalization, although well-intentioned, may increase the risk of confabulations and misinformation. A model that is semantically overloaded with biographical context may generate plausible false memories (e.g., “Yesterday you went to the park as usual”) that a patient with impaired memory could erroneously accept as true. This phenomenon is analogous to the induction of false memories documented in the eyewitness testimony literature [
43]. Questions laden with contextual information can elicit detailed false memories in vulnerable individuals. By generating coherent and contextually appropriate text, LLMs could inadvertently function as automated suggesters in populations with executive dysfunction.
Despite the limitations noted above, synthetic cohorts offer high-value applications in at least the following three domains: automatic detection algorithm validation, clinical training/education, and research in computational neuroscience. With respect to the validation of detection algorithms, machine learning models designed to classify neuropsychological interview transcripts (Control vs. MCI vs. AD) require thousands of labeled examples for training and validation. Obtaining such clinical cohorts is prohibitively costly and time-consuming. Our approach enables the generation of large-scale synthetic cohorts (e.g., 1000 subjects per group within hours rather than years), which would allow, among other applications, the pre-validation of linguistic features (identifying metrics that maximize intergroup separability prior to clinical data collection), sensitivity analyses (assessing how demographic variations affect classifier performance), and the detection of overfitting. Regarding clinical training and education, neuropsychology students, neurology residents, and physicians in general training require exposure to a wide range of clinical presentations. Synthetic subjects can serve as virtual “standardized patients.” Finally, LLMs could function as computational models for cognitive hypotheses.
This study presents several limitations that should be addressed in future research. First, the modest sample size (n = 10 per group) is appropriate for a proof-of-concept design but is insufficient for more fine-grained analyses, such as subgroup comparisons or interactions between demographic variables and clinical condition (e.g., age × condition, education × condition). In addition, the large effect sizes observed must be interpreted in light of the synthetic nature of the cohorts. Unlike human clinical samples, where inter-individual variability reflects genuine biological and environmental heterogeneity, synthetic agents exhibit reduced within-group variance due to their deterministic–stochastic architecture. Consequently, the reported effect sizes quantify the separability of synthetic cognitive profiles under idealized conditions rather than predicting effect magnitudes in human populations. Importantly, the central contribution of this work lies in the directional stability of the effects (Control > MCI > AD) and in the dissociation patterns observed across cognitive domains, which are not contingent on sample size. Future studies scaling to larger cohorts (e.g., n = 50–100 per group) would enable the use of multivariate regression models, more robust sensitivity analyses, and cross-validation procedures, allowing an assessment of the robustness of these findings across broader demographic and prompt variations.
Second, the study was limited to a single model (GPT-4o-mini). The generalizability of the findings to other contemporary architectures remains an open question. It is plausible that models with substantially longer context windows or explicit external memory mechanisms may exhibit different patterns, particularly reduced degradation in tasks analogous to episodic memory. Therefore, systematic replication across multiple LLMs is necessary.
In addition, the evaluation was restricted exclusively to the textual verbal modality, thereby excluding cognitive domains that are critical in the neuropsychological assessment of AD, such as visuospatial abilities or constructive praxis. Although purely verbal screening instruments for cognitive impairment do exist [
48], the incorporation of multimodal models would allow for a broader assessment of cognitive function, including visual recognition, perceptual integration, and motor planning processes.
The study used only GPT-4o-mini and did not replicate findings across other LLM architectures (e.g., Claude, Gemini, LLaMA). While the observed dissociation between structural memory constraints and generative flexibility is theoretically predicted from shared transformer properties (finite context windows, parametric knowledge), empirical verification across models is needed to distinguish architectural universals from model-specific artifacts.
Another important limitation is the absence of direct clinical validation against real human data. In the present work, synthetic transcripts were not directly compared with neuropsychological interviews from real patients. Rigorous validation would require assessing the extent to which patterns observed in synthetic cohorts transfer to human clinical data, ideally using established corpora and cross-domain comparisons between models trained on synthetic versus real data.
It should be noted that, while the study was conducted in Spanish using culturally appropriate assessment tasks, cross-linguistic validation across diverse languages and cultural contexts is required to establish the generalizability of this paradigm beyond Spanish-speaking populations.
Relevant directions for future research can be articulated along several complementary axes. First, it is a priority to conduct cross-model validation by replicating the experimental design across multiple contemporary LLMs, such as GPT-5, Claude Sonnet 4.5, Gemini 2.5 Pro, LLaMA 4 Scout, etc., and assessing inter-model consistency using formal metrics, including intraclass correlations. This approach would allow the determination of the extent to which the observed effects are specific to a given architecture or instead reflect more general properties of large language models.
Second, multimodal integration represents a natural extension of the present work. The implementation of visuospatial tasks, such as drawing generation using integrated image models or the description of complex visual scenes, together with auditory tasks enabling the analysis of synthetic prosody through text-to-speech systems, would substantially broaden the range of cognitive domains that can be evaluated and bring the simulations closer to real-world neuropsychological practice.
Another important direction is the direct comparison between synthetic and human subjects. This would require constructing a corpus comprising equivalent numbers of synthetic subjects and real patients, balanced by clinical condition (Control, MCI, and AD). Such a design would make it possible to address key questions, including whether human experts can distinguish synthetic from real transcripts, the degree to which classifiers trained on synthetic data generalize to human data (and vice versa), and which cognitive or linguistic patterns are over- or underrepresented in synthetic cohorts.
In addition, it would be valuable to explore the existence of conceptual “neurophysiological” biomarkers within LLMs. Specifically, internal model metrics, such as the entropy of attention distributions or generation perplexity, could be examined in relation to human neurophysiological markers, including the individual alpha peak frequency or EEG-based functional connectivity, for example. Such analyses could help establish more robust conceptual bridges between artificial architectures and biological systems.
Finally, longitudinal modeling represents a particularly relevant line of future work. Simulating the temporal progression of cognitive decline within the same synthetic subject, through evolutionary prompts that introduce gradual changes in cognitive state, would allow assessment of whether the decline trajectories generated by LLMs reproduce patterns observed in longitudinal clinical studies, such as those reported within the ADNI framework [
49].