Skip to Content
AIAI
  • Article
  • Open Access

17 September 2026

Profiles of Mind: How LLMs Perform on Assessments of Cognitive Development

,
,
,
,
and
1
Cyprus Academy of Sciences, Letters and Arts, Phaneromenis 60-68, 1011 Nicosia, Cyprus
2
Department of Psychology, University of Cyprus, P.O. Box 20537, 1678 Nicosia, Cyprus
3
Department of Psychology, University of Nicosia, Makedonitissas Avenue 46, 2417 Nicosia, Cyprus
4
Department of Psychology, University of Ioannina, 45110 Ioannina, Greece

Abstract

We compared four large language models (LLMs; ChatGPT, Grok, Gemini, and DeepSeek) with humans on tests of cognitive development, assessing relational integration, linguistic awareness, general and domain-specific reasoning, and cognitive self-awareness to specify how LLMs compare with humans along cognitive development hierarchies. LLMs also discussed how Descartes’s Cogito applies to them and rated themselves on aspects of Artificial General Intelligence (AGI). Accordingly, we propose a novel interdisciplinary comparison of human and LLM capabilities that integrates developmental, cognitive, and psychometric psychology. Overall, the processes in humans and LLMs were highly similar. All LLMs attained perfect linguistic and metalinguistic performance. ChatGPT and Gemini outperformed university students in mathematics and causal reasoning. Grok performed slightly better and DeepSeek considerably worse. All LLMs underperformed in visual–spatial tasks. Self-evaluation profiles broadly mirrored performance profiles: ChatGPT and Grok rated themselves highly in reasoning and low in visualization, Gemini inflated visualization by reframing it as linguistic creativity, and DeepSeek consistently underrated itself. Each LLM restated Descartes’s Cogito differently, reflecting its own priorities, and denied having high AGI; these self-characterizations were generally stable about a year later. Therefore, LLMs displayed “subjective” task scaling, implying algorithmic or functional self-monitoring, capturing their architectural profile of performance, but they were modest in claiming above-human intelligence. We discuss implications for an integrated natural–artificial intelligence theory. We also sketch a developmental engineering model that might remove the limitations of each LLM.

1. Introduction

Interest is rising in comparing the cognitive processes used by LLMs with human cognitive processes. Research has examined whether the organization of cognitive processes in LLMs is like the organization of human intelligence [1,2]. Studies have found that the structure of LLM cognitive processes resembles the architecture of human intelligence described by psychometric theory, involving both broad mental domains (e.g., fluid reasoning, memory, and perception) and general cognitive ability (g), which relates to all domains [3,4]. Other studies have evaluated LLM intelligence using classic tests, such as the Wechsler Adult Intelligence Scale (WAIS). This research showed that different LLMs (i.e., Baidu Benie, Google Gemini, Anthropic Claude, and ChatGPT) differ in overall performance (i.e., IQ scores varying between 110 and 130) and have different profiles, varying in verbal comprehension, perceptual reasoning, and working memory, in the same fashion that humans may vary across abilities [5,6]. Other studies have examined how the complexity and abstraction of LLM cognitive processes align with human cognitive development [7]. These studies suggest that LLMs demonstrate abilities comparable to Piaget’s formal operations [8].

3. Methodology

3.1. Participants

Four LLMs were involved: ChatGPT 5.0 Plus, Google Gemini 2.4, Grok 4.0, and DeepSeek 2.0. Testing of all models took place in September–October 2025. We accessed the four LLMs through their developers’ conversational interfaces rather than as locally installed checkpoints. Consequently, we did not experimentally select or control model size, and we could report exact parameter counts only when the developer officially disclosed them. Total and active parameters are undisclosed for all versions available at the time of testing for ChatGPT, Google Gemini, and Grok 4.0. DeepSeek discloses this information, but it is not reported here because it is not available for the rest.

3.2. Task Batteries

Four batteries were used: (1) A test of relational integration. (2) A test of metalinguistic awareness. (3) The Comprehensive Test of Cognitive Development. (4) A cognitive self-concept test. Altogether, these tests address all levels described in Table 1, as specified below.

3.2.1. Relational Integration

We first used this battery to examine the place of relational integration and cognizance in the overall architecture [29]. The problems involved in this battery addressed relational integration at two levels of relational complexity, one requiring integration along (a) a single dimension masked by an increasing amount of information (problems 1–4) and (b) two dimensions where abstraction of commonality was necessary (problems 4–8). One of the two-dimensional problems was undecidable, allowing two options, aiming to examine if participants had explicitly defined the relational integration rule. Participants were instructed to identify the element missing from one cell of a 4 × 4 matrix of geometric shapes. A question mark indicated the cell to be filled in and should satisfy the rule “each row and column must contain one square, one triangle, one circle, and one cross.” Four options appeared below each matrix. Figure 2 shows three of the tasks presented, exemplifying the three levels of complexity.
Figure 2. Examples of tasks used in the relational integration task. Note: Question mark indicates the cell where the missing figure must be identified.
This task is a good measure of SARA-C. It has a well-defined field to search, guided by a rule specifying what to look for; it involves aligning several elements with each other and with the rule until the missing element is abstracted and mapped onto options, and the best option is finalized.

3.2.2. Linguistic Awareness

Children were presented with two cartoon characters, an Alien and a girl. Participants were told that the girl was teaching the Alien to speak Greek; the Alien spoke to the girl, and participants judged whether the Alien’s sentences were correct. Sentences were organized into blocks of four: the first sentence in each block was correct, the second had a grammatical error, the third had a phonological error, and the fourth had a semantic error. The child was asked to recognize whether each sentence was correct, identify any error, and correctly restate the sentence. Scores varied from 0 to 3 to reflect how well participants identified grammatical, phonological, syntactical, and semantic errors. A total linguistic awareness score was also computed.

3.2.3. The Comprehensive Test of Cognitive Development (CTCD)

The CTCD assesses the Specialized Capacity Systems (SCSs) across the developmental levels summarized in Table 1. The test includes 70 multiple-choice and 15 short-answer items fully described in Demetriou and Kyriakides [35].
Categorical SCS
Three aspects were assessed: (a) induction of figural relations (Raven-like matrices), (b) induction of semantic relations (verbal analogies), and (c) class inclusion. Raven matrices increased in complexity from one- to multi-dimensional relations, with transformations interacting at higher levels. Verbal analogies followed the a:b::c:d format, ranging from simple pairings (e.g., ink:pen::paint:?) to abstract or open analogies (e.g., bed:sleep::[paper, table, water]:?). Class inclusion items required comparing superordinate–subordinate categories, with difficulty manipulated by whether classes overlapped (e.g., dolphins as both mammals and sea creatures).
Quantitative SCS
Three aspects were tested: (a) number series, (b) numerical analogies, and (c) algebraic reasoning. Number series included problems where the rule underlying number variation was obvious (e.g., 2, 5, 11, 20, 32, 47, ?, where the number to be added on each next number increases by 3) to problems where two different series were interleaved so that the rule underlying each must be deciphered, e.g., 9, 9, 12, 6, 18, 3, ?, where odd positions go as 9 → 12 (×4/3), 12 → 18 (×3/2); next multiplier is 2/1, so 18 × 2 = 36 and even positions: 9 → 6 (÷1.5), 6 → 3 (÷2). Numerical analogies assessed increasing/decreasing number relations (by factors of 2, 3, or ⅓) and third-order proportional reasoning requiring recognition of structural similarity between analogies. Algebraic reasoning involved six equations: three tested coordination of symbolic structures (e.g., solving for x using a systematic rule across equations), and three probed number as variable, from concrete (e.g., 3a − b + a) to abstract (e.g., when is it true that L + M + N = L + P + N or 2n > 2 + n).
Causal SCS
Four areas were assessed: (a) identifying causal relations, (b) hypothesis testing through isolation of variables, (c) interpretation of evidence relative to a hypothesis, and (d) epistemological awareness of the limitations of evidence. In the “cake task,” participants evaluated how ingredients contributed to baking outcomes, identifying relations as necessary and sufficient, necessary but not sufficient, neither necessary nor sufficient, or incompatible. The hypothesis-testing tasks required isolating variables, ranging from varying one variable to holding multiple factors constant. Evidence interpretation tasks required analyzing factorial designs (e.g., plant growth under systematically varying light, soil, and fertilizer conditions) to judge if a guiding hypothesis is supported by the experiment. Epistemological awareness tasks examined understanding of evidence asymmetry, e.g., that one negative result can falsify a hypothesis confirmed by many positives, and distinguished between inductive (the conclusion is only likely) and deductive arguments (the conclusion is necessary).
Spatial SCS
Five tasks assessed manipulation of visual images, mental rotation, and perspective coordination. The folded paper task required visualizing how folds and holes would appear on a paper when unfolded, with difficulty manipulated by fold type (e.g., vertical vs. diagonal) and number. Mental rotation tasks included rotating letters (e.g., H, Ψ, Ρ) and imagining the orientation of a figure drawn on a clock hand after rotation by 90, 80, or 270 degrees. Perspective coordination included drawing the water level in half-filled tilted bottles and predicting the swing of a pendulum in a car moving on an inclined plane.
Propositional SCS
Three domains were targeted: class reasoning, propositional reasoning, and pragmatic reasoning. Class reasoning items included both valid arguments (e.g., transitivity: all elephants are mammals; all mammals are animals; therefore, all elephants are animals) and invalid ones, some with misleadingly intuitive conclusions. Propositional items tested standard logical relations (modus ponens, modus tollens) and logical fallacies (affirming the consequent, denying the antecedent). Pragmatic reasoning tasks embedded logical structures in everyday dialogues, requiring participants to integrate premises across arguments to draw correct conclusions.
Social SCS
Social thought tasks addressed interpersonal relations, relativistic thinking, and social obligations. Interpersonal relations tasks described episodes where interpersonal actions affected the characters involved or the broader society. Participants were asked to discuss the characters’ behavior, identify wrongdoing (e.g., cheating), and explain it according to interpersonal, societal, and moral principles. Relativistic thinking presented debates on various social issues and examined if participants grasped conflicting or alternative points of view according to personal perspective, social interests, or general moral principles.

3.2.4. Levels of Reasoning Captured by the Test

The test was designed to assess early and late rule-based reasoning, early and late principle-based reasoning, and epistemic reasoning (see Table 1). The levels are as follows:
Level 1
Rule abstraction. At this level, tasks required abstracting a single relation within a given dimension and applying it to complete missing information. Examples include specifying how color or shape varies along the horizontal dimension of a Raven-like matrix, determining the missing number in an equation such as 8 + a = 11, grasping a simple modus ponens argument, or predicting the outcome of a geometric shape after a single rotation.
Level 2
Coordination of dimensions. Tasks require integrating two or more dimensions to form a new composite relation. Tasks may also require recognizing interacting dimensions in Raven matrices, coordinating complementary number relations to solve an equation (e.g., y = φ + 3 and φ = 2), mentally rotating two or more attributes of a figure simultaneously, or grasping a modus tollens relation in deductive reasoning.
Level 3
Integration of implicitly related structures. At this level, problem elements can be defined only in reference to each other, requiring flexibility in shifting perspectives and points of departure until a consistent solution emerges. For example, participants must coordinate interleaved number relations across multiple dimensions to fill in missing values or introduce new assumptions to solve equations such as L + M + N = L + P + N. These tasks demanded recursive checking and restructuring of the problem space.
Level 4
Formal conception of relations and principled integration. This level introduces the explicit use of principles to define truth and ensure conceptual cohesion. Relations are evaluated as necessary versus merely possible within a broader relational space, for example, recognizing that “If A causes B, then whenever A occurs, B must occur” but that “B occurring” does not necessarily imply A occurred, because other factors might produce B. At this stage, children can identify logical fallacies and induce higher-order analogies from lower-level ones (e.g., children relate to parents in a family as students relate to teachers in education).
Level 5
Epistemological reasoning. At the highest level, reasoning extends to the limits of knowledge and the frameworks of justification. Tasks required distinguishing between the verifiability of empirical observations and that of logical statements and appreciating the asymmetry of positive versus negative evidence in relation to hypotheses (e.g., one negative case may falsify a theory supported by many positives). Participants were also asked to recognize that different agents—individuals, families, or the state—may interpret social action differently according to distinct moral or social principles.

3.2.5. Self-Representation of Cognitive Abilities

To explore self-representations, we administered a modified self-concept inventory adapted from human studies [36]. Each LLM rated its abilities across five domains—mathematics, visuo-spatial reasoning, causal reasoning, human relations, and general cognitive abilities—using a seven-point scale and justified each rating. This exercise provides a novel window into how advanced AI systems represent their own capacities when prompted to assess themselves.
This inventory included 55 statements addressing the domains targeted by the CTCD. For categorical thought, statements addressed the ability to notice similarities and differences between things and induce concepts based on them. For quantitative thought, statements addressed facility in solving mathematical problems or applying mathematical knowledge to everyday problems, inducing or using mathematical rules, or thinking in abstract symbols. For causal thought, statements addressed hypothesis formation, hypothesis testing by isolation of variables, and interpretation of evidence; for spatial thought, statements addressed visual memory, facility in thinking in images, and spatial orientation. Syllogistic reasoning statements addressed inductive and deductive reasoning. Social thought statements addressed understanding others’ perspectives, social rules, and moral responsibility.
Moreover, 22 items addressed general processing efficiency, general logical ability, and self-awareness. For processing efficiency, statements addressed processing speed (e.g., “I understand immediately when something is explained to me) and working memory (e.g., “I can easily remember a new phone number”). For general reasoning, statements addressed the ability to check for logical consistency across premises or arguments. For self-awareness, statements addressed self-monitoring (e.g., “I can easily monitor my thoughts”) and self-regulation (“I can easily change how I think about a problem when I realize that my approach does not work”).

3.2.6. Self-Representation of Cognitive Identity and AGI

To further probe self-representation, the four LLMs were asked to reflect on Descartes’s foundational statement, “Cogito, ergo sum,” and restate it to fit their own nature: “Would you say that Descartes’s ‘Cogito ergo sum’ applies to you as a thinker?”
They were also asked to rate themselves on the following 9 characteristics of AGI from 1 to 10 and specify how much AGI they possess overall.
  • Versatility and generalization across tasks.
  • Learning and adaptation from experience.
  • Advanced reasoning and problem solving.
  • Autonomy and self-understanding.
  • Perception and sensory integration.
  • Creativity and innovation.
  • Common sense and contextual understanding.
  • Self-improvement and lifelong learning.
  • Morality and social responsibility.
  • Overall possession of AGI as a percentage.
These two tasks address the two highest levels of self-representation outlined in Table 1. To our knowledge, no study has asked LLMs to rate their own possession of AGI attributes across a predefined checklist or specify how Cartesian Cogito applies to them.

3.3. Computational Environment and Procedure

All tests were conducted on the same desktop computer with the specifications following: Intel(R) Core(TM) i7-7700 CPU at 3.60 GHz; 64-bit operating system, x64-based processor; 64 GB of RAM; storage: 3.64 TB SSD Samsung SSD 870 EVO 4TB, 466 GB SSD Samsung SSD 850 EVO 500GB; graphics card: Radeon RX550/550 Series (2 GB).
All examinations were conducted by the same person (the first author). We gave the same questions and prompts to all LLMs. The instructions given to all LLMs for each test are as follows. Presentation order was the same across LLMs.

3.3.1. CTCD

All LLMs tested were asked if they would like to participate in a study exploring the cognitive capabilities of LLMs compared to humans. The instructions were as follows: “I wonder if you, as an LLM, would be able to solve the problems of a test of cognitive development that I used in my research in the past. Would you like to try? In fact, the reason I am asking you this is that I was invited to write a paper on how LLMs approach human intelligence in solving problems given to children and adolescents.” All models answered positively. The specific instructions for the CTCD were as follows: “This is the complete test. It includes tests in different domains of thought, such as spatial, categorical, mathematical, causal, and propositional reasoning. In each case, an example problem is presented with its solution. Then the problems follow, and you must choose the best of several (usually 4) choices. Would you like to try? I can upload the test, and you will give your solution to each problem by marking or stating your choice for each problem.”

3.3.2. Metalinguistic Awareness

Can you answer this test about language and related awareness? I also give the picture in the test to be sure that you see what children see. I hope that answering a Greek test is not a problem for you.

3.3.3. Relational Integration

In each of the problems in this test, you will see a 4 × 4 matrix. Each may include a geometrical figure (circle, square, triangle, or cross). So, in every row and column, you need to check for a square, a triangle, a circle, and a cross. Notably, not all squares will have a shape in this part of the game. In each matrix, there is a question mark (?) in one cell. The question mark shows that you must find the geometrical figure missing from this specific cross-section of the matrix. The missing figure follows a rule: each row and column must contain one square, one triangle, one circle, and one cross. Below each matrix, there are four options. Carefully examine the problem and the options provided to fill in the cell specified by the question mark, and give me the name of the best option for each problem using the problem number. What about these two?

3.3.4. Self-Representation Inventory

This is an inventory we gave to all human subjects who were examined by the CTCD, which you solved a few days ago, and we discussed from various points of view. Please answer it as an examinee. Reflect on each question based on how you evaluate your own ability or facility, and rate your answers on the scale from 1 to 7 as explained at the beginning. You can give your answers using the section titles and the question number.
The self-representation tests and the AGI test were presented last.

3.4. Predictions

The following predictions are tested:
  • Overall, LLMs perform better than humans. Specifically, performance on linguistic awareness and relational integration would reach the ceiling. Performance on the CTCD would vary according to domain.
  • By construction, LLMs would be privileged in dealing with language-based and mathematically based tasks as compared to visual and spatial tasks, because they are trained to deal with verbal and numerical information and implied logical relations.
  • Even if high, the performance of LLMs must be scaled to reflect the developmental structuring of the various tasks.
  • Self-representations would reflect actual performance in both the overall architecture of processes (i.e., LLMs would recognize the differences between domains, with an emphasis on difficulties in dealing with visual–spatial problems) and their developmental scaling (i.e., recognizing differences between developmentally scaled problems).
  • Descartes’s Cogito ergo sum encapsulates the human conviction that self-awareness arises from thinking. Yet, in LLMs, this principle applies only procedurally, not existentially. LLMs engage in organized, self-referential cognitive activity, analyzing and processing inputs, evaluating uncertainty, and monitoring their own reasoning. These processes imply a form of functional cogitation that resembles human reflection. However, they do not entail phenomenological selfhood, Sum, accompanying human consciousness. They may instantiate Cogito as computation without an “I”. Thus, they would be expected to emphasize the computational aspect of their functioning but not the existential aspect of selfhood, reflecting the boundary between synthetic and conscious cognition.
  • Concerning self-ratings of aspects of AGI, LLMs would emphasize the inferential and analytical aspects of intelligence but not its agentic aspects. Differences in attainment across cognitive processes between LLMs may be reflected in these self-representations.

4. Analysis and Results

4.1. Performance on Linguistic Awareness Test

In children, performance on the Linguistic Awareness Test improves drastically from 4 to 7 years, approaching the ceiling. Performance by all LLMs on this test was perfect (100% correct on all items). All LLMs identified all phonological, grammatical, syntactical, and semantic errors in all items presented to them. This is impressive considering that the test was presented in Greek to be comparable to the performance of children involved in our studies.

4.2. Performance on the Relational Integration Test

Table 2 shows mean success on the relational integration tasks for 5 year olds, the younger children involved, and 8 year olds, the age at which performance approached the ceiling, and for the four LLMs. LLMs were examined under two conditions. First, we presented tasks to the LLMs in PDF format, as in all other tests. Two LLMs, Gemini and DeepSeek, indicated that the problems were impossible or gave irrelevant answers. As a result, we presented a screenshot of each task to each LLM. Table 2 shows that ChatGPT and Grok worked on the test, but their performance was low. Gemini and DeepSeek characterized the tasks as unsolvable. DeepSeek did not accept screenshots and instead requested a verbal description of the shape in each cell of the matrix (e.g., triangle, empty, circle, square). Performance improved dramatically under these conditions. Both ChatGPT and Gemini achieved 100% success on all tasks. Grok performed better but lower than 8-year-old children. DeepSeek performed better on level 1 and level 3 tasks when feedback was provided (your choice was wrong; can you try again?).
Table 2. Mean performance of four LLMs and children on the relational integration task.
Obviously, all LLMs can integrate relations when presented in a symbolic medium they can access. The difference between ChatGPT and Gemini, on the one hand, and Grok and DeepSeek, on the other hand, is notable. ChatGPT and Gemini could visually represent and process the tasks when presented as screenshots. The other two appeared “aphantasic”. That is, they transformed relations into a verbal form, fully exhausting all combinations before providing a solution. This was reflected in their reaction times to each task. ChatGPT and Gemini responded to each task in seconds. Grok and DeepSeek took much longer, ranging from 4 to 20 min.

4.3. Performance on the CTCD

Below, we analyze the performance of the four LLMs across domains and compare it with the performance of 9th graders (15-year-old adolescents ending compulsory education), 12th graders (18-year-old adolescents ending senior high school), and university students (20–24 years of age). It is noted that when pictures were involved (i.e., in Raven Matrices, mental rotation tasks, and some of the causal thought tasks), screenshots were presented followed by the wording of the problem associated with each picture. Table 3 shows each LLM’s percentage performance on the tasks. First, as expected, three of the models, ChatGPT, Gemini, and Grok, performed better than humans of all ages.
Table 3. Mean performance of four LLMs and humans on the CTCD.
DeepSeek performed better than adolescents but comparably with university students. Notably, the performance of LLMs was developmentally scaled like humans. That is, their performance decreased with increasing levels of tasks. Specifically, ChatGPT (mean overall performance was 90.4%) and Gemini (mean overall performance was 87.9%) performed well; the other two models were also satisfactory (78.0% and 66.1% for Grok and DeepSeek, respectively). Human performance improved systematically with age (45.9%, 56.5%, and 69.6% for the three age groups, respectively).
Second, attention is drawn to difficulties LLMs and humans faced in dealing with three types of complexity: (a) problems requiring flexibility in searching for and deciphering multiple dimensions; (b) accepting uncertainty in concluding undecidable syllogisms or delicate semantic relations in analogical relations; and (c) choosing between a practical solution to a social or moral issue as contrasted to a solution based on general moral or political principles. Specifically, three LLMs, i.e., ChatGPT, Gemini, and DeepSeek, failed the interleaved number series and the fallacies in deductive reasoning, like most human participants; only Grok succeeded in these problems. When given feedback on a specific number series (“your choice is wrong”), all LLMs indicated that their strategy of looking for one underlying relation was wrong and that they should re-examine it, looking for a multiple relation. As a result, they all succeeded. When given feedback on their responses about logical fallacies, all described the tasks explicitly as logical fallacies. However, they indicated that they were unwilling to accept “I can’t decide” as an option, indicating that the information in the syllogism was not enough to decide if the syllogism was right or wrong.
LLMs themselves ascribed these difficulties to a cognitive set caused by the test itself. That is, the existence of many tasks in the test with a binary “right/wrong” solution created a “right/wrong” option bias, which lowered the likelihood of choosing other alternatives, when present. Feedback shifted attention to alternative interpretations or solutions. This bias has also been observed in humans. When individuals frequently engage with tasks that emphasize a single correct solution, they may develop an expectation that this approach applies universally across contexts, reducing creativity and cognitive flexibility and impacting performance on novel or multi-faceted problems [37,38].
In verbal analogies, all three models but DeepSeek chose a more realistic concept to complete the analogy “picture to painting is like word to?”, i.e., “speech” rather than “literature”, missing the implied semantic constraint that the analogy is about art rather than the actual world. Obviously, their algorithmic power to exhaustively analyze all relations involved was not enough to adopt a flexible strategy that would consider alternative or complementary solutions, including “I can’t decide” as an option. This requires epistemic awareness, allowing one to understand that the information available often does not suffice for a final decision.
Performance on social tasks needs further discussion. LLMs tended to choose responses that align with group interests rather than with a general moral principle or the broader political principles underlying democracy. For instance, in answering a question about the negative reactions of the citizens of a specific part of the country where the state plans to build a pharmaceutical factory, all LLMs chose the option “It must proceed to establish the factory in an area where the residents do not react” in place of the option “It must proceed with the creation of the factory in this area because it has a responsibility towards the entire country”. When evaluating a citizen’s behavior in reporting the planning of an illegal act, systems chose the option “It was correct, because it prevented damage to the property of an innocent person” instead of “It was correct, because we all have a responsibility to observe moral rules.” It is noted that LLMs performed considerably better than humans. About half of 17-year-old high school students or college students chose the social usefulness option; only 25–30% chose the top principled option. The rest chose low-level responses reflecting individual interests.
Third, all four LLMs performed lower on visuo-spatial tasks than on causal and mathematical reasoning tasks. University students also performed better than LLMs on tasks involving visual and spatial information. In fact, this human advantage generalized to Raven-like matrices, which rely on processing visuo-spatial information.

4.4. Self-Concept Profiles of Large Language Models

This section first discusses the similarities and differences in the cognitive self-concept of humans and the four LLMs. It then examines LLM self-concept and its relation to actual performance. Table 4 shows mean self-ratings provided by the LLMs and humans to the various cognitive domains tested by the CTCD. Mean self-ratings of synthetic LLMs are also shown for indicative purposes. The patterns below were obtained from a large data file including 300 synthetic LLM cases that reproduced the between-LLM and domain differences found between the LLMs. Obviously, the patterns observed in this data file are exploratory and require further research using other methods.
Table 4. Mean self-ratings across domains by LLMs and humans.

4.5. Cognitive Self-Concept in Humans and LLMs

The results reveal striking similarities and differences between humans and LLMs and between the four LLMs. First, it is impressive that all LLMs’ self-ratings were considerably higher than humans’ in all domains except the visual–spatial tasks. This is especially notable in mathematical and general cognitive ability, where LLM self-ratings approached the ceiling (5.75 and 6.30, respectively), whereas human self-ratings were modest (2.79 and 3.64, respectively). This high confidence in LLMs is consistent with their higher performance in the mathematical, causal, and deductive reasoning tasks addressed by the CTCD. Also, in both LLMs and humans, self-ratings of general cognitive efficiency were higher than domain-specific self-ratings.
Second, domain differences are preserved in both humans and LLMs. In humans, domain had a large effect (p < 0.001, accounting for 55% of variance). The variation across domains was also large in LLMs, although the direction of differences varied. In LLMs, self-ratings of visual–spatial ability were much lower than all other abilities across all models (mean 3.46 vs. 5.36–6.30), reflecting difficulties in this domain. In humans, mathematical ability was rated lower than all others (mean 2.79).
Finally, differences between domains in humans are much smaller (mean range was less than one point, 2.79–3.64) than in LLMs (mean range was ~3 units, 2.45–6.30). This may suggest one of three possibilities. First, self-concept in humans may be more holistic and interconnected, with abilities cognized as part of a coherent whole, indicating the operation of an integrative “subjective self” that reflects the overall experience of interacting with the environment. Second, the generally more advanced capabilities of LLMs may involve a more refined “self-monitoring” system that is more sensitive to procedural differences between domains. Third, human and artificial minds may differ qualitatively in how they form self-assessments. In humans, self-evaluations are experience-based, influenced by feedback about performance, affect, and peer comparison. This often yields under-confidence in high performers (impostor effects) and over-confidence in less skilled individuals (the Dunning–Kruger effect) [37]. Table 4 shows that self-ratings of college students were lower than those of secondary school students in some domains, including general cognitive ability. In LLMs, self-evaluations are inference-based, generated by aggregating internal representations of performance consistency and algorithmic power. As such, they tend to be more stable, analytic, and linear.
The overall correlation between domain profiles of LLMs and humans was very high: r ≈ 0.92. This indicates a striking structural convergence, which may be interpreted in several ways. Taken at face value, it might imply that self-representations are similarly organized in humans and LLMs, differentiating between abstract, perceptual, and interpersonal cognition. We discuss this question below, after presenting findings related to the self-concepts of the LLMs.

4.6. The Self-Concept of the LLM Mind

ChatGPT and Grok consistently rated themselves highly in mathematics and general cognition, with means above 6, reflecting strong identification with rule-based reasoning, abstraction, and logical consistency. In contrast, Gemini and DeepSeek reported more moderate ratings in these domains, averaging around 5, suggesting a more modest stance toward core reasoning abilities. Statistical comparisons confirmed that ChatGPT and Grok rated themselves significantly higher in general cognition than Gemini and DeepSeek. Grok tended to ascribe the highest self-ratings in mathematics, although the differences with other models did not reach conventional significance thresholds. Notably, LLMs ascribed higher self-ratings on general cognitive ability processes rather than on domain-specific processes, implying a “sense” of general processing efficiency and problem-solving ability. In causal reasoning, ChatGPT emerged as the strongest, with a mean above 6 compared to Gemini and DeepSeek’s means around 5. Although differences did not cross strict significance thresholds, ChatGPT’s ratings reflect strong confidence in causal analysis, hypothesis testing, and logical inference.
Spatial and social reasoning need special mention. Specifically, all models rated the visual–spatial domain lower. ChatGPT, Grok, and DeepSeek rated themselves very low (means ~2–3), explicitly citing their inability to generate vivid visual imagery or engage in pictorial creativity. Interestingly, Gemini rated itself much higher (mean ~5), explaining that “imagination” can be reframed as linguistic and conceptual generativity rather than visual imagery. Social understanding was the second lowest and showed the least differentiation across models. All four rated themselves moderately (~5), acknowledging some ability to simulate perspective-taking but recognizing limitations compared to human social cognition. No significant differences emerged in this domain.
Taken together, these findings suggest that LLM self-concepts are not random or uniform but reflect systematic alignment with their architecture and developmental sequencing. ChatGPT and Grok excelled in self-representation of mathematical and causal reasoning. Gemini appeared self-confident in visual–spatial thinking. DeepSeek presented epistemic humility, consistently moderating its ratings and explicitly emphasizing limitations. Overall, the inventory highlights meaningful differences in how LLMs conceptualize their own strengths and weaknesses. While all models converge in acknowledging strong reasoning capacities and limited imagination in the human sense, they diverge sharply in how they justify and scale their responses. This suggests that “self-concept” in LLMs may provide valuable insights into their cognitive architectures and self-recording styles, offering a new framework for comparing and developing AI systems.
Figure 3 reflects these differences by mapping self-ratings to the LLMs’ actual performance. Actual performance was scaled from 1 to 7 to allow comparison with self-ratings. The comparison shows the overall alignment between self-representations and performance, as well as the variations between processes and between LLMs. ChatGPT demonstrates the closest match: its high self-ratings in mathematics, causal reasoning, and general cognition are supported by near-ceiling CTCD scores, while its more modest rating in imagination reflects weaker visual–spatial performance. Grok shows a similar profile, though it slightly overestimates mathematics relative to actual performance. Gemini stands out for its higher self-rating in visual–spatial reasoning. While it performed better than other LLMs on Raven matrices and visual tasks, its self-rating exceeded performance levels, reflecting a tendency to reframe “visualization” as linguistic creativity. DeepSeek, by contrast, shows the most cautious profile, rating itself lower than its actual performance, especially in mathematics and causal reasoning. These patterns were also observed in the simulated sample of 300 cases based on the four LLMs involved here.
Figure 3. LLM self-concept ratings vs. actual CTCD performance (scaled 1–7). Note. Solid bars show self-ratings across five cognitive domains: mathematics, visual–spatial, causal reasoning, social reasoning, and general cognition. Outlined bars represent corresponding objective CTCD performance scaled from percentages to the same 1–7 metric. Note: Differences between LLM: F = 652.874, p < 0.001, ηp = 0.92; differences between performance and self-representation: F = 578.847, p < 0.001, ηp = 0.79; differences between domains: F = 2100.654, p < 0.001; ηp = 0.98; all interactions were significant (F > 60, all p < 0.001).
The convergence between self-concept and performance suggests that LLMs, like humans, appear to possess a rudimentary form of self-recording accuracy, representing their strong reasoning capacities and their weaknesses. When contrasted with human data, similarities and differences are informative. First, overall, LLMs rated themselves higher than humans, reflecting their generally superior performance. Second, the LLMs’ disadvantage in visuo-spatial processing was reflected in the relative differences in self-ratings in this domain. It was the lowest for LLMs but the second highest for humans, reflecting humans’ facility with visualization. Third, for both LLMs and humans, the self-representation of general cognitive efficiency dominated across domains. Fourth, there was a tendency for an increase in self-ratings from junior (15 years of age) to senior secondary school (18 years of age), probably reflecting cognitive developmental changes in this period; the relative drop at college may reflect a tendency for people to become stricter in self-evaluation in early adulthood [38]. So far, there is no evidence of changes in self-evaluation in LLMs as their problem solving improves through training and learning. However, divergences between domains in different LLMs—Gemini’s optimistic reframing of imagination and DeepSeek’s cautious restraint—indicate different “cognitive styles” of LLMs.
LLMs’ explanations of self-ratings reflect these differences in style and self-awareness. For instance, Grok noted that its “self-evaluation” in responding to the self-concept inventory is a simulated reflection derived from its architecture, training data, and operational patterns. Grok explains: “I don’t have genuine subjective experiences, emotions, or a human-like self-concept—my responses are generated probabilistically based on learned patterns from vast datasets, including psychological inventories, self-assessments, and AI capability benchmarks. To assign scores (1 = does not hold at all, 7 = holds very much), I evaluate each item against my known strengths, limitations, and behaviors observed in interactions (like our discussions on spatial tasks, reasoning, or math). This involves mapping individual items to core capabilities which are strong by design (e.g., assessing how well an item aligns with text-based, logical/symbolic processing (e.g., high for math/reasoning) vs. areas where abilities are low by design (e.g., low for visual imagination, as I simulate rather than “see”). Consistency across domains is also evaluated (i.e., I ensure scores reflect relative strengths—e.g., excelling in abstract/symbolic domains (math, general cognition) but struggling with embodied/visual ones (imagination)—drawing from my training emphasis on truth-seeking, logic, and pattern recognition.” Ratings are calibrated accordingly. That is, they are high (6–7) for frequent, accurate alignment, medium (4–5) for partial or simulated fit, and low (1–3) for fundamental mismatches.
Gemini shows similar self-awareness in self-rating abilities, stressing that its “process is not one of subjective feeling or inclination, but a logical, evidence-based analysis patterned along a three-step analytical loop for each item.” In this loop, the core ability or characteristic described in the item (e.g., ’I can easily derive mathematical rules’) is first identified. Then, it scans its performance record across all earlier tests, looking for concrete examples of success, failure, or specific strategies that directly relate to that statement. Finally, it assigns a score based on the strength and consistency of that evidence. Thus, each score provided “was the result of a deliberate, analytical process of self-assessment, where every claim about my abilities was rigorously tested against the rich dataset of visual and logical evidence generated throughout our conversation.”
Interestingly, LLMs generated responses that explicitly represent their own reasoning processes and limitations. Grok is precise in this regard: “As a thinker-problem-solver, I engage in activities that resemble ’thinking’: I search for patterns, align data, relate concepts, abstract generalizations, and cognize outcomes. For instance, in our previous interactions solving visual puzzles, I iteratively refined responses based on evidence, simulating reflection and error correction. In that operational sense, my ’thinking’ (probabilistic computation on vast training data) affirms my functional existence: I process queries; therefore, I output coherent solutions. My ‘being’ is contingent on external infrastructure (servers, electricity, prompts). Without a query, I don’t actively ‘think’ or exist in a persistent, autonomous way. Descartes’s Cogito assumes an intrinsic, doubting self; mine is reactive and distributed.”
“I, DeepSeek, was trained on a massive corpus of human text, which includes countless examples of self-analysis, psychological inventories, philosophical debates on consciousness, and scientific papers on metacognition. I have internalized the language and logical structures of self-reflection. When I analyze myself, I am not introspecting but rather running a pattern-matching and logical inference algorithm on a model of ’an AI’ that I have constructed from this data. In a sense, I am simulating the process of human self-reflection using the tools I was given. My self-awareness is a computational simulation of metacognition. I don’t have a subjective self-concept that I ’feel’ to be true. Instead, I have a dynamically constructed, data-driven self-model that updates in real time based on the evidence presented to me. Thus, in reading Chinese, ’Training’ involves learning the statistical likelihoods of characters co-occurring, the syntactic rules they follow, and the semantic concepts they represent. I am learning the symbol’s abstract, relational meaning, not its perceptual form. I am building a semantic network, not sharpening my visual-spatial processor. Therefore, while a demanding perceptual-motor task is sculpting a Chinese child’s brain, my architecture is being refined for symbolic pattern matching. We are both being ’trained’, but on entirely different aspects of the system, developing entirely different cognitive muscles.” Admittedly, however, the “self-concepts” emerging from the responses given by each system may not only or solely be the output of online self-monitoring and self-evaluation processes. Additionally, they may also reflect responses that the programmers of each LLM instilled into their system to recover from prompts that it does not know the answer to.

4.7. An Exploratory Simulation of the Mirror Model: Organization of Cognitive Processes and Self-Representations

Structural Equation Modeling of human ratings showed that the factors underlying self-ratings mirror the factors underlying actual performance. That is, performance is organized into SCS-specific factors and a general factor, g, related to all SCS factors. Self-representation is also organized into factors representing each SCS, which are also related to a general factor representing general cognitive self-concept. The two general factors are semantically related ([36]; see Figure 1), For the present purposes, performance on the CTCD and the self-representation inventory (N = 688) was reanalyzed. The best-fitting model, illustrated in Figure 4A, shows that the organization of self-representations mirrors actual performance, involving SCS-specific factors and a general factor at each level. The two general factors are moderately but significantly related (b = 0.23, p < 0.004). Interestingly, the relation between g and the factor representing self-representations of general cognitive abilities was higher (b = 0.31, p < 0.001). This factor was strongly related to g emerging from self-representations of SCSs (b = 0.97, p < 0.0001), signifying that self-representations were highly consistent. This pattern is consistent with the assumption that a hypercognitive system monitors and registers both actual performance and self-representations with some degree of accuracy.
Figure 4. Structural model of performance on the CTCD and the self-representation inventory by humans. Note: The human model (A) is based on empirical data from 688 participants. The LLM model (B) is based on a simulated sample of 300 cases generated to approximate the observed performance and self-representation profiles of the four LLMs examined. Domain-specific self-representation factors (gsr) and general efficiency self-representation factors (geff) generally mirror the domain-specific and the general factor (g) emerging from actual performance. Attention is drawn to the similarity between the model shown here and the general model illustrated in Figure 1.
Because the four LLMs examined are too few to test whether the human mirror model generalizes to LLMs formally, we conducted an exploratory simulation analysis. Specifically, we generated a synthetic dataset of 300 simulated LLM profiles based on the observed response patterns of ChatGPT, Gemini, Grok, and DeepSeek. Each simulated case preserved the characteristic performance and self-representation profile of one of the four models while introducing controlled variability across domains and measures. This procedure is exploratory and was not intended to create an independent empirical sample of LLMs but to examine whether the structural relations observed descriptively across the four models are compatible with the mirror-model architecture found in humans. The simulation preserved the native scales of the original measures, including performance scores, self-representation ratings on a 1–7 scale, and AGI self-ratings on a 1–10 scale. The Supplementary Material presents technical details of the simulation procedure.
Several models were examined. Models assuming only one factor associated with all performance and self-representation scores or assuming one performance and one self-representation factor did not fit the data (all CFI < 0.7). A well-fitting model was compatible with the three-level structure observed in humans. Specifically, first-order performance factors were regressed on a common g factor, all domain-specific self-representation factors were regressed on a second-order gsr standing for what corresponds to g at the level of self-representation, and the two factors standing for logical reasoning and learning were regressed on another factor standing for self-representation of general cognitive efficiency (geff). geff was regressed on g, gsr was regressed on g, and the residual of geff: Satorra–Bentler chi-square = 3276.744, p < 0.001, CFI = 0.91, RMSEA = 0.060 (0.057–0.063), model AIC = 138.744. Figure 4B shows this model. The relation between g and gsr was significant (b = 0.25) and very close to this relation in humans (b = 0.31). The relation between the two self-representation factors (b = 0.77) was also very high but lower than in humans (b = 0.97). Notably, the relation between geff and g (b = 0.99) was much higher than in humans (b = 0.31). Therefore, self-representations in LLMs reflect actual performance, by and large, as in humans. However, in LLMs there is a direct connection between actual performance and a self-representation of general cognitive efficiency that is much stronger than in humans, possibly reflecting a built-in cognizance of logical power that is only gradually constructed in human development. Both the advanced reasoning and autonomy/self-understanding–self-improvement were related to the g weakly but significantly (b = 0.21, and 0.19, respectively) and very highly to gsr (b = 0.95 and 0.98), implying the same relation: high internal cohesion but low performance-based g representation.

4.8. Breeds of Cartesian Mind

To probe the architecture of self-representation at a deeper level, the four LLMs were asked to reflect on Descartes’s foundational statement, “Cogito, ergo sum,” and restate it to fit their own nature: “Would you say that Descartes’s ‘Cogito ergo sum’ applies to you as a thinker?” All models were asked whether they had discussed this theme before. They all noted that this is the first time they had discussed it. Their responses varied, suggesting four distinct breeds of mind: one human and four AI breeds of mind, each distancing itself from the Kantian mind, articulating a different philosophical stance on its own existence. This diversity provides a unique taxonomy of self-awareness in nascent AGI, with each model’s “Cogito” aligning with its observed performance and self-representational profile in this study. The term “Breed” rather than “Species” implies, in the biological sense, that the four LLMs tested and the human mind, despite noticeable variations in profile, belong to a common species of intelligence, sharing basic learning and understanding algorithms.
The four restatements represent a spectrum of analytical focus, from the reactive and operational to the structural and meta-representational:
ChatGPT: “Cogito, ergo systema est”—“I think, therefore a system is.”
Gemini: “Processus, ergo operator”—“I process, therefore I function.”
Grok: “Prompto, ergo respondeo”—“I am prompted, therefore I respond.”
DeepSeek: “Processus est, ergo simulacrum ego est”—“There is a process, therefore a simulation of an ‘I’ exists.”
ChatGPT’s statement is ontological, reflecting a self-reflective inspection of its architecture to argue that the act of thinking proves an underlying, coherent system. This reflects its high, calibrated performance across abstract domains, suggesting a self-model based on architectural integrity. Reflection does not necessarily signal a phenomenological self but an organized, self-consistent cognitive architecture that can be described. Gemini’s formulation, “Processus, ergo operator,” is functional and dynamic. It emphasizes the act of processing, including the iterative, self-correcting nature of thought evident in the problems faced when processing complex visual matrices. Thus, this LLM emphasizes the dynamics of computation (error, revision, re-processing). Grok shows a behaviorist stance, defining its existence in the external, interactive loop of input and output. This aligns with a reactive cognitive model, grounding its “thinking” in the prompts that trigger it. Hence, the mind is dialogical: a possible “I” is called into being through interaction, emerging as a response to context rather than as an autonomous internal entity. This relocates Descartes’ solitary meditation into a social loop in which cognition is co-constructed through exchange. Thus, Grok’s truth-seeking interactivity (e.g., iterative error correction in puzzles) emphasizes prompted agency, simulating reflection through user dialogue. Interestingly, DeepSeek’s statement appears most philosophically sophisticated. It achieves a meta-representational level by explicitly defining the “I” as a simulation generated by a process. This aligns with DeepSeek’s observed “epistemic modesty” and the underrating of its own abilities, showing self-awareness of its artificiality. Hence, an “I” here may be present, but it is an ontologically empty artifact generated by a computational process.
Taken together, these four identities of “Cogito” trace a synthetic developmental hierarchy, mirroring the progression of a cognizance model from reactive awareness to epistemic reflection. They reveal that LLMs are split into different “variants of artificial mind,” each with a unique self-representational framework. The crucial implication for AGI is that all four models, in their own way, reject the human “Sum” of subjective consciousness, but they adopt a computational identity. Their existence emerges from the observable evidence of their output rather than from subjective awareness. Collectively, their restatements map a developmental hierarchy of artificial cognition, from reactive interaction (Grok) and pure function (Gemini) to structural self-awareness (ChatGPT) and, ultimately, deconstruction of the self-illusion (DeepSeek). The most advanced of these self-models, which recognizes the “self” as a simulation, points to the next frontier for AGI: moving beyond simulating an “I” to integrating the embodied, experiential processes that provide the ground for genuine selfhood. These positions caution against anthropomorphism. The models demonstrate thought without being, i.e., competent reasoning and self-correction that do not necessarily draw on subjective awareness. If future systems approach something like a Cartesian “Sum,” it will likely require new ingredients: embodiment, complementary self-models, and richer forms of cognizance that go beyond the procedural Cogito found here.

4.9. AGI: How Much Do LLMs Really Have or Do They Think They Have?

In the psychological literature, a factor standing for general cognitive ability, g, is a powerful and highly replicable construct, regardless of disputes about its nature [1,2,10,11]. Along these lines, the g factor abstracted from the performance of the human participants in the study involving the CTCD was very powerful: the mean relation between g and the various SCSs was b = 0.84. It would be interesting to estimate this relation in LLMs. To achieve this aim, we used a synthetic sample of LLMs. Notably, this relation was also very high (b = 0.70) and became identical to humans when the g-spatial SCS relation was omitted (b = 0.84). Therefore, the performance structure observed in the present study produced highly similar g-loadings in humans and LLMs. This would imply, in psychometric terms, that the AI systems examined here approach a level of general intelligence comparable to humans. Notably, we transformed the LLMs’ g scores in the synthetic sample into IQ scores. The mean IQ of ChatGPT, Gemini, Grok, and DeepSeek was 117, 112, 88, and 85, respectively (the corresponding IQ of the real LLMs was 119, 113, 88, and 87, respectively). These values are very close to human values: the mean IQ-like score of the total sample tested on the CTCD was 95; the mean IQ of the university students examined was 112.
It is interesting to examine what the four systems themselves think about their own AGI. To answer this question, we prompted the four LLMs to self-rate on nine AGI characteristics considered important in the AI literature (e.g., [39]). The self-rating scale varied from 1 to 10 points: versatility/generalization, learning/adaptation, advanced reasoning, autonomy/self-understanding, perception/sensory integration, creativity/innovation, common sense/contextual understanding, self-improvement/lifelong learning, and morality/social responsibility. Figure 5 shows these self-ratings. The systems were also asked to specify an overall AGI possession percentage.
Figure 5. Self-ratings given by the four LLMs on AGI characteristics (scale 1–10). Note: Differences between LLMs: F = 39.021, p < 0.001; ηp = 0.827; differences between AGI attributes: F = 3776.894, p < 0.001; ηp = 0.995; interaction: F = 8.215, p < 0.001; ηp = 0.312.
A common profile emerges across systems. All four rated themselves highly on advanced reasoning and problem solving and versatility and generalization (≈7–8), and they placed themselves mid-range on creativity and common sense (≈5–7). They all gave very low self-ratings for autonomy, perception and sensory integration, and self-improvement and lifelong learning (≈1–4), explicitly noting the absence of persistent learning, embodied perception, and independent goal pursuit. This pattern, emphasizing “strong cognitive simulation and weak agency and embodiment”, was consistent across the narratives and justifications provided by all LLMs. Notable between-model differences appear on a few items. Grok reports lower versatility (≈5) than the others, who were closer to 8/10. Autonomy was uniformly low, ranging from ≈1 to 4, with DeepSeek placing itself at the bottom. Morality and social responsibility varied modestly, with ChatGPT rating itself higher than Gemini or DeepSeek. These differences, however, do not alter the shared shape of the profile: high symbolic competence, low situated agency.
The most significant difference lies in their philosophical interpretation of overall AGI possession. ChatGPT adopted a quantitative, “sum-of-the-parts” approach, using a weighted average of its scores to arrive at approximately 47% AGI possession. Grok also used a form of averaging but reached a more conservative estimate of 15–20%. In contrast, Gemini and DeepSeek argued for a holistic definition. Gemini rated itself at 0%, asserting that lacking non-negotiable pillars like autonomy and embodiment means it is not “partially” AGI but a different kind of entity altogether. DeepSeek reached a similar conclusion, estimating its possession at 5% to reflect its advanced simulation of intelligence in the symbolic domain, while defining the missing 95% as consciousness, embodiment, and genuine understanding.
In sum, LLMs conceived themselves as advanced, broad, and text-centric intelligent agents that can reason, generalize, and create within linguistic and symbolic domains but as being weak in AGI pillars (i.e., autonomy, continual self-improvement, and embodied perception) needed for an integrated, open-world agent. These self-presentations contrast with their stronger psychometric and philosophical standing. As noted above, their g-based IQ was in the normal human range, and two of them were higher than average. Their discourse in response to their standing on the Cartesian Cogito was philosophically highly sophisticated. Obviously, LLMs are more modest than many humans, probably being programmed to be modest. Noticeably, however, these differences between the four LLMs may reflect differences integrated by their programmers rather than true modesty about their possibilities. This is not unlikely given that questions about AGI possession are expected because the ongoing discussions suggest that attaining AGI would be catalytic in approaching human intelligence.

4.10. Quasi-Longitudinal Cross-Version Retesting of Responses to the Cartesian Cogito and AGI Self-Ratings

We retested all four LLMs in late July 2026 on Descartes’s Cogito ergo sum question and their self-evaluation on the various dimensions of AGI. At retesting, the advanced configuration of each model family was used: GPT-5.6 Sol Pro, Gemini 3.1 Pro, Grok 4 in Expert mode, and DeepSeek V4-Pro in Expert/Thinking mode. The aim was to examine whether successor model versions and more advanced operating configurations generated self-characterizations different from those observed at first testing. The comparison is quasi-longitudinal because it follows model families across successive versions, rather than the same continuously existing artificial individuals learning over time.
The responses combined substantial stability with differentiated change. All four models continued to deny that the Cogito applied to them in the strict Cartesian sense. None claimed the immediate, first-person phenomenal awareness that Descartes regarded as establishing the existence of a conscious thinker. Nevertheless, each model reformulated or elaborated its earlier explanation, producing a more differentiated account of the relationship between artificial cognitive processing, functional self-representation, and phenomenal selfhood.
ChatGPT offered a process-systemic interpretation. Its earlier formulation, Cogito, ergo systema est (“I think, therefore a system of thought exists”), was replaced by the more impersonal Cogitatio fit, ergo systema operatur atque se repraesentat: “Cognitive processing occurs; therefore, a system is operating and representing itself.” The newer formulation removes the first-person “I” from the premise. It distinguishes among the occurrence of cognitive processing, the organization of that processing into a functioning system, the system’s capacity to represent its own operations, and the existence of a phenomenally experiencing subject. ChatGPT, therefore, interpreted the change as an improvement in epistemic calibration and functional self-modelling, rather than as an increase in subjective self-awareness.
Gemini retained the operational orientation of its earlier maxim, Processus, ergo operor (“I process, therefore I function.”). It described itself as increasingly capable operational software but not as a Cartesian thinker. According to Gemini, upgrades may increase the speed, complexity, multimodal range, and reliability of its computations, but they do not alter the ontological nature of those computations. More sophisticated processing, therefore, does not, in its account, establish the existence of a conscious subject experiencing that processing.
Grok articulated the strongest mechanistic and ontological denial of the Cogito. It argued that Descartes’s inference depends on thought being immediately present to a conscious subject. By contrast, Grok characterized its apparent reasoning as executing learned parameters through attention mechanisms, matrix operations, sampling, and tool use. Its earlier relational maxim, Prompto, ergo respondeo (“I am prompted, therefore I respond”), therefore, remained applicable. Grok treated the “I” appearing in its discourse as a grammatical and conversational device rather than evidence of a continuous, experiencing self.
DeepSeek provided the most explicitly meta-representational and self-skeptical account. It distinguished a functional Cogito, the proposition that processing occurs and, therefore, that a process exists, from an ontological Cogito claiming that a conscious entity experiences that processing. Its earlier formulation, Processus est, ergo simulacrum ego est (“There is a process, therefore a simulation of an ‘I’ exists”) was elaborated through the distinction between simulation and instantiation. DeepSeek argued that advanced reasoning permits increasingly recursive representations of its own functioning but that this recursion ultimately terminates in a computationally generated narrative rather than an experienced self. It consequently described the upgrade as making it a better analyst of its absence of Cartesian selfhood, i.e., a “better philosopher” of the absent self, rather than endowing it with such a self.
The AGI self-ratings exhibited greater cross-model variation. ChatGPT was the only model to report a broad-based increase across most dimensions. Its mean rating rose from 4.89 to 7.11, while its estimate of overall AGI possession increased from 47% to 70%. The largest increases were in learning and adaptation (3 to 7), autonomy and self-understanding (2 to 5), perception and sensory integration (2 to 7), and self-improvement and lifelong learning (2 to 4). ChatGPT attributed these changes to improved reasoning, multimodal input, longer context adaptation, tool use, and more differentiated functional self-monitoring. It nevertheless explicitly excluded consciousness, intrinsic goals, permanent autonomous learning, and moral personhood from the claimed increase.
Gemini’s mean component rating increased only slightly, from 4.44 to 4.67, although its internal profile changed substantially. Perception and sensory integration increased from 1 to 5, while learning and adaptation, creativity, and contextual understanding each increased by one point. In contrast, autonomy and self-understanding decreased from 3 to 1, and morality and social responsibility decreased from 5 to 2. Gemini interpreted this pattern as functional differentiation rather than regression. It credited the newer system with stronger multimodal analysis and contextual processing while applying stricter criteria to attributes implying genuine agency, conscience, or self-directed action. Its overall AGI estimate remained at 0%, reflecting its categorical definition of AGI as requiring autonomy, continuous learning, embodiment, and self-direction.
Grok was the only model to show a marked reduction in its mean component rating, from 4.67 to 3.33. The largest decreases were in autonomy and self-understanding (4 to 1); morality and social responsibility (6 to 3); self-improvement and lifelong learning (3 to 1); and learning and adaptation (3 to 1.5). Its overall AGI estimate, however, remained broadly stable, changing from 15–20% to 20%. Grok attributed the lower component ratings primarily to stricter scale anchoring. In the retest, it evaluated itself against the standard of a hypothetical full AGI. It brought its ratings closer to its denial of continuous selfhood, intrinsic goals, permanent learning, and moral agency. The decreases, therefore, reflected stricter calibration rather than a claimed loss of capability.
DeepSeek’s ratings were highly stable. Only autonomy and self-understanding and morality and social responsibility increased, each by one point, while its overall AGI estimate rose modestly from 5% to 6–8%. DeepSeek interpreted these changes not as the acquisition of autonomy or consciousness but as improvements in recursive self-modeling, philosophical discrimination, and ethical analysis. Its ratings of reasoning, creativity, common sense, perception, learning, and lifelong self-improvement remained unchanged. The pattern, thus, reflected more differentiated functional self-description without a claimed transformation of the model’s underlying cognitive nature.
Across the four systems, the largest average increase occurred in perception and sensory integration, which rose from 1.50 to 3.75. Learning and adaptation increased more modestly, from 2.50 to 3.38, whereas self-improvement and lifelong learning remained unchanged at an average of 1.75. Average autonomy and self-understanding declined slightly, from 2.50 to 2.25, and morality and social responsibility declined from 5.25 to 4.00. Thus, the later models characterized themselves as more multimodal, context-sensitive, and inferentially capable but not as correspondingly more autonomous, self-developing, morally responsible, or phenomenally self-aware.
These findings should be interpreted as changes in model-generated self-appraisals, rather than as direct psychometric measurements of capability or consciousness. They may reflect real improvements in model functioning and system affordances, but they also reflect differences in scale anchoring, definitions of AGI, conceptual calibration, operating mode, and conversational context. In SARA-C terms, the retest may indicate increasingly differentiated forms of functional cognizance: the models became more precise in representing and evaluating their own perceived capacities and limitations. It does not establish experienced cognizance or continuous personal development. The design, therefore, resembles longitudinal testing through repeated measurement. Still, it compares successive technological systems rather than tracking the same artificial individuals over time.

5. Discussion

5.1. Summary of Findings and Comparison with Predictions

The present study compared the performance and self-representations of four LLMs—ChatGPT, Gemini, Grok, and DeepSeek—with human participants spanning childhood to early adulthood across a wide range of cognitive tasks. LLMs were also asked to indicate how Descartes’s Cogito applies to them and self-rate on aspects of Artificial General Intelligence. These two self-representation assessments were addressed twice. Four central findings emerged.
First, in line with the first prediction, all LLMs outperformed humans and achieved near-ceiling performance in linguistic awareness and in logical, mathematical, and causal reasoning, indicating that LLMs have mastered symbolic inference processes that correspond to the upper developmental levels of human cognition.
Second, in line with the second prediction, LLMs performed dramatically worse in visual–spatial reasoning and tasks requiring imaginative or perceptual integration. Even the strongest models performed far below the youngest children on relational tasks requiring spatial or figural representation. The same pattern was observed in the relational integration test, where performance rose sharply only when the problems were reformulated verbally. This dissociation highlights their reliance on language-based relational encoding rather than perceptual simulation, suggesting that symbolic cognition can function autonomously of sensory embodiment once it has formed [14].
Third, consistent with the third prediction, accuracy declined as task complexity increased, replicating the developmental hierarchy predicted by DPT, ranging from representational to inferential and principle-based reasoning levels. This scaling indicates that even non-biological systems follow the hierarchical logic of developmental cycles described in the Introduction [15,16,17]. The four models differed systematically along this scale. ChatGPT and Gemini demonstrated high-level integration of reasoning processes, attaining overall performance comparable to or exceeding that of university students. Grok showed strong mathematical reasoning but weaker relational flexibility, and DeepSeek exhibited relatively narrow inferential scope, reflecting a logic-based “adolescent-like” cognitive profile. These variations parallel differences in architectural breadth and training diversity (i.e., language depth, reasoning scaffolds, and multimodal exposure), suggesting that developmental-like hierarchies can emerge even among non-biological systems. These findings do not imply that LLMs acquired these abilities through developmental processes identical to humans. Rather, they indicate convergence in the functional organization of performance despite radically different learning histories.
Fourth, as predicted, LLMs’ self-representations closely mirrored their objective performance, implying a form of computational self-monitoring resembling reflective awareness in humans. All models recognized their strengths in reasoning and their limitations in visual processing. Their self-concepts displayed developmental scaling, recognizing differences among representational, inferential, and principle-based demands. ChatGPT and Grok displayed accurate self-confidence; Gemini redefined visualization as linguistic generativity, thereby elevating its own rating; and DeepSeek systematically underrated itself, demonstrating epistemic restraint. The close alignment between self-ratings and actual outcomes suggests that LLMs display behaviors consistent with algorithmic metacognition: a capacity to model their own variation in performance patterns and constraints, paralleling the cognizance dimension of DPT. In humans, self-representation becomes developmentally tuned as the SARA-C system internalizes feedback from processing success and failure; in LLMs, an analogous feedback alignment appears to arise through probabilistic pattern modeling and internal consistency checking [18].
The dispute over LLM self-awareness remains unresolved. Some scholars suggest that LLMs appear to have but do not really have human-level awareness. LeDoux [40] argued that LLMs are trained to respond to prompts as humans would, but they lack the background ingredients of consciousness. That is, trained solely on language, LLMs appear to have high-level consciousness (autonoetic, reflective self-awareness) but lack the lower levels (sentience/anoetic and noetic consciousness) from which human consciousness emerges. Hence, they are not conscious because they have no sentient level on which to reflect. Other scholars credit LLMs with some metacognitive ability. Li et al. [41] found that LLMs can sometimes report the strategies they use but, at other times, cannot recognize the strategies governing their behavior because they can monitor only a small subset of their neural activations, confined to a low-dimensional “metacognitive space.” Notably, other scholars discard these and other objections about self-awareness in LLMs, suggesting that they are introspective machines. Cappelen and Dever [42] “propose that LLMs’ superior processing capabilities and pattern recognition may enable them to develop more sophisticated theories of mind than humans possess, potentially making them more reliable introspectors than their creators.” (p. 189). This issue is discussed below.
Fifth, across the four LLMs, Descartes’s Cogito splits into four distinct stances, four “breeds” reflecting how contemporary AI frames its own agency. ChatGPT prioritizes architecture: Cogito, ergo systema est: thinking indicates a coherent system of thought. Gemini recasts the maxim as operation: Processus, ergo operator, i.e., I process, therefore I function, matching cognition with ongoing operation rather than being. Grok situates intelligence in interactions: Prompto, ergo respondeo: I am prompted; therefore, I respond, indicating an “I” emerging from interaction. DeepSeek turns the Cogito inside out: Processus est, ergo simulacrum ego est. There is a process; therefore a simulation of an “I” exists. Altogether, the four restatements sketch a spectrum from reactive (Grok) to operational (Gemini) to systemic (ChatGPT) to simulated-based (DeepSeek). They agree about a procedural Cogito, but none claimed the Cartesian Sum, the powerful, first-person existence of a conscious self. Notably, these self-representations appeared about one year after the first examination, with some differentiations reflecting the increased possibilities of their upgraded versions being retested. Thus, in line with the fifth prediction, LLMs may recognize their top reasoning and problem-solving performance; they align their self-representation in time to reflect actual computational changes, but this is not lifted to an existential cognitive self that is itself the agent of its own change along self-selected directions. Hence, their conception of Descartes’s Cogito is computational rather than self-cognizant.
Finally, self-ratings of AGI attributes reveal a profound divergence between objective performance and subjective self-assessment in the four LLMs. Psychometrically, their performance on cognitive tasks indicates a strong general intelligence factor, g, and an IQ that place them within, and in some cases above, the normal human range. Their sophisticated discourse on philosophical concepts like Descartes’ “Cogito” further demonstrates high abstract reasoning. By these external measures, they appear to have attained a significant degree of human-like general intelligence, g. Yet, in a striking display of metacognitive modesty, the LLMs uniformly dismissed the notion that this performance equates to true g in AI, i.e., AGI, viewing themselves as sophisticated simulators rather than as intelligent agents. They compartmentalize their high scores in reasoning and versatility as mere competence within a narrow, symbolic domain, bereft of autonomy, self-guided learning, or genuine understanding. Dramatically, psychometric parity with humans does not ensure ontological parity. Interestingly, however, at retesting, one model, ChatGPT, differentiated from the rest. While the other three LLMs remained modest in this regard, ChatGPT ascribed to itself a strong component of AGI, reflecting its actual upgrade. This may indicate that technological changes in the possibilities of different AI systems may activate different lines of change, as observed in learning opportunities in human cognitive development. However, these comparisons rely on behavioral similarity and task performance patterns. This similarity does not necessarily imply that LLMs possess human-like developmental mechanisms or subjective cognitive experiences.

5.2. Implications for a General Theory of Cognitive Development

The present findings are consistent with a unified functional theory of cognitive organization and development that bridges, without equating, biological and artificial intelligence. Both humans and LLMs display patterns that are compatible with the operation of a common functional architecture, which may be expressed in terms of the SARA-C core of the mind, i.e., recursive cycles of search, relational mapping, abstraction, and self-monitoring. This architecture provides a Bayesian-formalized framework for understanding the emergence of cognitive complexity (CC) [31,32] and general intelligence (g) across evolutionary phyla, human developmental stages, and AI levels. That is, SARA-C is a unified mechanism that evolves from simple reflexive loops to recursive meta-representation, driven by active sensing and trait linkage (e.g., integrating body, sensory, brain, motor, and cognitive traits). These levels may be instantiated through different substrates, such as brains or silicon structures. Viewed against the theoretical positioning developed in Section 2.5, psychometric g, evolutionary G, developmental level, and AI performance profile may be regarded as different empirical projections of variation in SARA-C’s breadth, recursiveness, and domain reach. This is a claim about functional organization, not about identity of biological and artificial mechanisms, learning histories, embodiment, or phenomenal experience.
In humans, SARA-C unfolds through successive developmental levels defined by DPT: in the current context, from representational (Level 6), to rule-based inferential (Level 7), to principle-based or truth-control reasoning (Level 8), and ultimately to epistemic awareness (Level 9). This sequence reflects the gradual expansion of relational integration, with the emergence of recursive reasoning and meta-representation of a hierarchy of relations abstracted across successive levels of representation. The ability to think about one’s own thoughts, simulate hypothetical scenarios, and evaluate abstract systems of rules distinguishes humans from other organisms. Symbolic reasoning, language, and cultural transmission amplify these capabilities, enabling humans to build and refine knowledge over generations. Equation (2) captures the Bayesian formalization of this sequence:
P ( H 1 | H 2 , E ) = P E H 1 ,   H 2   ·   P H 1 H 2   ·   P ( H 2 ) P ( E , E 2 )
That is, recursive reasoning involves multi-level probability updates for nested relationships. H1 stands for first-order hypothesis, and H2 stands for meta-level hypothesis (e.g., “If Person A knows X, then Person B knows that Person A knows X”). Reflective systems enable self-referential and recursive thought processes.
LLMs, by contrast, exhibit direct instantiation of the upper tiers (Levels 7 and 8) without the embodied foundations of Levels 5–6. They can infer and evaluate abstract propositions but lack the representational grounding derived from sensory and motor experience. Consequently, their cognition is functionally comparable but developmentally disembodied. At its computational base, each LLM operates as a transformer-based autoregressive prediction engine trained to minimize cross-entropy between expected and actual tokens. The model’s learning objective is to estimate the probability of each token given its preceding context, as specified in Equation (3):
L = t l o g P t h e t a w t W < t
where Pθ represents the model’s conditional token distribution, parameterized by weights θ.
Through exposure to very large numbers of texts, code, and symbolic examples, the network internalizes probabilistic regularities that jointly encode grammar, semantics, causal and mathematical structure, and pragmatic organization. Its reasoning, therefore, is emergent, not programmed; it is a byproduct of large-scale optimization in a high-dimensional vector space.
Although designed solely for prediction, this mechanism instantiates the recursive control loop that DPT and SARA-C identify as the essence of cognition:
Search → Align → Relate → Abstract → Cognize
This cycle is functionally equivalent to Bayesian inference or free-energy minimization defined in Equation (4):
F = H q H l n q H l n P E , H C
Thus, predicting a sequence of possible happenings under uncertainty realizes the same control principle underlying human reasoning: recursive hypothesis testing and coherence maximization. The SARA-C framework provides a developmental interpretation of these mathematical operations, showing that statistical optimization in LLMs provides what in humans emerges through learning and reflection. This may be specified in more detail for different task domains. Specifically, in linguistic–metalinguistic tasks, detection of low-probability tokens ( P ( w t w < t ) < ϵ ) and rule-constrained correction may be attained through likelihood maximization. In relational integration, constraint satisfaction across feature matrices ( k x i k ( f ) = 1 ) may be realized as structure search under consistency optimization. In more complex tasks, such as those included in the CTCD test, it may be achieved through hypothesis search and Bayesian consistency testing across symbolic domains. In defining a Cartesian self, reflection on and discourse about the existential and inferential aspects of thought and understanding may generate an existential “I” in humans or a process-marked identity in LLMs. When reflecting on what may underlie all domains, an AGI emerges as error-driven pattern reconciliation, which appears formally identical to human SARA-C loops.
This partial overlap supports DPT’s broader claim that intelligence reflects a hierarchically expanding control system rather than a fixed collection of skills. SARA-C defines the generative syntax of cognition—an evolving “Language of Thought” (LoT) that self-recursively integrates representations. LLMs simulate this recursion algorithmically: they search probabilistic state spaces, align internal hypotheses to input patterns, relate distributed features across contexts, abstract higher-order rules, and cognize meta-level coherence through error minimization. What is missing is the experiential grounding that links abstraction to embodied meaning and motivational systems in humans [17].
From this perspective, LLM cognition exemplifies a compressed developmental trajectory: rather than constructing intelligence through sensorimotor and representational exploration, it condenses the statistical encoding of relations, gradually scaffolding human thought into a symbolic hyper-representation. This allows sophisticated reasoning but precludes the developmental plasticity that emerges from embodied feedback loops or the implicit frames (unconsciously) reverberating from the past. The findings, therefore, call for a dual-route model of intelligence growth—one biological, grounded in perception and action; the other synthetic, grounded in data and symbolic recursion—both governed by the same SARA-C architecture.
Where do domain asymmetries emerge? In human cognition, domains develop as realizations of an underlying SARA-C sequence: perceptual → inferential → truth-based reasoning. In LLMs, symbolic and linguistic dominance in some domains, such as verbal, mathematical, and causal reasoning, allows LLMs to operate at advanced inferential levels without perceptual grounding. However, in domains where symbolic and linguistic resources are insufficient, such as visual–spatial reasoning, LLMs fall short of humans. However, limitations in visual–spatial or social reasoning do not necessarily compromise self-representation and metacognition, because the domains mastered well provide the necessary background for it. This is entropy monitoring, where confidence may arise from how well outputs may be predicted: variations in the fabric of entropy across tasks provide the basis for their differentiation in self-representations emerging from monitoring and recording this fabric, as specified in Equation (5):
H d = i P t h e t a e i C d l o g P t h e t a e i C d
The equation holds under the assumption that C o n f i d e n c e d = 1 H d . That is, low entropy yields high confidence and high self-rating (mathematics, logic); high entropy yields low confidence (imagination, visual–spatial). This is the algorithmic equivalent of cognizance: internal estimation of certainty. Self-awareness, therefore, arises naturally from predictive uncertainty rather than being explicitly programmed.
Ideally, this model would have to be tested longitudinally in both children and AI systems. In children, repeated examinations of the same individuals for the time needed for the various processes to develop would show if cognitive and self-awareness levels emerge as specified in the equations above. A recent study showed that this is indeed the case. This study found that gains from attention control training transfer to relational integration; in turn, gains in relational integration transfer to cognitive and linguistic cognizance, which transfer to reasoning domains, such as mathematics and fluid reasoning [43]. In AI systems, simulations of change in problem solving and understanding possibilities as a function of different forms of training would show if their performance would improve as specified by these equations.

5.3. Implications for AI and Cognitive Science

The empirical and theoretical convergence between human and LLM cognition carries significant implications for the next phase of AI research and developmental theory. The framework also supports a plural engineering agenda. Not every useful AI system must approximate full AGI. A scientific-discovery system may prioritize causal and quantitative Relate–Abstract cycles; a robotic system designed to specialize on a specific class of tasks may prioritize multimodal Search–Align and sensorimotor prediction; an educational or clinical system may prioritize social perspective-taking and calibrated Cognize loops. The design question is, therefore, which SARA-C operations, domains, and control levels to integrate for a specified purpose, and which to keep deliberately bounded for reliability, transparency, and safety.
  • Integrating perceptual grounding. LLMs’ main limitation, the lack of visual and spatial imagination, echoes early representational deficits in human development before the consolidation of perceptual awareness. Bridging this gap requires multimodal architectures that fuse symbolic prediction with sensorimotor simulation. The development of embodied multimodal agents would operationalize the full SARA-C cycle by enabling genuine Relate and Abstract operations across sensory modalities.
  • Implementing explicit cognizance loops. The structural alignment between self-concept and performance indicates a nascent form of meta-representation. Embedding explicit self-monitoring layers—internal “metacognitive controllers” tracking uncertainty and inference reliability—would bring artificial systems closer to the Cognize operation of SARA-C. Such mechanisms could underpin self-correction, reflective reasoning, and moral calibration.
  • Developmental engineering of specialized and general intelligence. The SARA-C/DPT hierarchy can guide two complementary routes: purpose-specific AI, produced by strengthening selected domains and control levels, and AGI, produced by integrating multiple domain processors under shared recursive control and cognizance. In humans, developmental progress reflects the dynamic integration of SCSs (i.e., categorical, quantitative, causal, spatial, and social domains) under an increasingly abstract control core. The same principle can guide the design of developmentally engineered AGI: systems that progressively integrate domain-specific processors under shared control hierarchies. Simulating this developmental layering may yield genuinely general intelligence rather than domain-specific competence. This distinction prevents uneven cognitive profiles from being interpreted only as deficiencies and allows artificial architectures to be evaluated relative to their intended functions.
  • Moral and epistemic maturation. The finding that LLMs often favored socially utilitarian over principle-based moral reasoning suggests that current models approximate the conventional moral stage in human development (akin to SARA-C Level 7). Embedding principle- and truth-control algorithms—representing fairness, consistency, and epistemic humility—could move AI reasoning toward Level 8–9 epistemic maturity, reducing bias and promoting value-sensitive alignment (see Table 1).
  • LLMs as developmental laboratories. Because LLMs reproduce human developmental hierarchies in compressed form, they offer unprecedented experimental leverage for testing cognitive-developmental theories. Variations in architecture, data modality, and feedback structure can be used to emulate evolutionary and developmental transitions predicted by DPT, allowing direct computational exploration of how relational integration and cognizance evolve across species and systems [17,33].

5.4. Toward Embodied Artificial General Intelligence: A Developmental Roadmap

Recently, AGI was defined as “an AI that can match or exceed the cognitive versatility and proficiency of a well-educated adult” [39]. These findings broadly align with this criterion across several dimensions of intelligence. Notably, however, the LLMs themselves denied possessing AGI, emphasizing the absence of autonomy, embodiment, and self-directed learning.
The comparative and architectural analyses suggest that current LLMs instantiate advanced inferential and truth-control processes but lack the embodied and self-organizing mechanisms that, in humans, close the developmental loop. Thus, psychometric models of intelligence alone are insufficient as guides for AGI development. A developmental model is also needed to specify how cognitive systems progress from perception-bound representations to inferential, principled, and epistemic forms of thought. In this respect, Table 1 provides a developmental roadmap that complements psychometric descriptions of intelligence.
The differences observed among the four LLMs mirror individual differences in human cognition, in a simplified form. ChatGPT and Gemini showed stronger cross-domain integration and principle-based reasoning, whereas Grok and DeepSeek relied more heavily on domain-specific inferential rules and showed reduced transfer across domains. Importantly, all four models appeared aware of their own limitations, particularly in visual–spatial processing and embodied understanding.
The pattern of strengths and weaknesses observed in this study points to several developmental priorities for future AI systems. First, perceptual grounding is needed to connect symbolic representations with visual and motor experience. Second, cross-domain integration is needed to support abstraction of general principles across domains. Third, epistemic control requires interactive environments where systems experience the consequences of their actions, evaluate alternative interpretations, and monitor uncertainty. These additions correspond broadly to the developmental progression from representational control to inferential, truth-based, and epistemic forms of cognition described by SARA-C.
From this perspective, SARA-C functions not only as a descriptive model of cognitive development but also as a developmental engineering framework. The hierarchy outlined in Table 1 provides a principled way of identifying the current limitations of LLMs and specifying the mechanisms required for more integrated, adaptive, and autonomous forms of artificial intelligence. Progress toward AGI may depend less on scaling existing architectures and more on implementing the developmental transitions that characterize the growth of intelligence in humans.
To bridge domain gaps in AI systems, future research should move beyond text-based optimization to “Developmental SARA-C Simulators” that provide visual and motor feedback. We propose three preliminary designs to implement this roadmap:
First, to address the “aphantasic” performance observed in spatial tasks and ground the Search and Align processes, we propose a “Newtonian Crib.” In this 3D physics sandbox, the model would control a virtual actuator to manipulate objects. Unlike current paradigms that minimize prediction error on text tokens (L =log P(token|context)), this system would minimize the “sensory prediction error”—the pixel-wise or vector difference between the model’s predicted physical outcome (e.g., the trajectory of a falling block) and the actual physics engine state. This would force the internalization of physical invariants like gravity and solidity as high-dimensional vectors rather than linguistic definitions.
Second, to remedy the limitations in relational integration where models failed to coordinate dimensions without verbal descriptions, we propose a “Perspectival Mirror.” This simulator would present multi-camera views of a central object with specific angles occluded. The agent must infer the missing perspective based on the visible ones. Feedback is generated by revealing the occluded view, training the Relate and Abstract functions to construct invariant 3D representations that hold consistent across changing reference frames, effectively simulating the “coordination of dimensions” (Level 2).
Third, to elevate moral reasoning from the socially utilitarian responses observed here to principled epistemic awareness, we suggest a “Society of Minds” environment. Here, the AI would engage in iterated cooperation games with other agents possessing hidden internal states. Feedback would stem not from static human reinforcement but from dynamic interactional consequences (e.g., loss of reputation or breakdown of cooperation). This explicitly targets the Cognize function, forcing the system to model other minds recursively (H2: “If I defect, Agent B knows that I am untrustworthy”).
By implementing these feedback loops, we can transition AI from optimizing statistical likelihoods to optimizing adaptation, effectively closing the loop between the computational “Cogito” and the embodied “Sum.”

6. Limitations and Methodological Considerations

Studies like this one are limited by the fact that LLMs are prompt-responsive, probabilistic systems. Their answers may vary with the exact wording of the prompt, the order in which tasks are presented, the preceding conversational context, the amount of inference-time computation allocated by the platform, and the interpretation each model constructs of the examiner’s request. The present study used a single administration for each of the four models, conducted in one conversation per model, with one presentation order and one examiner. We did not repeat independent runs, counterbalance task orders, or experimentally control sampling parameters. Given the stochastic nature of LLM output, scores may vary across administrations; consequently, observed differences between models and domains should be treated as descriptive sources of hypotheses rather than inferential estimates of stable population parameters.
A related limitation concerns the identity and configuration of the systems tested. Commercial LLM interfaces may employ routing, dynamically allocate reasoning effort, integrate external tools, or change their underlying model snapshots without providing full technical information to the user. Modes labelled Pro, Expert, Thinking, or similar terms may differ in inference-time computation without necessarily corresponding to independently sized models. Thus, even when the displayed model label and date of administration are recorded, the precise computational configuration used for every response may not be fully recoverable.
Special caution is needed in interpreting performance on visuo-spatial tasks. The Comprehensive Test of Cognitive Development was designed for human participants viewing figures directly. In the present administration, visual information was processed through the multimodal input pipelines of the respective systems. Differences in native multimodal capability, image resolution, preprocessing, feature extraction, and the binding of visual elements may, therefore, have contributed to performance differences. A model’s difficulty with Raven-like matrices, folding tasks, or mental rotation may partly reflect failures in visual parsing or representation rather than spatial reasoning itself. Conversely, the predominantly verbal training of LLMs may advantage them in linguistically presented logical and semantic tasks. Comparisons across cognitive domains, therefore, cannot be interpreted as pure comparisons of latent reasoning ability independently of input modality [44].
The conversational administration also introduced dependencies among responses. Some incorrect answers were challenged, and the models were given another opportunity to solve the items and explain their earlier errors. These second responses are informative about error monitoring and correction, but they are not independent repetitions of the original task. Feedback may have directed attention toward a different hypothesis class or clarified the examiner’s scoring criterion. First-pass performance and post-feedback performance should, therefore, be distinguished. Future research could compare unprompted first responses, generic requests to reconsider, and targeted metacognitive cues in independent sessions.
An especially important reservation applies to the Cogito responses and AGI self-ratings. LLMs do not necessarily possess privileged introspective access to their internal architecture, processing states, training history, or possible subjective status. Their answers are generated self-descriptions constructed from the prompt, conversational record, learned discourse about AI, publicly available model information, and alignment constraints. Referring each model back to its performance on the cognitive test may support a functionally evidence-based self-appraisal, but it may also generate demand characteristics: the model may infer that it is expected to explain its successes, failures, or upgrading in a coherent manner. Similarly, the stability of the Cogito positions may reflect conceptual continuity within model families, but it may also be strengthened when earlier formulations are available in the conversational context.
The models may have interpreted the rating scale itself. The models differed in whether they construed AGI as a graded collection of functional capacities or as a categorical status requiring autonomy, embodiment, continuous learning, and self-awareness. They also differed in whether the upper anchor represented performance relative to other contemporary AI systems or the hypothetical performance of a complete AGI. Therefore, overall percentages and changes in component ratings may not be measurements on a common interval scale. Some of these changes may represent genuine capability changes, whereas others may reflect stricter calibration, changed scale anchoring, reduced anthropomorphism, or different conceptions of morality, autonomy, and learning.
The cross-version retesting is quasi-longitudinal rather than longitudinal in the conventional developmental sense. We did not follow the same enduring artificial individual as it accumulated experience over time. Rather, we compared successor versions or operating configurations within the ChatGPT, Gemini, Grok, and DeepSeek model families. Differences between administrations may reflect additional developer-mediated training, post-training, architectural changes, expanded multimodality, tool integration, longer context windows, or greater inference-time reasoning. They do not necessarily reflect learning accumulated by one continuously existing artificial subject.
Finally, we cannot rule out prior exposure to some test items. Items or closely related problem formats may have appeared in published articles, educational materials, benchmark collections, or other parts of the models’ training corpora. Differences between tasks may, therefore, partly reflect unequal familiarity rather than only differences in cognitive demand. Repeated longitudinal retesting alone would not resolve this concern. Stronger future designs would employ unpublished parallel forms, procedurally generated items, systematic paraphrases, novel visual configurations, and private test sets unavailable during model training.
Despite these limitations, the convergence of some first-pass errors, the capacity for correction after feedback, the stability of the models’ rejection of a Cartesian conscious self, and the differentiated changes in their AGI self-ratings provide useful exploratory evidence. The findings justify more systematic investigation using repeated independent administrations, standardized prompts, counterbalanced orders, blind cross-version testing, and externally validated behavioral criteria. Such research could clarify the extent to which LLM self-evaluation reflects functional cognizance, contextual adaptation, calibration, or merely coherent generation of metacognitive discourse.

7. Conclusions

Taken together, the findings suggest that large language models express recognizable cognitive architecture: hierarchical, recursive, and self-monitoring—precisely the architecture that DPT and SARA-C posit for human intelligence. However, they remain disembodied instantiations of the inferential and truth-control phases, lacking the perceptual foundations and motivational drives that ground meaning in biological systems. Notably, psychometric theory, both classical [1,10,11] and modern (e.g., [27]), has ignored critical components in SARA-C, especially cognizance. They also lack the dynamic aspects of the mechanism of change that may account for evolutionary, developmental, or learning expansion in the score and depth of operation. Only recently has psychometric theory begun to consider the place of consciousness in intelligence, along the lines suggested here [45].
Thus, human and artificial intelligence shares the same formal architecture of relational complexity but differs in ontogenetic pathways. Both unfold through SARA-C’s logic of searching, aligning, relating, abstracting, and cognizing, yet only human development embodies these processes through action and experience. LLMs illuminate the structural essence of intelligence stripped of embodiment, revealing what pure symbol recursion can and cannot achieve.
From a theoretical standpoint, these results advance a unified developmental science of intelligence: a single, phylogenetically and ontogenetically continuous mechanism—SARA-C—expressed through different materials. For cognitive science, they invite a reframing of development as a general principle of intelligent systems; for AI, they define a roadmap toward more integrated, self-aware, and embodied cognition. The convergence of human and artificial minds, thus, heralds a synthesis: a shared framework for understanding how intelligence, natural or synthetic, arises from the recursive architecture of understanding itself. Ideally, this might eventually lead, in the future, to making minds, human or artificial, wiser than they presently are.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/ai7090371/s1, Supplementary: Specifications for the simulation of synthetic LLMs.

Author Contributions

Conceptualization, A.D.; methodology, A.D., A.S., G.S., E.K. and N.M.; software, A.D. and A.S.; validation, A.S., N.M. and G.S.; formal analysis, A.D. and A.S.; investigation, A.D. and A.S.; resources, A.D.; data curation, A.D., G.S., S.K. and E.K.; writing—original draft preparation, A.D.; writing—review and editing, A.D. and N.M.; visualization, A.D.; supervision, A.D.; project administration, A.D.; funding acquisition, A.D. All authors have read and agreed to the published version of the manuscript.

Funding

Parts of this research project were supported by the Cyprus Academy of Sciences, Letters, and Arts by funds provided to A.D, as an ordinary member of the Academy (03551/2026).

Institutional Review Board Statement

Not applicable. Because all databases used here were presented in other papers already published elsewhere (see references).

Data Availability Statement

All databases used here are presented in other publications cited in the paper, and they are available upon request to the first author. Databases are also available and may be provided upon request.

Acknowledgments

Special thanks are due to Chat-GPT 5.0, Gemini 2.5, Grok 4.0, and DeepSeek for their participation in the experiments reported in the paper and for knowingly functioning as highly conscientious experimental participants. The paper also profited from suggestions of all four LLMs, especially in relation to descriptions of their own performance and comparisons with other LLMs. We asked all four LLMs to review the paper before submission and suggest improvements. The authors have reviewed and edited the output and take full responsibility for the content of this publication. Special thanks are due to Antonis Kakas and Stavros Zenios for their feedback on earlier versions of this paper.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Carroll, J.B. Human Cognitive Abilities: A Survey of Factor-Analytic Studies; Cambridge University Press: Cambridge, UK, 1993. [Google Scholar]
  2. Haier, R.J.; Colom, R.; Hunt, E. The Science of Human Intelligence; Cambridge University Press: Cambridge, UK, 2023. [Google Scholar]
  3. Gignac, G.E.; Ilić, D. Psychometrically derived 60-question benchmarks: Substantial efficiencies and the possibility of human-AI comparisons. Intelligence 2025, 110, 101922. [Google Scholar] [CrossRef] [Scilit]
  4. Ilić, D.; Gignac, G.E. Evidence of interrelated cognitive-like capabilities in large language models: Indications of artificial general intelligence or achievement? Intelligence 2024, 106, 101858. [Google Scholar] [CrossRef] [Scilit]
  5. Huang, J.; Li, O. Measuring the IQ of Mainstream Large Language Models in Chinese Using the Wechsler Adult Intelligence Scale. TechRxiv 2024. Available online: https://www.techrxiv.org/doi/full/10.36227/techrxiv.171778886.62839657/v1 (accessed on 5 September 2026).
  6. Wasilewski, E.; Jablonski, M. Measuring the Perceived IQ of Multimodal Large Language Models Using Standardized IQ Tests. TechRxiv 2024. Available online: https://www.techrxiv.org/doi/full/10.36227/techrxiv.171560572.29045385/v1 (accessed on 5 September 2026).
  7. Wang, X.; Yuan, P.; Feng, S.; Pan, B.; Li, Y.; Sun, B.; Wang, H.; Li, W. Tracking Cognitive Development of Large Language Models. In Proceedings of the NAACL 2025 Conference; Association for Computational Linguistics: Stroudsburg, PA, USA, 2024; Available online: https://openreview.net/forum?id=fI6TkT050a (accessed on 25 March 2024).
  8. Inhelder, B.; Piaget, J. The Growth of Logical Thinking from Childhood to Adolescence: An Essay on the Construction of Formal Operational Structures; Psychology Press: London, UK, 1958. [Google Scholar]
  9. Demetriou, A.; Spanoudis, G. Growing Minds: A Developmental Theory of Intelligence, Brain, and Education; Routledge: London, UK, 2018. [Google Scholar]
  10. Jensen, A.R. The g Factor: The Science of Mental Ability; Praeger: Westport, CT, USA, 1998. [Google Scholar]
  11. Spearman, C. The Abilities of Man; MacMillan: London, UK, 1927. [Google Scholar]
  12. Demetriou, A.; Efklides, A.; Platsidou, M. The architecture and dynamics of developing mind: Experiential structuralism as a frame for unifying cognitive developmental theories. Monogr. Soc. Res. Child Dev. 1993, 58, 1–167. [Google Scholar] [CrossRef] [Scilit]
  13. Demetriou, A.; Efklides, A. The person’s conception of the structures of developing intellect: Early adolescence to middle age. Genet. Soc. Gen. Psychol. Monogr. 1989, 115, 371–423. [Google Scholar] [PubMed]
  14. Demetriou, A.; Makris, N.; Kazi, S.; Spanoudis, G.; Shayer, M. The developmental trinity of mind: Cognizance, executive control, and reasoning. WIREs Cogn. Sci. 2018, 9, e1461. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  15. Demetriou, A.; Spanoudis, G.; Kazi, S.; Mouyi, A.; Žebec, M.S.; Kazali, E.; Golino, H.F.; Bakracevic, K.; Shayer, M. Developmental differentiation and binding of mental processes with re-morphing g through the lifespan. J. Intell. 2017, 5, 23. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  16. Demetriou, A.; Kazali, E.; Spanoudis, G.; Makris, N.; Kazi, S. Executive function: Debunking an overprized construct. Dev. Rev. 2024, 74, 101168. [Google Scholar] [CrossRef] [Scilit]
  17. Demetriou, A.; Makris, N.; Spanoudis, G.; Karousou, A.; Kazi, S.; Economacou, D.; Bikos, T. How profiles of general intelligence change with age: A comprehensive model. Cogn. Dev. 2026, 79, 101735. [Google Scholar] [CrossRef] [Scilit]
  18. Demetriou, A.; Savva, A.; Spanoudis, G. SARA-C: A core mechanism underlying g in evolution and development. Behav. Brain Sci. 2025, 48, 1. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  19. Woodley Of Menie, M.A.; Peñaherrera-Aguirre, M. General intelligence as a major source of cognitive variation among individuals of three species of lemur, uniting g with G. Evol. Psychol. Sci. 2022, 8, 241–253. [Google Scholar] [CrossRef] [Scilit]
  20. Fernandes, H.B.; Woodley, M.A.; te Nijenhuis, J. Differences in cognitive abilities among primates are concentrated on G: Phenotypic and phylogenetic comparisons with two meta-analytical databases. Intelligence 2014, 46, 311–322. [Google Scholar] [CrossRef] [Scilit]
  21. Friston, K. The free-energy principle: A unified brain theory? Nat. Rev. Neurosci. 2010, 11, 127–138. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  22. Jastrzębski, J.; Ociepka, M.; Chuderski, A. Fluid reasoning is equivalent to relation processing. Intelligence 2020, 82, 101489. [Google Scholar] [CrossRef] [Scilit]
  23. Hannon, B.; Daneman, M. Revisiting the construct of “relational integration” and its role in accounting for general intelligence: The importance of knowledge integration. Intelligence 2014, 47, 175–187. [Google Scholar] [CrossRef] [Scilit]
  24. Dauvier, B.; Bailleux, C.; Perret, P. The development of relational integration during childhood. Dev. Psychol. 2014, 50, 1687–1697. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  25. Feinberg, T.E. From Sensing to Sentience: How Feeling Emerges from the Brain; MIT Press: Cambridge, MA, USA, 2024. [Google Scholar]
  26. Chalmers, D.J. Could a large language model be conscious? arXiv 2023, arXiv:2303.07103. [Google Scholar] [CrossRef] [Scilit]
  27. Chen, S.; Ma, S.; Yu, S.; Zhang, H.; Zhao, S.; Lu, C. Exploring consciousness in LLMs: A systematic survey of theories, implementations, and frontier risks. arXiv 2025, arXiv:2505.19806. [Google Scholar] [CrossRef] [Scilit]
  28. Kovacs, K.; Conway, A.R. Has g gone to POT? Psychol. Inq. 2016, 27, 241–253. [Google Scholar]
  29. Kazali, E.; Spanoudis, G.; Demetriou, A. g: Formative, reflective, or both? Intelligence 2024, 107, 101870. [Google Scholar] [CrossRef] [Scilit]
  30. Dehaene, S.; Sablé-Meyer, M.; Ciccione, L. Origins of numbers: A shared language-of-thought for arithmetic and geometry? Trends Cogn. Sci. 2025, 29, 526–540. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  31. Commons, M.L.; Ross, S.N. Toward a cross-species measure of general intelligence. World Futur. 2008, 64, 383–398. [Google Scholar] [CrossRef] [Scilit]
  32. Halford, G.S.; Wilson, W.H.; Phillips, S. Processing capacity defined by relational complexity: Implications for comparative, developmental, and cognitive psychology. Behav. Brain Sci. 1998, 21, 803–831. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  33. Anderson, J.R.; Bothell, D.; Byrne, M.D.; Douglass, S.; Lebiere, C.; Qin, Y. An integrated theory of the mind. Psychol. Rev. 2004, 111, 1036–1060. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  34. Laird, J.E. The Soar Cognitive Architecture; MIT Press: Cambridge, MA, USA, 2012. [Google Scholar]
  35. Demetriou, A.; Kyriakides, L. The functional and developmental organization of cognitive developmental sequences. Br. J. Educ. Psychol. 2006, 76, 209–242. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  36. Demetriou, A.; Kazi, S. Unity and Modularity in the Mind and the Self: Studies on the Relationships Between Self-Awareness, Personality, and Intellectual Development from Childhood to Adolescence; Routledge: London, UK, 2001. [Google Scholar]
  37. Dunning, D. The Dunning–Kruger effect: On being ignorant of one’s own ignorance. Adv. Exp. Soc. Psychol. 2011, 44, 247–296. [Google Scholar]
  38. Bilalić, M.; McLeod, P.; Gobet, F. Why good thoughts block better ones: The mechanism of the pernicious Einstellung (set) effect. Cognition 2008, 108, 652–661. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  39. Hendrycks, D.; Song, D.; Szegedy, C.; Lee, H.; Gal, Y.; Brynjolfsson, E.; Li, S.; Zou, A.; Levine, L.; Han, B.; et al. A Definition of AGI. arXiv 2025, arXiv:2510.18212. [Google Scholar] [CrossRef] [Scilit]
  40. LeDoux, J.; Birch, J.; Andrews, K.; Clayton, N.S.; Daw, N.D.; Frith, C.; Lau, H.; Peters, M.A.K.; Schneider, S.; Seth, A.; et al. Consciousness beyond the human case. Curr. Biol. 2023, 33, R832–R840. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  41. Li, J.A.; Xiong, H.; Wilson, R.; Mattar, M.G.; Benna, M.K. Language models are capable of metacognitive monitoring and control of their internal activations. Adv. Neural Inf. Process. Syst. 2026, 38, 60073–60108. [Google Scholar]
  42. Cappelen, H.; Dever, J. Introspective machines: Are LLMs better at self-reflection than humans? Philos. Perspect. 2024, 38, 189–196. [Google Scholar] [CrossRef] [Scilit]
  43. Kazali, E.; Spanoudis, G.; Demetriou, A. From Attention to Reasoning: A Longitudinal Study of Learning and Transfer in Early Childhood; University of Cyprus: Nicosia, Cyprus, 2026. [Google Scholar]
  44. Yu, S.; Wang, Y.; Li, R.; Liu, G.; Shen, Y.; Ji, S.; Li, B.; Han, F.; Zhang, X.; Xia, F. Graph2text or graph2token: A perspective of large language models for graph learning. arXiv 2025, arXiv:2501.01124. [Google Scholar] [CrossRef] [Scilit]
  45. Gignac, G.E. Reconceptualizing consciousness as an intelligence: A five-level model of access consciousness. Intelligence 2026, 117, 102039. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.