The present systematic literature review aimed to map the current state of adaptive gamification and adaptive game-based learning in preschool and early childhood education. More specifically, it examined how these interventions have been deployed and studied, which content areas and educational contexts have received most attention, what theoretical frameworks and adaptive mechanisms have been used, and what kinds of learning and motivational outcomes have been reported. Overall, the 19 included empirical studies suggest that adaptive gamification and adaptive GBL can be a promising direction for young learners. At the same time, the evidence does not yet form a fully coherent or mature field. It is still characterised by methodological variety, uneven theoretical grounding, and several practical questions that remain open.
6.1. Key Findings and Implications
Regarding the first research question, which concerned methodology and assessment tools, the reviewed studies showed considerable methodological variety. Experimental and quasi-experimental designs were present in a large part of the identified studies (14 out of 19), which indicates that researchers are not relying only on descriptive or exploratory accounts. However, this should not be interpreted as evidence that the field has already reached methodological maturity. Several studies still depended on small samples, short implementation periods, researcher-developed instruments, or feasibility-oriented designs, and five studies used randomised controlled trials, with only one however from them incorporating mixed-methods components. That combination is also familiar in the wider adaptive and tailored gamification literature: promising designs are often tested alongside narrow samples, brief interventions, and a still limited accumulated evidence base [
16,
17].
The sample sizes give a clear example of this problem. Eleven of the 19 studies included fewer than 100 participants, while the whole range extended from very small feasibility or usability studies to one large-scale trial. In early childhood research, this is partly understandable, because recruiting young children, obtaining permissions, and implementing interventions in real classroom settings are not simple procedures. Even so, small samples restrict the level of generalization that can be made. Thus, the present evidence can support careful interpretation, but it is not yet strong enough to produce broad claims for all preschool and early primary contexts.
The assessment tools used across the studies also revealed a useful tension. Some mathematics-focused studies used standardised instruments, such as the Test of Early Mathematics Ability, whereas others relied on structured questionnaires, observational measures, or tools developed by the researchers for the specific intervention. Standardised measures give researchers a shared reference point, which makes findings easier to compare across studies. At the same time, they are often too broad to catch what changes during play: a child may try a new strategy, slow down, ask for help differently, or interact with peers in a new way. Intervention-specific tools can record those details more closely, but they bring their own trade-off, since they are harder to compare across studies and may reflect the assumptions of the research team. This issue is especially relevant in game-based learning, where gains may appear in the learning process rather than only in broad test-score changes [
77].
One of the strongest aspects of the reviewed studies was the extensive use of system-generated data. Seventeen studies incorporated some form of learning analytics, including item-response logs, time-on-task metrics, accuracy traces, performance trajectories, response latency, or affective indicators. This is important because adaptive systems do not only present learning tasks. They also record how children respond to difficulty, feedback, game elements, and pacing while the learning activity is taking place. In this sense, log data can show whether children are learning more efficiently, even when a simple pre–post comparison would not make this visible [
24]. At the same time, these traces should not be treated as self-evident proof of learning. For log data to carry real evidential weight, it has to be read against the intended learning goal; a score, click pattern, time measure, or engagement signal does not by itself show that learning has occurred [
8].
Concerning the second research question, mathematics was the dominant content area, appearing in 11 out of the 19 studies (58%). That pattern is understandable. Early mathematics often breaks down into ordered skills, and those skills can be represented fairly easily as levels, mastery checks, or changes in difficulty. This structure also makes live adjustment easier: while a child is still playing, the system can use a recent answer, pause, or error to decide whether to offer a new prompt, repeat a concept, or move to another level [
24,
36,
41,
66]. The concentration on mathematics also shows the limits of the current evidence. Science education, literacy, executive function, computational thinking, creative thinking, and broader developmental skills have not yet received the same systematic attention.
This imbalance should not be treated as a minor detail. Broader early childhood GBL literature indicates that games can support more than early numeracy. Existing reviews report benefits for cognition, motivation, social interaction, emotional development, and engagement. Those findings, however, do not transfer automatically from one setting to the next. A design that supports one objective or age group may lose much of its value when children bring different prior knowledge, when the classroom routines differ, or when the task asks for a less structured type of thinking [
14,
33,
82]. For this reason, the strong presence of mathematics in the present review can be read in two ways. It gives the field a relatively coherent evidence base, but it also leaves several important areas of early childhood learning underrepresented.
The even distribution between preschool-kindergarten children (ages 3–6) and early primary learners (ages 6–9) is also more than a demographic observation. Children in the younger age band usually have limited reading ability, shorter attention spans, developing fine motor skills, and a greater need for adult mediation. Studies involving very young preschoolers make this point concrete: autonomy, instructions, and interaction demands can become barriers in themselves, while developmentally tuned gamified assessment may reduce time-outs and help children stay with the activity [
32,
78]. Therefore, adaptation in early childhood cannot be reduced to the technical adjustment of difficulty. It also has to involve interface simplicity, feedback modality, adult support, emotional load, and the physical or motor demands of the learning environment.
The geographic distribution of the studies was mainly European and North American. The pattern was uneven: eight studies came from Europe, four from North America, six from Asia, and one from Africa, while no studies from South America or Oceania were identified. Combined with the near-exclusive focus on neurotypical children, that distribution limits how confidently the findings can be applied across cultures or to inclusive classrooms. None of the 19 studies focused specifically on children with identified special educational needs, and only one [
80] involved a single child with learning difficulties. This is a significant gap, because adaptive systems are often promoted as a way to support children who do not follow one standard learning path. Research with children who have reading difficulties complicates that claim. Evidence from slightly older children with reading difficulties suggests that this promise is not straightforward [
75], as instructional feedback can be experienced differently depending on children’s needs and prior support. A feedback cue that helps one child notice an error may impose additional processing demands on another child, especially when the task is demanding or the feedback format is hard to interpret [
75]. However, as no studies were identified for the specific educational level and age group, more research with diverse learner populations is needed, not only to support inclusion but also to understand how adaptivity operates when learner variability is more visible.
The strong reliance on researcher-developed platforms (16 out of 19) is another finding with two sides. On the positive side, it shows that researchers are increasingly able to design complex adaptive systems for young learners, something that was considerably more difficult only a few years ago. On the other side, it creates questions about scalability, sustainability, classroom adoption, and long-term maintenance. Only three studies used a commercial product, in all cases My Math Academy by Age of Learning. This is consistent with concerns from the broader game-based assessment literature, where many educational game tools are difficult to access, reuse, or replicate [
8]. For this reason, future work needs to examine not only whether an intervention can work under research conditions, but also whether it can be maintained and used in ordinary classrooms.
The third research question focused on theoretical frameworks, adaptive mechanisms, and game elements. The theoretical grounding across the reviewed studies was useful but uneven. Vygotsky’s Zone of Proximal Development and scaffolding theory was the most frequently used framework, which is expected, since adaptive learning is closely connected with matching instruction to the learner’s current level. Self-Determination Theory was also present, especially where motivation and game element design were central. Flow theory, mastery learning, constructivism, and evidence-centered design appeared in more specific cases. However, three studies did not state an explicit theoretical framework, and several others referred to theory without clearly showing how it shaped the adaptive mechanism.
This lack of explicit theoretical grounding is not a small issue. When theory is only loosely connected to design, it becomes hard to know why a system succeeded, why it failed, or whether the same approach could transfer to another content area. Reviews of tailored and adaptive gamification raise much the same problem. They describe a number of systems that appear to have been built around local design decisions, with limited follow-up evidence about whether any benefits survive after the initial intervention period [
15,
16,
17]. The field therefore needs to spell out more clearly how theory, learner modelling, adaptive mechanisms, game elements, and outcomes fit together.
The adaptive mechanisms were classified into 10 categories, from mastery-based dynamic difficulty adjustment to newer AI and machine learning approaches such as deep reinforcement learning, graph neural networks, procedural content generation, and mood-based adaptation. The later studies in the corpus, especially those from 2024 onward, show a visible shift toward AI and machine learning: instead of relying only on preset rules, several systems attempt to model the learner and adjust content through more complex computational architectures. This development is important, but it should be interpreted cautiously. Many of these systems remain short-term, technically complex, sample-limited, or tied to a very specific context. The broader AI-supported gamification literature also indicates that early childhood still receives much less attention than higher education or general school-level settings [
4].
A central contribution of the present review is the distinction between adapting learning content and adapting game elements themselves. Most included studies adapted content difficulty, pacing, feedback, or learning sequence while keeping the broader game structure fixed. Some studies implemented partial adaptation, where task difficulty changed according to performance, but the gamification framework remained the same. By contrast, only a small set of studies adjusted game elements in response to learner characteristics, and adaptation based on player type was especially uncommon [
43,
45]. For that reason, a large share of the work labelled adaptive gamification in early childhood looks closer to adaptive learning wrapped in game features [
36,
63,
65,
83].
This distinction matters in practice as well as in theory, because game elements carry effects of their own. In a preschool classroom, a badge, timer, avatar, reward, or leaderboard is not just decoration; it changes what children attend to and how they feel while working. One child may treat a reward as a small invitation to keep trying; another may read it as pressure to hurry, beat peers, or collect points before thinking through the task [
45,
84,
85]. The same design feature can feel playful, irrelevant, or stressful depending on the activity, the classroom climate, and the child’s earlier experiences with games or competition [
22,
46,
86,
87]. Research on intelligent game-based learning gives a similar warning: incentives, personalised agents, and navigation tools can make an activity easier to enter, but they can also become so visible that the learning goal moves into the background [
88]. The included studies show this tension as well, with positive motivational responses sometimes appearing alongside stress linked to badges, currency, cooperation, or adaptive challenge [
20,
45]. Adaptive gamification therefore needs a more precise question than whether game elements are motivating. It needs to ask which element supports which child, in which classroom situation, and at which level of challenge [
46,
87].
Regarding the fourth research question, the learning and motivational outcomes were encouraging, but uneven. Most studies reported positive learning trends, whether through statistically significant gains, improved task accuracy, learning efficiency, or descriptive performance improvements. Where effect sizes were reported, they ranged from small to quite large. Still, these findings should not be overstated. Some studies were mainly feasibility, usability, or assessment-oriented studies rather than controlled intervention trials. For that reason, they should be interpreted as part of a cautiously positive evidence base, rather than as proof that adaptive gamification is already effective for all young learners [
83].
The null and mixed findings are particularly useful for interpreting the field. In at least one carefully designed comparison, the adaptive version of a game did not clearly outperform the non-adaptive version on cognitive or affective outcomes, even though both groups improved over time [
77]. Several explanations could account for this result: the training may have been too brief, the standardised tools may have been too blunt, or the algorithm may not have changed the learning path enough to matter. Even so, the result is a useful reminder that adaptivity is a design hypothesis, not a built-in advantage [
83]. What matters is the object being adapted, the way the mechanism operates, whether children notice and understand the change, and whether the selected outcome measure is sensitive to that change [
65,
89].
Learning effectiveness and learning efficiency also need to be separated. In some studies, children in the adaptive condition finished with outcomes close to those of the comparison group, but they got there with less time or fewer unnecessary steps [
24,
63]. For classroom practice, that difference is not trivial. A system that reduces repeated attempts, waiting time, or unnecessary challenge may be valuable even when final scores look similar. Future studies should therefore examine score gains alongside pacing, cognitive load, and the amount of classroom time needed to reach a comparable learning outcome [
36,
90].
Prior knowledge was the most frequently identified moderating variable. Children with lower or moderate prior knowledge sometimes appeared to gain more from adaptive systems, especially when the challenge stayed near what they could manage with support. The pattern, however, does not support a simple rule in which less prior knowledge always means greater benefit. A child who lacks basic concepts may need modelling, shorter steps, or more explicit prompts. A child who already understands part of the material may need the opposite: fewer interruptions, less explanation, and a task that stretches existing knowledge. Feedback can therefore act as support, distraction, or overload depending on the child’s starting point [
40,
63,
66]. Prior knowledge is better treated as part of a wider profile that includes feedback, scaffolding, task difficulty, and the specific adaptive algorithm [
63].
The findings regarding gender in education are also worth mentioning, but with caution. The gender-related result is suggestive rather than settled. In one science education study, adaptive gamification appeared to reduce or remove gender differences in learning gains compared with traditional inquiry-based instruction. The authors link this result to adaptive narratives and positive role models, which may have made the science task feel more supportive for female students [
43]. However, this remains a limited finding, and the mechanism is not yet clear. It should therefore be treated as a promising direction for future research rather than as a firm conclusion about gender equity in adaptive gamified environments.
Motivational and affective outcomes were assessed in 16 out of 19 studies, although the instruments used were less standardised than those used for cognitive outcomes. Most findings moved in a positive direction, with enjoyment, engagement, preference, satisfaction, persistence, or willingness to continue reported across several studies. In several cases, children who spent longer with the adaptive activity also advanced further, suggesting that persistence may be part of the learning pathway rather than only a separate motivational outcome [
36,
41,
66,
78]. Still, enjoyment and engagement should not be read as automatically positive. Children can be engaged and still experience stress, anxiety, cognitive overload, or reward-focused participation. Future work should therefore examine which game elements generate meaningful engagement and which ones may create affective load.
Across the four research questions, the findings suggest a field that is active and promising, but also still unsettled. The presence of game elements, personalisation, or AI-supported adaptation is not enough on its own to guarantee stronger learning or motivation. The value of an intervention depends on several interacting conditions: the granularity of the adaptive mechanism, the developmental appropriateness of the interface, the theoretical grounding of the game elements, the sensitivity of the assessment tools, the learner’s prior knowledge, and the classroom context. This reading is close to the broader gamification literature, where mixed results are often explained by how well the intervention is implemented, what motivational assumptions guide it, and whether the mechanics are actually tied to the learning objective [
13,
17,
88].
These findings also point to some new perspectives regarding practice. For preschool educators, adaptive systems should not be seen as tools that aim to replace teacher judgement or role, but rather as tools that can help teachers notice where children need support [
22]. This is extremely valuable since children’s prior knowledge and developmental stage seem to affect how adaptivity works [
77]. Also, faster progress inside a system does not always lead to stronger final learning outcomes [
24]. Therefore, teachers need information that is simple, understandable and usable, such as whether a child is progressing because their learning has truly improved, or their task sequence is more efficient, or they have simply remained engaged longer. This is vital in early childhood settings, where digital games should also have child-friendly interfaces, clear visual instructions, appropriate touch interaction, and, in many cases, adult mediation [
33]. As such, the practical value of adaptive gamification and game-based learning in early childhood seems to depend not only on the algorithm itself, but also on whether teachers can interpret the adaptation and connect it with classroom decisions [
91].
6.2. Limitations
The present systematic literature review contains certain limitations that should be acknowledged. First, it was restricted to English-language studies, and this may have excluded relevant work published in other languages. This point is important for early childhood education, because practices around play, technology use, and teacher mediation are often culturally and linguistically situated. Some approaches may therefore be present in local research traditions without appearing in the English-language literature. In addition, for the sources that returned very large result sets, screening was limited to the first 1000 entries ordered by relevance. However, relevance is affected by the platform and can change over time, which can affect the exact reproducibility of the search the longer the time passes. What is more, the search focused on studies published between 2016 and 2026 and used selected academic databases. This date range kept the review close to current work on adaptive gamification and AI-supported personalisation, but it may have missed older foundational studies or papers that described related ideas with different terminology [
45].
Regarding the process of how the SLR was conducted, the present review indeed holds some limitations as the main screening, selection, and extraction process was initially done by a single researcher, potentially increasing the risk of false exclusions, extraction errors, or unconscious bias, which can be near 5% [
74]. To reduce this restriction, a second researcher with experience in systematic literature reviews checked the search results, screening decisions, extracted data, and coding decisions afterward. However, this procedure is weaker than fully independent dual screening. Therefore, the findings should be interpreted with this methodological limitation in mind.
Moreover, the included studies were heterogeneous in design, population, duration, content area, adaptive mechanism, and outcome measures. Some compared adaptive and non-adaptive versions of the same game, others compared adaptive systems with traditional instruction, and others examined feasibility, usability, or engagement without a direct comparison condition. The MMAT appraisal identified recurring uncertainties concerning sampling procedures, measurement validity, attrition, baseline comparability, and the reporting of intervention procedures. Larger cluster-randomised studies tended to report clearer procedures, baseline checks, and retention, while smaller pilot, usability, or platform studies often provided less detail in these areas. This diversity reflects the innovative character of the field, but it also limits direct cross-study comparison and makes formal meta-analysis difficult. In addition, many studies used small samples and short intervention periods. In early childhood research, that constraint is understandable, but it still narrows generalisability and leaves open the question of whether reported benefits would last over time. Gamified interventions add another complication. Badges, rewards, or classroom routines may feel exciting during the first sessions and far less meaningful once children know what to expect [
15]. The review also combined journal articles and conference papers, two formats that can differ in peer-review rigour and reporting detail. In some cases, theoretical frameworks, adaptive mechanisms, or assessment procedures were not described with enough detail in the original publications, making coding and interpretation more difficult. Therefore, the findings and conclusions of the present review should be read as a synthesis of the available evidence, showing promise but not as a definitive statement on the effectiveness of each adaptive approach.
Publication bias should also be considered. Because positive or statistically significant results are easier to publish and index, the available corpus may overstate the field’s level of success. The limited number of null or mixed findings in the final corpus may therefore say as much about what reaches publication as about how often adaptive gamification succeeds in practice [
20,
77].
Finally, the field is moving quickly. AI, deep reinforcement learning, mood recognition, and multimodal learner modelling are now entering this area, and they may change what counts as adaptive design in the next few years [
92,
93,
94]. Some technologies included in this review may therefore date quickly, while studies published after the search period may introduce more advanced adaptive architectures [
4,
93,
95,
96]. Future reviews will need to return to the evidence base and examine whether these newer systems produce outcomes that are not only stronger, but also durable, inclusive, and workable with young children in real classrooms [
14,
93,
94].