1. Introduction
The integration of artificial intelligence (AI) resources and tools in education can no longer be described as a peripheral innovation; on the contrary, it is structurally altering the ways in which knowledge is produced, validated, and acquired. In the specific context of music education, the “algorithmic shift” is not limited to incorporating tools that assist in the creation, recording, or editing of sound; it also has a decisive impact on the most essential and unique aspects of music learning: understanding the relationship between sound events, the attribution of aesthetic-cultural meaning to musical discourse, and the construction of musical narratives at the intersection of theoretical-technical knowledge, perception, and musical experimentation. It is shifting music pedagogy from a purely formalist and structuralist vision toward a phenomenological and cognitive understanding that places the subject (whether performer, listener, or student) at the center of the creative process. From this perspective, the penetration and influence of AI in the music education ecosystem represents an unprecedented vector that promises new opportunities to personalize, provide feedback, and support creative processes, reconfiguring roles, and epistemological or methodological perspectives (
J. F. Merchán Sánchez-Jara et al., 2024). In this context, this paper proposes to analyse the possibilities of integrating AI into the music education ecosystem as a pedagogical agent capable of deploying unprecedented and innovative potentialities aimed at reinforcing the understanding and comprehension of the musical phenomenon and the construction of meanings, converting complex concepts, formal structures, and sound narratives into a more intelligible and accessible insights, which ultimately contribute to enhancing the creativity and cognitive development of students.
The paper analyses these issues from a framework in which music teaching and learning constitutes an ecological environment, in an epistemological sense, which arises in response to the tendency in academia to divide knowledge into discrete and autonomous units.
Boyce-Tillman (
2004) proposes thinking of music education as a system of relationships where technique, perception, culture, and experience are interrelated facets that define each other and function as an indivisible whole. When these connections are broken or weakened, learning is reduced to isolated, disconnected, and fragmented skills that favor half-hearted musical learning, with no significant impact on the understanding and development of musical thinking.
Traditionally, access to in-depth musical knowledge has been associated with barriers inherent to musical knowledge related to notational literacy, the development of instrumental technique over long periods of time, or the critical appropriation and rigorous handling of a complex conceptual and terminological field in the context of highly abstract theoretical frameworks. In many academic contexts, these barriers have operated as filters that separate musically competent students from non-competent students, reinforcing previous inequalities and limiting equal participation in music. The very discussion about the role of notation in teacher and school training shows how teaching the reading of texts in common Western notation systems can become an obstacle when it is presented as an isolated prerequisite rather than being integrated into meaningful musical activities (
Fautley, 2017). This issue is central to the argument posited throughout this paper: if the functionalities of AI resources are integrated solely as a circumstantial and instrumental element, geared towards technical tasks or automating exercises, their impact will be superficial; on the other hand, their integration as an agent of semiotic mediation (as a means of working with conceptualizations and representations) can contribute to reconfiguring music as a horizon of meanings shared by more students.
The central hypothesis defending this perspective is that AI can act as a “cognitive artifact” (
Cassinadri, 2024;
Norman, 1991;
Rivera-Novoa & Duarte Arias, 2025) that makes highly abstract and difficult-to-objectify aspects of musical logic visible and intelligible under the supervision of the teacher: harmonic patterns, rhythmic hierarchies, narrative tensions, timbral configurations, and formal processes whose perception may be intuitive for the novice listener but whose analytical understanding is neither immediately accessible nor easily articulated. This type of mediation does not aim to replace the teacher’s judgment or the physical experience of making music; it aims to redistribute the cognitive load so that students invest mental resources in higher-order processes: listening (
Gordon, 1989), interpreting, comparing, justifying, imagining alternatives, and making informed decisions in the artistic–aesthetic sphere. In terms of cognitive load theory, the aim is to reduce extrinsic load and free up resources to construct musical schemas (
Sweller, 1988), drawing on multimedia learning principles that combine verbal explanation with visual and auditory representations that promote conceptual understanding (
Mayer, 2008).
The approach is also articulated with the precepts that inform educational ecosystems in the field of musical practice in connected environments, based on a conception of the ecology of learning, understood as a dynamic network of knowledge, agents, experiences, and mediations (
Leman, 2007) that transcends the mere accumulation of technological resources. Reviews of “educational ecosystems” on the social web show that musical practice expands into learning networks that transcend the classroom through communities, platforms, repositories, and spaces for collaborative creation (
J. F. Merchán Sánchez-Jara, 2016;
J. Merchán Sánchez-Jara et al., 2022). Generative AI and learning analytics systems enable an unprecedented qualitative leap by offering evidence-based personalized scaffolding (
Giannakos & Cukurova, 2023). In this context, personalized scaffolding refers to educational support strategies adapted to the needs, pace, and learning styles of each student, facilitating understanding and independent practice. These strategies are dynamically designed based on multimodal data that reflect individual performance, interaction, and preferences, allowing for more precise and effective interventions. From this perspective, the design of these scaffolds should not be understood as a simplification that dilutes the optimal challenge inherent in the creative process, but rather as a strategy for managing its complexity.
Recontextualizing Elliott’s (
1995) proposals, AI mediation must ensure that the learning process maintains a level of demand that stimulates the student’s musical thinking, preventing technology from acting as a substitute that reduces the experience to passive or superficial learning.
In the context of AI applied to music, and in relation to data analysis, there is a crucial need to move beyond sound pattern modeling to address the interpretive dimension of the musical phenomenon: the construction of significance.
Steels (
2021) argues that data-driven approaches often reduce intelligence to mere pattern prediction, thereby sidestepping more humanistic perspectives on musical semantics. He calls for a shift toward tools oriented toward the development and application of shared understandings. In this sense, and in a complementary way,
Clancy (
2023) proposes the “ecosystem” as a transdisciplinary network of actors and practices (developers, students, and artists) in which the construction of musical meaning becomes an explicit axis of analysis and action, highlighting how interaction between different perspectives and disciplines is key to not limiting AI to data processing, but rather integrating it into creative and cognitive processes that consider the aesthetic and expressive experience of music as a fundamental dimension. Based on these premises, it is possible to develop an interpretive framework based on educational intervention strategies that allow AI to be positioned as a facilitator of processes of understanding and constructing musical meaning.
The horizon proposed by
Steels (
2021) on the construction of meaning takes on a critical dimension by recognizing that the musical elements analysed through AI are not treated as neutral atoms of information, but as repositories of the history and social tensions of their time. Under this premise,
Adorno (
1970/2004) postulates on the historicity inherent in all artistic material resurfaces. The risk of a purely data-driven approach lies in the dehistoricization of the material, reducing the sound components to discrete, empty data stripped of the layer of aesthetic memory. Consequently, it is essential that AI be integrated as an actor that allows formal data to be clearly mapped and structured so that students can focus on unraveling the historical density during the process of musical analysis itself. This line of thinking is in line with
Anders’s (
2021) ideas about proposing an interaction where technology serves to reveal the network of cultural relationships that define the musical work. Under this premise, technology transcends its role as a standardized generator to emerge as a hermeneutic prism that facilitates the recovery of the historical and semantic meaning of music. This new approach acts as a catalyst for students to interact with musical material in a conscious way, integrating the historicist strata that make up the musical text.
Several studies confirm that facilitating this expanded understanding can increase the quality and diversity of students’ musical creativity by enabling informed exploration without requiring prior theoretical and technical mastery (
Burnard & Younker, 2004;
Lam, 2024) and/or enhancing cognitive benefits associated with music education by intensifying processes of attention, working memory, executive control, or schema formulation (
Habibi et al., 2018;
Schellenberg, 2004,
2005).
This study is situated within a theoretical–conceptual research approach, aimed at constructing interpretative frameworks and developing theoretical propositions on an emerging object of study (
Jaakkola, 2020). In this sense, it is conceived as a theoretical–conceptual (theory-building) study on AI in music education, aimed at articulating and extending existing conceptual frameworks in the field. We adopt an interpretive, sociocultural epistemological stance, viewing learning as the construction of meaning mediated by cultural context. The objectives focus on developing an integrative explanatory model of AI in music education through conceptual analysis, literature synthesis, and theoretical model construction, without collecting new empirical data. The scope is limited to elaborating an expansive theoretical framework, which implies a limitation in terms of empirical validation but allows the clarification of key concepts for future empirical research.
2. Theoretical Framework: Mediation, Comprehension, and Musical Meaning
The real impact and significance of AI in music education can only be understood through a theory of mediation that explains how tools, signs, and cultural practices develop cognitive activity and transform aesthetic experience. From a sociocultural perspective, learning is understood as an active process in which students construct meaning through social interaction and the mediation of signs and symbolic tools. In this framework, the notion of scaffolding (
Wood et al., 1976) allows us to conceptualize temporary supports that make it possible to tackle tasks that exceed the current competence of students. When AI acts as adaptive support (rather than a mere automatic corrector), it acts as a mediator within the “zone of proximal development” (
Vygotsky, 1978), promoting students’ progression from guided execution to the development of interpretive and creative autonomy. In this way, AI not only facilitates the acquisition of technical skills, but also contributes to the development of more complex cognitive and expressive processes. The idea of AI as a “cognitive artifact” (
Norman, 1991) is particularly eloquent in this context, because much of musical knowledge is shaped as a relational, hierarchical, and multimodal entity. Understanding a work involves simultaneously and symbiotically handling melodic, rhythmic, and formal lines and structures; stylistic, historical, and cultural expectations and considerations; expressive-emotional links; and temporal narratives (tension-relaxation, direction, contrast, dynamics). Traditional music education has tended to translate this complexity into notational and/or terminological simplifications, ignoring the fact that the phenomenon transcends notation, which is merely a culturally situated representation based on a reductionist process (
J. Merchán Sánchez-Jara et al., 2017). Limiting music education to the mere reading of scores can impoverish the processes of understanding, by turning notation into an end in itself rather than a tool for interpreting, reflecting on, and constructing musical meanings (
Fautley, 2017). From this perspective, a technological-pedagogical framework based on AI technologies can promote the integration of various forms of representation of the same sound object: sheet music, recordings, spectrograms, harmonic diagrams, semantic annotations, body gestures, or code.
Nattiez’s (
1990) tripartite proposal, which describes the musical event through three interrelated dimensions—the poietic process (composition, performance, creative decisions); the trace or neutral level (material support of the musical phenomenon in the form of text or score, recording, digital file, data); and the aesthetic process (processes of reception, perception, and interpretation by the listener)—allows us to delve deeper into how these multiple forms of representation contribute to the construction of musical meaning. This model is particularly useful due to its ability to delineate the areas of intervention of AI in the different facets of the musical event; it demonstrates its capacity to influence poietic processes through mediation in creative decision-making (harmonization suggestions, generation of variations, assistance in orchestration); to reconfigure the trace through the translation, structuring, and formalization of sound objects into digital representations (converting audio into data; segmenting, labeling, representing); or to transform aesthetic processes by modulating the conditions of listening, interpretation, and attribution of meaning (focused listening, isolation of voices, visualization of components, didactic recommendation). In purely pedagogical terms, the essential question is not to verify the possibilities of AI as a generative agent (creating music), but to analyze the functionalities (and their didactic implementation) that allow students to gain an in-depth understanding of the structural, expressive, and stylistic principles that organize a musical work, identify the compositional, interpretive, and technological decisions that shape it, and reflectively explore the possible aesthetic alternatives within a style or at the frontier of its conventional frameworks. This issue can also be read as a tension amidst opposed epistemological perspectives.
Boyce-Tillman (
2004) distinguishes amongst analytical knowledge (A) and experiential knowledge (B). Analytical knowledge is oriented toward the production of concrete results, emphasizes objectivity, is based on impersonal logic, and is validated through tests, verifiable criteria, or external evidence. This type of knowledge is characteristic of contexts where precision, replicability, and methodological clarity are essential, such as in formal science, and where rigorous control over processes and final products is sought. In contrast, experiential knowledge is oriented toward being, personal involvement, and the emotional dimension of understanding. Such knowledge is constructed through direct experience, sensory perception, and subjective reflection, favoring associative and contextual ways of understanding reality. This modality predominates in fields such as musical performance, arts education, and creative practice, where intuition, empathy, and emotional involvement enrich both the learning process and the significance of the practice.
In AI-driven musical environments, the bias toward the formalization of analytical knowledge is intensified when systems are used, in a decontextualized manner, as performance correctors or optimizers, prioritizing objectifiable results and statistical patterns over aesthetic–artistic experience. In contrast, when AI performs a semiotic mediation function, it is possible to establish areas of concomitance between sensory and performative experience (characteristic of experiential knowledge) and analytical representations, without reducing music to the merely quantifiable. This consideration of the tendency to privilege the formalizable finds a paradigmatic parallel in the position of
Steels (
2021), who points out that the shift toward behaviourist approaches in AI (statistical learning, pattern recognition, and prediction) displaces goals, intentions, and symbols, relegating reflection on meaning to ostracism. This perspective suggests that the application of educational AI in music should not be limited to analytical knowledge of structures and rules, but should also enhance interpretive understanding and expressive sensitivity, integrating technical precision with artistic meaning and awareness. In this sense, AI-mediated environments should promote musical learning that considers both the internal logic of music and its aesthetic and meaningful dimension. For this mediation to be effective, it is essential to neutralize the “black box” nature of current generative systems (
Arrieta et al., 2020). In this light, the pedagogical value is linked to the tool’s ability to constitute itself as an “explainable AI” (XAI) that provides transparency about its inference processes. By making explicit the weights of the structural variables and the hierarchy of the formal features detected, the system allows students to audit the technical validity of the data before proceeding to their own aesthetic interpretation.
To overcome the opacity of current architectures, it is interesting to consider the paradigm of human-centered AI proposed by
Shneiderman (
2022). The author argues that technology should serve to shield the hegemony of the individual’s conscious decision-making and ensure understanding of the different levels of analysis used by AI in its decision-making. In the music education ecosystem, this ontological shift implies that AI transcends its role as a provider of closed solutions to become a partner in co-creativity (
Agres et al., 2021). From this premise, AI functions as an interface for reflection that provides students with an enriched perspective, broadening their creative horizons and allowing artistic intentionality to emanate from a deeper and more meaningful aesthetic understanding.
This position is supported by the literature on technologies in music education, which shows how digital tools can mediate between analytical knowledge and experiential knowledge. Even before the widespread use of AI, paradigm shifts were observed in connection with the curricular use of certain digital tools that appeal to both the structural understanding of music and the capacity for interpretive expression (
Ho, 2004).
Buonviri and Paney (
2020) in aural skills highlight how digital tools not only facilitate auditory training and feedback (aspects linked to analytical knowledge) but also enable students to develop their musical sensitivity and recognize expressive nuances, dimensions that are characteristic of experiential knowledge. Similarly, in the field of composition,
Kardos (
2012) emphasizes how technology can “make different sound worlds accessible” to students with less theoretical and technical competence, offering interfaces that simplify the manipulation of musical materials and allow them to concentrate on exploring timbres and the significance of different creative ideas, thus integrating formal analysis and musical experimentation. The educational value of these contributions lies not in the novelty of the device, but in its ability to reconcile technical mastery and experiential understanding of music in a balanced way.
Two particularly relevant contributions in the field of pedagogical theory add to this argument. First, cognitive load theory (
Sweller, 1988) explains why the overload of peripheral processes (copying symbols, handling opaque interfaces, performing mechanical tasks) acts as a barrier that relegates conceptual learning and meaning construction to the background. Similarly, the principles of multimedia learning and instructional design (
Mayer, 2008) propose strategies for organizing information in order to enhance perception and processing. More specifically, these principles advocate combining different information channels (audio, visualization tools, or natural language) in a coordinated manner to facilitate the understanding of complex concepts, reinforce memory, and promote the transfer of knowledge to other contexts. In the specific context of music education, these issues are addressed by the possibility of graphically representing harmonic structures or complex rhythms while enabling simultaneous listening and reinforcing the narrative with textual annotations, so that students can integrate analytical understanding with direct musical experience. Applied to AI-mediated environments, these fundamentals legitimize the use of technology as a strategy for translating complex musical structures into conceptually understandable and narratively meaningful representations.
In short, AI can act as a cognitive and semiotic mediator, helping to interpret and construct musical meaning from its different forms of representation (trace) (notation, sound interfaces, visualizations), deepen understanding and aesthetic perception (aesthetic), and amplify the development of creativity in students (poietic). The challenge is to transform this potential into a coherent thread within curriculum design and integration, so that technology is not limited to reproducing practices associated with the development of analytical knowledge (A), but rather facilitates an iterative cycle that integrates understanding, meaning construction, musical creation, and interpretation, reconciling both technical development and the expressive and cultural experience of students.
Learning Ecologies in Music Education: From the Classroom to the Expanded Educational Ecosystem
At a time like the present, the conception of the music classroom as an educational ecosystem necessarily implies shifting the focus from a closed physical space to a dynamic network of interconnected interactions, resources, repertoires, technologies, and social and cultural practices. From an ecological perspective, learning is not understood as a mere transmission of predefined meanings, but as an emerging process that is constructed in the relationship between musical activity, technological mediators, analytical knowledge, and the expressive experience shared in the learning community. Recent studies further conceptualize AI-mediated musical environments as distributed ecosystems in which learning and meaning emerge through the interaction of human actors, technological systems, and cultural practices, rather than through linear transmission models (
Zhao et al., 2025). In this framework, artificial intelligence does not act solely as a tool for automation or technical correction, but as a mediation device that can promote structural understanding, aesthetic interpretation, and creative exploration.
This orientation is supported by various studies that emphasize that the educational potential of AI is fundamentally based on the creation of open environments that promote reflection, the negotiation of meanings, and the collaborative construction of musical knowledge through the mediation of computational assistance (
Holland, 2000). The integration of AI into a learning ecology based on pedagogical criteria should contribute to coherently linking the formal dimension of music with its experiential dimension, favoring dynamics in which understanding, interpretation, and creation are configured as interdependent practices at the service of students’ cognitive, artistic, and cultural development (
Lee, 2025). Recent empirical research further supports this perspective, showing that AI-mediated musical environments facilitate the integration of theoretical knowledge with experiential and creative practice, enabling learners to explore, experiment, and construct meaning within authentic musical contexts (
Zhuang & Li, 2025).
The design and implementation of these ecological environments must consider the balanced and symbiotic articulation of multiple dimensions related to musical learning and practice. These include critical listening practices, creation, editing, and production of musical materials; generation of systemic spaces for debate, reflection, and metacognition; the study, analysis, and progressive acquisition of skills related to musical languages (notation, theoretical concepts, interpretive metaphors, etc.); meaningful interaction and functional mastery of various artifacts (instruments, software, digital platforms, etc.), as well as knowledge, understanding, and enculturation in cultural frameworks that legitimize criteria for evaluation, interpretation, and appreciation of different musical manifestations. Whilst similar forms of educational ecology were already present in the social web environment before the widespread use of generative AI, where digital platforms and networks expanded and normalized spaces for informal learning, collaboration, and shared music production (
J. Merchán Sánchez-Jara et al., 2022), current and emerging developments have substantially deepened and reconfigured this ecosystem. In this new context, the learning ecology incorporates an additional layer of mediation, reconfiguring itself as an interpretive and explanatory environment, insofar as systems can infer patterns, model musical behaviours, and generate representations that function as mediators of meaning. These forms of mediation not only support the structural understanding of music, but are also key to analyzing how students construct, negotiate, and attribute value and value to what they hear, interpret, and produce (
Zhao et al., 2025).
In this context, it is possible to optimize these educational ecosystems by paying attention to the different actors involved in the configuration of algorithmic mediation. This does not depend solely on the practices of teachers and students, but also on the (determining) decisions incorporated into the design of systems, platforms, and computational models, in which different biases are introduced and coded: what is considered “good music,” what forms of participation are most valued, or what dimensions of learning are of interest to systematize as measurable. In this sense, the perspectives of developers, students, and artists converge in the same educational ecosystem (the classroom and its extension), in which different biases are iteratively integrated, based on external design decisions that influence the construction, negotiation, and validation of musical meaning (
Clancy, 2023). In this order, teaching designs that can be oriented toward
empowering teachers and students, rather than imposing predefined criteria, are imposed. Thinking of teaching as interconnected and interdependent implies organizing what occurs before, during, and after the classroom, strategically distributing lectures, guided practices with clear objectives, reflective exchanges about music, and offering feedback that enables autonomous understanding. In this vein, experiences with AI-supported flipped classrooms show that engagement and performance improve when technology reorganizes study and practice routines and frees up face-to-face time for metacognitive activities (analytical listening, argumentation, interpretation, and creation), rather than simply adding auxiliary tools (
Lv, 2023).
4. The Instrumented Classroom: In Search of Understanding Musical Processes Through Multimodal Data
Multimodal analytics in learning environments proposes integrating audio, video, software interaction, and, in some cases, biometric or motion signals to understand how students learn (
Giannakos & Cukurova, 2023). In music, this perspective opens up an unprecedented and disruptive avenue that allows for the evaluation or analysis of performance, not only by the final result (validation of the final interpretive product) but also by the trajectory of decisions, strategies, revisions, tests, and justifications that inform or justify this resulting product. This range of new possibilities enables evidence-based multimodal analysis that allows, in the field of improvisation, for example, to detect whether the student always resolves phrases with the same cadence, whether they avoid modulations due to insecurity or insufficient competence, whether they tend to simplify rhythms in specific passages, or whether their intonation or technical precision deteriorates in areas of high motor load.
Beyond the mere observation of technical or gestural patterns, multimodal analytics allows us to explore the cognitive and procedural dimensions of musical learning. Every gesture, pause, correction, decision, or interpretive choice becomes a significant piece of data in the cognitive map that guides the student’s decision-making. These reflect technical skills, as well as planning strategies, approaches to performance, and conceptual understanding of language and discourse. This wealth of data makes it easier for teachers to identify not only “what” problems arise, but “why” they occur. In creative practice (creation, performance, or improvisation), multimodal analytics allows us to visualize the progression in the construction of meanings or the development of musical ideas over time, detecting how a student experiments with variations, modulations, or textures, or how they modulate the set of decisions made in response to challenges associated with sound perception or the interpretive process.
This broadening of focus, from product to process, from technical gesture to the cognitive map that motivates it, places multimodal analytics at the heart of a learning ecology where data not only describes how music is performed or produced, but also how the metacognitive processes that articulate its meaning are articulated. However, this ability to make trajectories, decisions, and patterns at the procedural level visible is not susceptible, as is the case with all resources and artifacts that mediate sociological and/or cultural processes, to introducing biases and privileging or denigrating different perspectives or positions; this question depends largely on which repertoires, practices, and cultural frameworks are integrated into the systems that analyze and provide feedback on learning. Thus, the same technology that allows for an in-depth understanding of creative processes can also delimit which types of music are legitimized and which forms of meaning are considered valuable. By virtue of this consideration, a classroom as an ecological system is also a space for cultural selection. AI can reinforce a canon if it is fed with homogeneous datasets, but it might also enable additional plural spaces of meaning if it is designed to explore diverse repertoires and to contextualize heterogeneous musical narratives. In this sense, it is necessary to pay attention to how state-of-the-art music education frameworks emphasize the need to question dominant epistemologies and to review how musical knowledge is legitimized (
Hess, 2015;
Kallio, 2020). An educational ecosystem mediated by AI technologies aimed at promoting understanding must therefore critically and deliberately incorporate the sociocultural contexts, situated practices, and systems of meaning that shape the musical experience, rather than limiting itself to decontextualized stylistic classifications or abstract formal taxonomies (genre, form, style, era, etc.). Recent studies further stress that AI-mediated musical environments are not neutral but are embedded in sociocultural and epistemological frameworks that shape what is recognized as valid musical knowledge and practice, thus requiring critical attention to issues of representation, inclusion, and meaning-making (
Zhao et al., 2025).
5. From Musical Meaning to Metacognitive Creativity
As recent research points out, musicians’ interaction with artificial intelligence tools is significantly correlated with the perception and development of creativity in areas such as musical composition, by integrating cognitive and meaning-construction processes (
Ma et al., 2025). The mediation of AI resources in musical understanding makes creativity a constitutive dimension of the act of understanding, not as a circumstantial addition, but as a facet inherent to the very process of meaning construction. In this framework, AI is not limited to facilitating sound production, but organizes and conditions the processes of decision-making, exploration, and meaning construction. In this way, creativity is placed at the core of the cognitive and semiotic process of music, emerging as an inherent consequence of technological mediation in the musical experience.
By the same token, creating does not consist (solely) of producing original sound artifacts, but rather involves constructing coherent discourses (conveyors of meaning) through conscious decisions within a system of constraints (stylistic, functional, instrumental, cultural, social) while exploring, imagining, and evaluating possible alternatives. From this perspective, creativity is a situated process of informed decision-making, in which understanding acts as an enabling condition for generating coherent and meaningful discourses.
Various studies on music education show that the development of creativity in students rarely unfolds a linear or predictable manner, but instead emerges through idiosyncratic trajectories that combine exploration, contrast, evaluation, and constant refinement (
Burnard & Younker, 2004). This iteration and reflective nature place the act of creation as a process guided by metacognitive control mechanisms, in which the subject consciously supervises, adjusts, and redefines their own creative strategies. From this perspective, musical understanding does not necessarily precede creation, but is dynamically configured in the creative process itself, reinforcing the concept that all mediation (including technology) directly affects the ways in which students construct meaning and, therefore, the way in which they elaborate and develop musical proposals consistent with their own aesthetic perception and with the educational and experiential framework in which they participate. Along the same lines, recent work on creativity and technology in music education points out that digital tools enhance creativity when they facilitate open exploration, collaboration, and critical reflection, and not when they are reduced to offering guides, templates, or closed structures that limit the student’s margin of decision and autonomy (
Lam, 2024).
From the perspective of procedural development, a frequent criticism of certain technology-mediated practices is that they encourage the development of a type of superficial creativity based on mechanical interactions that are cognitively decontextualized and driven by trial-and-error processes.
Bown (
2021) posits, in this regard, that typical interaction with generative systems often consists of sequenced cycles that are limited to adjusting parameters and cherry-picking results as a kind of trial-and-error search that can produce plausible results without affecting or promoting understanding. In loop-based composition, for example, the student creates sound discourses by dragging preconfigured blocks, listens to the results, and chooses what “sounds good,” guided primarily by the most immediate aesthetic impact of the sound. In this case, the creative response is based on a primary perceptual reaction, linked to the sensory effect, rather than on a reflective aesthetic judgment capable of integrating structural, stylistic, or expressive criteria into a deeper understanding of musical discourse. This can be further enriched by considering the notion of scenario-based interaction proposed by
Nika et al. (
2017), in which human–computer musical improvisation is structured through predefined yet flexible frameworks that guide interaction without constraining creativity. In this model, AI systems do not merely react to user input, but operate within evolving scenarios that encode stylistic constraints, interaction rules, and temporal developments, enabling a form of situated co-improvisation. From a pedagogical standpoint, this approach is particularly relevant because it externalizes the underlying logic of musical interaction, making explicit the conditions that shape creative decision-making. Rather than leaving students in an unstructured exploratory space or reducing creativity to trial-and-error processes, scenario-based systems provide a scaffolded environment in which learners can engage with musical form, anticipation, and responsiveness in a controlled yet open-ended manner. This facilitates the development of higher-order skills such as stylistic awareness, temporal structuring, and interactive listening, reinforcing the idea of AI as a mediator of musical understanding rather than a mere generative tool.
This perspective aligns with various theoretical postulates that conceive generative systems as epistemic instruments, constituted through interfaces that allow artistic reasoning to be materialized and tested; that is, transforming creative ideas and decisions into tangible sound elements that can be evaluated, explored, and modified iteratively, rather than being limited to autonomous generators oriented toward randomness or the mere reproduction of paradigmatic structures within a style (
Nika et al., 2025). From this perspective, interaction with AI systems should not be understood as a process of selecting outputs, but as a process of progressive formalization of musical thought, in which the student externalizes hypotheses, contrasts alternatives, and refines criteria through iterative feedback loops. In this sense, AI environments function as spaces for artistic reasoning in action, where the relationship between intention, action, and result becomes explicit and open to reflection. This shift is particularly relevant in educational contexts, as it enables students to transition from intuitive or reactive forms of creativity toward more conscious, structured, and critically grounded forms of musical decision-making (
Nika et al., 2025). In the educational sphere, this shift does not imply that creativity should dispense with intuition; rather, it suggests that intuition must be translated into musical hypotheses that are open to exploration and critique. Current and emerging AI developments represent a paradigm shift in this regard, empowering students to probe the structural logic of their work: Which melodic variants preserve a phrase’s identity while altering its character? How is the perception of phrasing transformed by modal exchanges or subtle modulations? In this framework, AI provides the generated examples and counterarguments necessary to justify creative choices, transforming creativity from the mere selection of preconfigured options into a rigorous practice of reflective musical reasoning. In addition,
Nika et al. (
2025) emphasize the role of interactive generation as a paradigm in which musical agents are not conceived as autonomous creators, but as responsive systems that evolve through continuous interaction with the user. This approach reframes AI not as a producer of finished musical artifacts, but as a dynamic partner in the co-construction of musical discourse, capable of adapting to the user’s inputs, constraints, and evolving intentions. Such interaction allows for the emergence of situated musical knowledge, where understanding develops through action, negotiation, and real-time feedback. In pedagogical terms, this opens up the possibility of designing learning environments in which students do not merely consume or evaluate generated material, but actively shape and interrogate it, fostering a deeper integration between perception, action, and reflection in musical learning processes.
In the specific field of composition,
Kardos (
2012) demonstrates how technology unveils new ‘sound worlds’, fostering student autonomy and creative confidence. By externalizing structural relationships, these tools allow students to witness in real time how a melody interacts with underlying harmony, or how a variation maintains its thematic essence despite shifts in rhythm, texture, or register.
6. Conclusions
This paper highlights that the integration of artificial intelligence into music education reaches its maximum transformative potential when it is part of an ecological framework, as a semiotic and cognitive agent that transcends its instrumental use. This perspective is consistent with recent approaches that conceptualize AI not as an isolated tool but as part of a broader pedagogical and relational framework that integrates technological mediation with reflective and meaning-oriented learning processes (
Yoo, 2026).
Guo et al. (
2026) observe that AI’s impact spans all areas of music education and hence has a truly transformative pedagogical potential.
J. F. Merchán Sánchez-Jara et al. (
2024) similarly found that AI is paving the way toward a more personalized, interactive and efficient learning experience. Far from conceiving AI as a mere technical resource aimed at automating tasks or correcting errors, its potential as a mediating agent capable of making visible structural, narrative, and expressive aspects of music that have traditionally belonged to musicians highly competent in notational or technical aspects has been defended. Correspondingly, the central hypothesis proposes that AI can function as a cognitive artifact (
Cassinadri, 2024;
Norman, 1991;
Rivera-Novoa & Duarte Arias, 2025), redistributing the cognitive load and facilitating access to higher-order processes such as comparing, inferring, anticipating, or justifying musical decisions.
From a theoretical point of view, this proposal is based on a sociocultural conception of mediation, where musical learning is understood as the construction of meaning through symbolic tools as pedagogical resources that operate in relation to the students’ zone of proximal development. When AI acts as adaptive scaffolding rather than as a corrector or automatic and decontextualized generator, it favors progression from guided performances to more autonomous forms of interpretation and creation, reconfiguring both creative decision-making and forms of representation and reception processes.
The work has also emphasized that creativity is not an afterthought to understanding, but rather an inherent dimension of the very process of constructing meaning. From this perspective, creating involves making conscious decisions within systems of cultural and stylistic constraints, in non-linear trajectories marked by exploration, contrast, and revision. As
Mazlan et al. (
2026) note, such AI tools enhance practice efficiency, personalize instruction and improve assessment objectivity, and they flag creativity and feedback as key pedagogical impacts. In this framework, technological mediation can enhance metacognitive reflection as long as it promotes open exploration and is not reduced to closed templates that limit autonomy. Likewise, the nexus between musical comprehension, creativity, and cognitive development reinforces the educational relevance of this integration process. Musical practice is associated with executive functions in the field of brain plasticity development, demonstrating how technological mediation can amplify executive and metacognitive processes widely recognized as central to the artistic development of students. Recent studies further indicate that interaction with AI-supported musical environments is associated with increased cognitive engagement, creative reasoning, and metacognitive awareness, reinforcing its potential to support higher-order executive processes in artistic learning (
Ma et al., 2025). However, taking advantage of these affordances requires conscious and coherent pedagogical articulation: the effectiveness of AI depends on its symbiotic and indispensable integration as a constituent part of curriculum design under criteria defined by the teacher and aligned with educational objectives. In this context,
Gisbert-Caudeli (
2025) similarly concludes that ‘a critical and humanistic integration of AI can enrich the educational process facilitating to a large extent the adaptation and personalization of the learning process. Finally, the work highlights that any AI-mediated educational ecosystem also constitutes a space for cultural selection. Technology can reinforce monolithic canons or, conversely, open up spaces for more pluralistic meanings, depending on the repertoires and epistemological frameworks taken into consideration. Therefore, critical and contextualized integration is essential to avoid formalistic reductionism and preserve the aesthetic, cultural, and semantic dimension of the musical experience. In short, AI can contribute to reconfiguring music education as a shared horizon of understanding, creativity, and cognitive development, provided that its mediation is articulated on the basis of coherent and culturally committed pedagogical criteria.
Future research should move toward the systematic empirical validation of the ecological and mediational framework advanced in this study, conceptualizing AI-mediated environments as complex sociotechnical systems that reorganize the conditions of musical cognition, meaning-making, and creative agency. From this perspective, design-based and longitudinal methodologies are particularly suited to examining how AI, as a cognitive artifact, dynamically redistributes cognitive load and reshapes the development of higher-order processes such as reflective judgment, anticipatory reasoning, and metacognitive regulation over time (
Rivera-Novoa & Duarte Arias, 2025). Special attention should be paid to the analysis of multimodal interaction processes, investigating how learners engage with heterogeneous representations (visual, auditory, symbolic) and how these interactions contribute to the construction of musical meaning within expanded learning ecologies. In parallel, further empirical work is needed to elucidate the pedagogical conditions under which AI functions as adaptive scaffolding rather than as a decontextualized generator, particularly in relation to teacher orchestration and curricular design. Experimental and quasi-experimental studies may provide robust evidence regarding the impact of AI-mediated environments on executive functions, creative reasoning, and the development of situated musical thinking (
Ma et al., 2025). Moreover, future research should critically address the sociocultural and epistemological dimensions of these ecosystems, examining how algorithmic mediation participates in processes of cultural selection, inclusion, and exclusion. Finally, advancing toward explainable, human-centered, and culturally responsive AI systems constitutes a necessary condition to ensure that technological mediation remains aligned with principles of agency, interpretability, and meaningful learning, thus consolidating AI not as an external tool but as an integral component of educational ecologies.