Next Article in Journal
A Preliminary Data-Driven Competency Mapping Study for Modular Construction Designers: Exploratory Korean Validation Using Bayesian BWM and Fuzzy DEMATEL
Next Article in Special Issue
Toward Sustainable Digital Education in Biology: Evaluating Educators’ Perceptions and Adoption Intentions for a Virtual Laboratory Toolkit from Four European Contexts
Previous Article in Journal
Coupled Thermal Desorption–Thermal Plasma Methods for Diesel-Contaminated Soil Remediation and Syngas Production
Previous Article in Special Issue
Offline Technology for Rural AI Literacy: Steps Towards a Holistic Educational Solution
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

How Pre-Service Elementary Teachers Develop Scientific Concepts in AI-Integrated Lesson Designs: Implications for Sustainable Teacher Education

Department of Gifted Education, Seoul National University of Education, 96, Seochojungang-ro, Seocho-gu, Seoul 06639, Republic of Korea
Sustainability 2026, 18(10), 5211; https://doi.org/10.3390/su18105211
Submission received: 29 April 2026 / Revised: 9 May 2026 / Accepted: 20 May 2026 / Published: 21 May 2026
(This article belongs to the Special Issue Sustainable Digital Education: Innovations in Teaching and Learning)

Abstract

As AI and digital tools become more widely adopted in school education, integrating them sustainably into teacher preparation has become a central concern for sustainable teacher education. This study examined how pre-service elementary teachers develop scientific concepts within AI-integrated lesson plans and how those patterns change within each case following teaching demonstrations and instructor feedback. Qualitative content analysis was conducted on twelve lesson plans—initial drafts and revised versions from six groups across two science units—produced within an elementary science methods course. Plans were analyzed along three dimensions of conceptual development (conceptual structuring, generalization, and conceptual explicitness) and three functional roles of AI and digital tools. In draft plans, tools were predominantly used for learner engagement and artifact production, with scientific concepts embedded in activity contexts. Following feedback, conceptual explicitness was the dimension most frequently revised, while changes in conceptual structuring and generalization appeared in fewer cases. Cases in which conceptual development reached higher levels in revised plans shared a common design feature: AI outputs were repositioned within the consolidation stage in connection with explicit concept statements, rather than serving as content presentation. These findings suggest that pedagogical judgment about positioning AI outputs within lesson stages, reflected across design–demonstration–feedback–revision cycles, is central to the quality of AI-integrated science lesson design and offers implications for sustaining teacher preparation in the era of AI.

1. Introduction

The rapid proliferation of generative AI tools has substantially altered how teachers design instruction, develop materials, and provide feedback to learners [1,2]. In science education, AI and digital tools offer particular affordances for supporting inquiry-based learning, enabling real-time formative assessment, and mediating between observation and conceptual explanation [3]. These developments have intensified interest in how pre-service teachers learn to integrate AI and digital tools into lesson design, as teacher education programs are increasingly expected to cultivate not only technical familiarity with such tools but also the pedagogical judgment needed to deploy them effectively [4,5]. The capacity of pre-service teachers to integrate AI tools in pedagogically sound ways is increasingly understood as a foundation for sustainable teacher education, since the quality of teacher preparation in this domain shapes how AI is taken up in classrooms over time [6].
Existing research on AI integration in pre-service teacher education has primarily focused on attitudes, perceived efficacy, and general patterns of tool use [2,5,7,8]. These studies indicate that pre-service teachers tend to adopt AI tools for motivational or presentational purposes rather than for deepening conceptual learning. Recent work has begun to examine how AI tools function in science lesson planning specifically, drawing on pedagogical content knowledge frameworks to interpret teachers’ integration choices [9]. However, less attention has been paid to how pre-service teachers organize and develop scientific concepts within AI-integrated lesson designs—specifically, whether AI tools are positioned to support conceptual structuring, generalization, and explicit concept formulation, or whether they remain peripheral to the core instructional logic. A related gap concerns how these design patterns change following teaching demonstrations and instructor feedback, a process through which the way AI tools are positioned within lesson designs may shift in relation to content-specific learning goals.
This study addresses these gaps by examining how pre-service elementary teachers develop scientific concepts within AI-integrated lesson plans and how those design patterns change within each case following teaching demonstrations and instructor feedback. Qualitative content analysis was conducted on twelve lesson plans—initial drafts and revised versions produced by six groups across two science units within the same elementary science methods course. The analysis focused on three dimensions of conceptual development—conceptual structuring, generalization, and conceptual explicitness—and on the functional roles of AI and digital tools in relation to those dimensions. By making patterns of conceptual development within lesson plans analytically visible, this study aims to offer implications for pre-service teacher education curricula that connect subject matter knowledge with the pedagogical use of AI tools, contributing to discussions on how teacher preparation can be sustained in pedagogically meaningful ways alongside the rapid integration of AI in education.

2. Materials and Methods

2.1. Research Design

This study employed qualitative content analysis to examine lesson plans produced by pre-service elementary teachers. Qualitative content analysis enables systematic categorization and interpretation of textual meaning within a defined analytical framework and is particularly suited to analyzing instructional design artifacts, as it allows conceptual patterns to be identified and compared across cases [10]. The analytical focus was not on actual classroom implementation or student learning outcomes, but on the ways scientific concepts were organized and developed within the lesson plans themselves. Within this focus, the primary unit of analysis was within-case change between draft and revised plans for each group, rather than cross-case comparison of absolute coding levels, given that the structural characteristics of each unit’s learning activities may shape the starting point of conceptual coding. This approach is grounded in the view that lesson plans function as artifacts that indirectly reveal pre-service teachers’ pedagogical content knowledge and instructional reasoning [11,12]. The English-language manuscript was prepared with the assistance of a generative AI tool (Claude, Anthropic, San Francisco, CA, USA) for language editing. All analytical decisions, coding, interpretations, and conclusions were made by the author.

2.2. Research Context and Data Sources

The data were drawn from an elementary science methods course offered at a university of education in Republic of Korea during the second semester of the 2025 academic year. The course centered on science lesson design and teaching demonstrations and was organized around units from the elementary science curriculum. Six groups of third-year pre-service elementary teachers participated, with three to four members per group. Each group selected one science unit from the matter domain of the 2022 revised national curriculum and designed an AI-integrated lesson plan for a specific class session within that unit.
Each group first designed an initial draft lesson plan, then conducted a teaching demonstration, received instructor feedback, and submitted a revised version. Following each teaching demonstration, the instructor provided oral feedback to the group on the entire lesson plan. Specifically, instructor feedback addressed three aspects of each lesson plan. First, the pedagogical role of AI and digital tools was examined: whether each tool was assigned a clear function within the lesson—such as visualizing scientific phenomena, deepening conceptual understanding through interactive tasks, or providing individualized support—rather than being included without explicit instructional purpose. Second, the appropriateness of explanations for the target learners was reviewed: whether the language, analogies, and level of abstraction used in AI-generated outputs and teacher-planned explanations were suitable for the developmental level of elementary students. Third, the alignment of scientific concepts with curriculum standards was evaluated: whether the concepts presented through the lesson activities and AI outputs were consistent with the achievement standards specified in the national curriculum for the relevant unit and grade level. Feedback sessions lasted approximately ten minutes per group and were delivered by the same instructor across all six groups to ensure consistency in evaluative criteria. Each group revised their lesson plan based on this feedback and submitted a final version. This design–demonstration–feedback–revision cycle provided the basis for comparing draft and revised plans within each group.
The twelve lesson plans analyzed in this study—one draft and one revised plan per group—each contained lesson objectives, activity sequences, teacher questioning strategies, student tasks, consolidation stage concept statements, and descriptions of AI and digital tool use.
The two units selected for analysis are part of the matter domain in the elementary science curriculum of Republic of Korea [13]. The Matter domain encompasses both observable phenomena (e.g., color changes in indicator solutions) and processes requiring connection to underlying explanations (e.g., changes in the states of water), making it a productive context for examining how AI tools are positioned in relation to different types of conceptual development [14]. Research has shown that children’s conceptual understanding of matter develops progressively and involves multifaceted patterns across grade levels [14,15], further supporting the selection of this domain for analyzing how lesson designs scaffold conceptual development. Changes in the State of Water is taught in Grade 4 and centers on connecting observable changes in water’s physical states with explanations of underlying processes. Prior research has shown that students’ reasoning about water state changes develops progressively across grade levels in elementary school [15], making this unit a productive context for analyzing how lesson designs scaffold conceptual development. Acids and Bases is taught in Grade 6 and is organized around the classification of solutions using indicators, alongside experimental observation of solution properties. These two units were selected because they share the same content domain while differing in the structural characteristics of their learning activities, allowing analysis of AI-integrated lesson designs across activity types within the elementary science curriculum.
For clarity, cases are labeled accordingly—SW-A, SW-B, and SW-C for the three groups addressing the Changes in the State of Water (hereafter SW) unit and AB-A, AB-B, and AB-C for the three groups addressing the Acids and Bases (hereafter AB) unit.

2.3. Analytical Framework

2.3.1. Conceptual Development Categories

Three categories were used to analyze how scientific concepts were developed across lesson plan stages: conceptual structuring (CS), generalization (GEN), and conceptual explicitness (EXP). These categories were grounded in concept-based approaches to curriculum design, which argue that meaningful learning is organized hierarchically—moving from facts and activities toward relational organization of concepts and the formulation of generalizations that articulate conceptual relationships under specified conditions [16,17]. The framework also draws on the broader view that explicit articulation of concepts and their relationships is central to scientific concept formation [18].
CS was coded at three levels: CS1 (activity-centered, with no explicit relational structuring of concepts), CS2 (comparative structuring using criteria such as similarities, differences, or conditions), and CS3 (definition- and generalization-centered, with condition-inclusive rules explicitly presented). GEN was coded as GEN0 (case-level only), GEN1 (general statements present but conditions unspecified), or GEN2 (condition-inclusive generalization in the form “when [condition], [outcome]”). EXP was coded as EXP0 (key terms absent or unclear), EXP1 (terms present but meaning not structured), or EXP2 (terms and meanings formulated in explicit sentences or through structured summary activities). Table 1 summarizes the analytical categories.

2.3.2. AI and Digital Tool Function Categories

AI and digital tool use was analyzed not in terms of frequency or tool type, but according to the pedagogical function each tool was designed to perform within the lesson. Three functional categories were applied: concept presentation (AI-PRES)—presenting concepts or phenomena visually or verbally; concept relational structuring (AI-STR)—supporting comparison or relational organization of concepts; and concept refinement and generalization (AI-REF)—supporting concept refinement, elaboration, or generalization [3,4]. Table 2 summarizes the functional categories.

2.4. Analytical Procedure

Analysis followed the general procedures of qualitative content analysis, including iterative reading, coding, category refinement, and cross-case comparison [10]. Each lesson plan was first segmented into activity-stage units (introduction, development, consolidation, application), then coded using the CS, GEN, EXP, and AI function categories. Cases where category application was ambiguous were flagged and reviewed against category definitions, with boundary cases adjudicated in consultation with an expert in elementary science education. To check coding consistency, four of the twelve lesson plans (one draft–revised set per unit, approximately 33% of the total corpus) were selected, and all activity-stage units within these plans were independently coded by the researcher and an expert in elementary science education. Cohen’s kappa coefficient was 0.748, indicating acceptable intercoder reliability. Disagreements were resolved through discussion, and the agreed-upon criteria were applied to the full dataset. Final codes were compared across draft and revised versions within each group to identify patterns of change.

3. Results

The following sections present the coding results for the twelve lesson plans across the three analytical dimensions of conceptual development and the three functional categories of AI and digital tool use. As noted in the methods Section 2, the primary unit of analysis was within-case change between draft and revised plans for each group, rather than cross-case comparison of absolute coding levels.

3.1. Conceptual Development Patterns in Draft Lesson Plans

Analysis of the six draft lesson plans revealed variation in conceptual development across cases. Table 3 summarizes the coding results for each case.
In terms of conceptual structuring, four of the six draft plans were coded at CS1, indicating that science concepts were embedded within activity contexts without explicit relational organization. The remaining two cases (AB-B and AB-C) were coded at CS2, reflecting the presence of comparative criteria in the lesson structure. No draft plan reached CS3.
The two CS2 cases shared a structural characteristic: the classification task itself required students to apply criteria for distinguishing acidic from basic solutions, which meant comparative organization was inherent to the activity design. Among the cases organized around other types of learning activities, only SW-B included a comparative inquiry structure between evaporation and boiling; the remaining cases organized the lesson around artifact production (SW-A), a boiling experiment with expressive output (SW-C), or a solution-mixing observation with quiz-based consolidation (AB-A) without stages for relational concept structuring.
Generalization levels followed a related pattern. Cases coded at CS1 were also coded at GEN0–1, indicating that concept statements did not extend beyond observation-level descriptions or partial generalizations without specified conditions. Among CS2 cases, AB-B and AB-C reached GEN1–2, reflecting the presence of partial condition-inclusive language in board-writing plans. SW-B reached GEN1.
Conceptual explicitness also varied across the cases. The three Changes in the State of Water drafts (SW-A, SW-B, SW-C) and AB-A were coded at EXP1—key terms appeared but were not formulated in explicit definitional sentences. AB-B and AB-C were coded at EXP2, as their board-writing plans contained structured classification statements linking indicator color changes to solution categories.
Regarding AI and digital tool function, all six draft plans incorporated concept presentation (AI-PRES) as the primary functional role—tools were used to present phenomena, generate visual materials, or support artifact production. Two cases additionally demonstrated concept relational structuring (AI-STR) through real-time result-sharing platforms that facilitated inter-group comparison (AB-B: Google Slides; AB-C: Padlet). One case (AB-C) showed partial concept refinement and generalization (AI-REF) through the use of Claude as a prediction tool prior to experimentation.

3.2. Changes in Conceptual Development Patterns in Revised Lesson Plans

Following teaching demonstrations and instructor feedback, all six groups submitted revised lesson plans. Changes in conceptual development across the six cases were uneven. Table 4 presents a comparison of draft and revised coding results for each case.
Conceptual explicitness (EXP) showed change in the largest number of cases. Among the four cases coded at EXP1 in the draft stage, three (SW-A, SW-B, and SW-C) reached EXP2 in their revised plans. The remaining case at EXP1, AB-A, showed no change. The two cases already at EXP2 in the draft stage (AB-B and AB-C) maintained that level.
Changes in conceptual structuring (CS) were observed in fewer cases. SW-A moved from CS1 to CS2 through the introduction of a prompt template that required students to articulate state-change principles as part of an artifact production task. SW-C showed the largest shift, moving from CS1 to CS3, as the revised plan explicitly structured the consolidation stage around condition-inclusive statements and a concept summary table. AB-C moved from CS2 to CS2–CS3 through the introduction of a machine learning classification task, in which students trained a model to distinguish acidic from basic solutions based on indicator color patterns. The remaining three cases (SW-B, AB-A, AB-B) showed no change in CS level between draft and revised plans.
Generalization level (GEN) changed in two cases. SW-C moved from GEN1 to GEN2, as the revised board-writing plan included condition-containing statements specifying the locations and conditions under which evaporation and boiling occur. AB-C moved from GEN1–2 to GEN2, as the classification model-building activity required students to make the criteria for distinguishing solution types explicit. The remaining four cases showed no change in GEN level.
AB-A showed no substantive change across all three dimensions. The primary modification in this case was a reallocation of instructional time between activities, without alteration to the conceptual structure, consolidation design, or AI tool use.

3.3. Changes in AI and Digital Tool Function

The AI and digital tools used across the twelve lesson plans included generative AI chatbots (Claude, Anthropic, San Francisco, CA, USA; Gemini, Google, Mountain View, CA, USA), a Korean generative AI platform (Wrtn, Wrtn Technologies, Seoul, Republic of Korea), a Korean educational AI chatbot (OurChild AI, Ingradient Inc., Seoul, Republic of Korea), simulation platforms (PhET, University of Colorado, Boulder, CO, USA; JavaLab, an independently developed science simulation website, Republic of Korea; Vivasam, Chunjae Education, Seoul, Republic of Korea), a metaverse-based learning platform (ZEP, ZEP Inc., Seoul, Republic of Korea), a real-time quiz and response analysis platform (Sim Ground, V_LAB Inc., Seoul, Republic of Korea), a web-based machine learning platform (Teachable Machine, Google, Mountain View, CA, USA), collaborative tools (Padlet, Padlet Inc., San Francisco, CA, USA; Google Slides, Google, Mountain View, CA, USA; Canva, Canva Pty Ltd, Sydney, Australia), and an AI music generation tool (Suno AI, Suno Inc., Cambridge, MA, USA). Table 5 summarizes the AI and digital tool use across draft and revised plans for all six cases.
In draft plans, concept presentation (AI-PRES) was the predominant functional role across all six cases. Tools were used primarily to present phenomena visually, generate images, or support artifact production. Two cases additionally demonstrated AI-STR function through platforms facilitating inter-group result sharing and comparison (AB-B: Google Slides; AB-C: Padlet). One case showed partial AI-REF function through the use of Claude as a prediction tool prior to experimentation (AB-C).
In revised plans, the functional role of AI tools shifted in four of the six cases, though the nature of the shift differed across cases.
In SW-A, the primary tool structure was retained, but a structured prompt template was introduced for Gemini, requiring students to embed state-change terminology in their prompts. This added a partial AI-REF function to the existing AI-PRES role. In SW-B, Claude was added to generate visualizations of microscopic particle movement during evaporation and boiling. These visualizations were directly integrated into comparative tasks in the consolidation stage, extending tool function from AI-PRES + AI-STR to AI-PRES + AI-STR + AI-REF. In SW-C, the consolidation stage was restructured around Sim Ground, a platform used to collect real-time student response data; incorrect responses were analyzed to address misconceptions. This repositioned tool use from an expression-centered function toward AI-REF + AI-STR in the consolidation stage. In AB-C, Claude was replaced by Teachable Machine, a web-based machine learning platform. Students used photographs of experimental results to train a classification model, connecting the sorting of solutions into acidic and basic categories with the process of pattern recognition. This shift moved tool function from AI-PRES + partial AI-STR to AI-STR + AI-REF. AB-A and AB-B showed no substantive change in AI tool function between the draft and revised plans.

4. Discussion

This study examined how pre-service elementary teachers develop scientific concepts within AI-integrated lesson plans, with the analytical focus placed on within-case change between draft and revised plans rather than on cross-case comparisons of absolute coding levels. The findings reveal several patterns that warrant interpretation in light of existing research on technology integration in teacher education.
Among the three dimensions of conceptual development, conceptual explicitness (EXP) showed change in the largest number of cases. Of the four cases coded at EXP1 in the draft stage, three reached EXP2 in their revised plans. In contrast, changes in conceptual structuring (CS) and generalization (GEN) appeared in fewer cases. This distribution suggests that adding explicit concept formulation to the consolidation stage was a more frequently adopted design choice when revising lesson plans following feedback. This observation is restricted to the present six cases and does not support broader claims about hierarchical relationships among the three dimensions; nonetheless, it indicates that explicit sentence-level concept formulation may function as a more accessible entry point for design revision than restructuring conceptual relations or generalizing across cases.
The two cases that achieved CS3 or GEN2 in their revised plans (SW-C and AB-C) shared a common design feature: both repositioned AI tool outputs within the consolidation stage in connection with structured concept statements. SW-C introduced Sim Ground for real-time response data collection and combined this with a concept summary table containing condition-inclusive statements. AB-C replaced its prediction tool with Teachable Machine and structured the consolidation activity around interpreting a student-trained classification model. While both revisions involved tool changes, the more consequential design decision was the placement of tool outputs at the consolidation stage in ways that required students to articulate or apply explicit conceptual criteria. This finding aligns with Cuban’s argument [19] that the educational value of technology depends not on the tool itself but on its integration into instructional purpose and extends this argument by showing that tool placement within specific lesson stages is a key locus of pedagogical decision-making in AI-integrated lesson design.
A common feature across the advanced strategies that emerged in the revised plans is worth noting. The prompt template in SW-A, the microscopic visualization linked to comparative tasks in SW-B, the real-time misconception correction in SW-C, and the machine learning classification model in AB-C differ in surface-level form. However, all four repositioned AI tools so that their outputs were directly connected to students’ cognitive products—what students input as prompts, what they compared visually, what misconceptions they exhibited, or what classification criteria they encoded into a model. This pattern suggests a shift in how AI was conceived: from a tool that presents content for students to view toward a tool that processes or responds to students’ own thinking. This conceptual shift reflects the integration of technology with content-specific pedagogical reasoning that the TPACK framework identifies as central to effective technology integration [4]. It also resonates with recent calls for artifact-based assessment of AI-integrated teaching that focuses on how teachers position AI tools within instructional designs rather than on their general familiarity with tools [20,21].
The patterns identified above are inseparable from a methodological observation: the type of learning activity central to each case appeared to influence the starting point of conceptual coding. The two cases organized around classification activities (AB-B and AB-C) were coded at CS2 and EXP2 in their draft stage, while cases organized around observation, inference, or artifact production showed lower starting points. This is consistent with the structural characteristic of classification tasks, in which comparative criteria are inherent to the activity itself. Accordingly, the central object of analysis in this study is not the relative level of coding across cases but the trajectory of change within each case. Case AB-A is notable for showing no substantive change across any of the three dimensions between the draft and revised plans. The primary modification in this case was a reallocation of instructional time between activities, without restructuring the conceptual logic or the role of AI tools. Several features of the draft plan may have contributed to this outcome. The initial lesson was organized around a mixing experiment followed by a digital quiz (Kahoot) in the consolidation stage, a structure that did not inherently require comparative criteria or condition-inclusive concept statements. Unlike the classification-oriented cases (AB-B, AB-C), where comparative structuring was embedded in the activity itself, AB-A’s activity structure may have offered fewer entry points for the type of conceptual restructuring that feedback was intended to prompt. Additionally, the consolidation stage relied entirely on a quiz format, which assessed recall rather than requiring written concept formulation. However, these observations remain tentative; the present data do not allow inference about group-level factors such as engagement with feedback, collaborative dynamics, or prior knowledge that may have also influenced the outcome. These observations point to a broader implication: the development of pre-service teachers’ AI-integrated lesson design competence cannot be assessed independently of the structural characteristics of the science learning activities they encounter and may benefit from continued iteration across diverse activity types.
Taken together, these patterns carry implications for sustainable teacher education in the context of AI integration. The findings suggest that the educational value of AI tools is not stabilized through frequency of use or familiarity with specific tools but through pedagogical decisions about where AI outputs are positioned within lesson stages and how they are connected to students’ conceptual reasoning. The present study does not claim that six cases are sufficient to establish a generalizable pathway for sustainable teacher education. Rather, the analytical framework and the within-case trajectories observed here illustrate how the quality of such pedagogical design choices can be made visible and examined within a structured preparation cycle. It is this capacity for structured visibility—enabling teacher educators to identify where and how AI outputs are positioned in relation to conceptual development—that constitutes the study’s contribution to discussions of sustainable teacher education.
This study has several limitations that delimit the scope of its conclusions. Given the small number of cases—six groups within a single course at one institution in Republic of Korea the findings are intended as exploratory observations rather than generalizable claims, and the sustainability implications discussed in this study should be understood as directions for further investigation rather than as established conclusions. The two units examined—organized primarily around observation-inference and classification activities—represent a limited range of the science learning activity types available in the elementary curriculum. Activity types such as modeling, controlled experimentation, and investigation-based inquiry were not included and may yield different patterns of AI tool integration and conceptual development. The analytical framework focused on three dimensions of conceptual development and three functional roles of AI tools, but real-time formative assessment functions emerging in some revised plans were not fully captured by these categories. Finally, the analysis was conducted on lesson plan artifacts rather than on enacted classroom teaching, which limits inferences about how design choices translate into student learning.

5. Conclusions

This study contributes to research on AI integration in pre-service teacher education by analyzing instructional design artifacts rather than self-reported attitudes or perceived efficacy. The analytical framework distinguishing three dimensions of conceptual development (conceptual structuring, generalization, and conceptual explicitness) and three functional roles of AI and digital tools (concept presentation, relational structuring, and refinement and generalization) made patterns of conceptual development within lesson plans analytically visible and supported within-case comparison of how those patterns changed following teaching demonstrations and instructor feedback. By placing the unit of analysis on within-case change rather than on cross-case comparison of absolute coding levels, the study offers a methodological orientation that accounts for variation introduced by different science learning activity types.
Two implications follow for pre-service elementary teacher education. First, beyond familiarity with AI tools, pre-service teachers need to develop pedagogical judgment about where to position AI outputs within lesson stages and how to connect those outputs to students’ conceptual reasoning. The cases in which conceptual development reached higher levels in revised plans were characterized not by the choice of advanced tools, but by the placement of tool outputs within the consolidation stage in connection with explicit concept statements. Second, pre-service teachers need opportunities to design AI-integrated lessons across diverse types of science learning activities, since structural features of activity types may shape how scientific concepts can be developed in lesson designs. Embedding design–demonstration–feedback–revision cycles within teacher education programs offers a structural context within which pedagogical design choices can be made visible and revised.
These implications connect to broader discussions of sustainable teacher education in the era of AI. The findings suggest that the educational value of AI tools in science lesson design is not stabilized through tool acquisition alone but through pedagogical judgment that can be examined, articulated, and revised across iterations. From this perspective, sustaining the quality of teacher preparation alongside the rapid evolution of AI tools depends less on training pre-service teachers in particular technologies and more on building structural opportunities for them to engage in revisable, content-aware design practice [22].
As noted in the discussion, this study is based on a small number of cases within a specific institutional and curricular context, and the analysis was conducted on lesson plan artifacts rather than enacted classroom teaching. Future research could extend this line of inquiry by analyzing AI-integrated lesson design across a broader range of science learning activity types, by tracing changes in design competence over longer periods, and by connecting lesson design analysis with observation of enacted classroom teaching.

Funding

This research received no external funding.

Institutional Review Board Statement

Ethical review and approval were waived for this study by Institution Committee due to Legal Regulations (This study qualifies for IRB exemption under Korean national legislation. According to the Bioethics and Safety Act of the Republic of Korea (Article 15, Paragraph 2) and its Enforcement Rule (Article 13), research that does not collect or record personally identifiable information from human subjects, and analyzes existing instructional artifacts produced as part of regular university coursework, is exempt from formal IRB review).

Informed Consent Statement

Verbal informed consent was obtained from the participants. Verbal consent was obtained rather than written because the study posed minimal risk and did not involve sensitive personal data.

Data Availability Statement

The data presented in this study are available on request from the corresponding author due to privacy and confidentiality considerations regarding the participating pre-service teachers’ coursework.

Acknowledgments

During the preparation of this manuscript, the author used Claude (Anthropic) for the purposes of improving English phrasing and clarity. The author has reviewed and edited the output and takes full responsibility for the content of this publication.

Conflicts of Interest

The author declares no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AIArtificial Intelligence
TPACKTechnological Pedagogical Content Knowledge
PCKPedagogical Content Knowledge
CSConceptual Structuring
GENGeneralization
EXPConceptual Explicitness
AI-PRESConcept Presentation
AI-STRConcept Relational Structuring
AI-REFConcept Refinement and Generalization
SWChanges in the State of Water (unit)
ABAcids and Bases (unit)

References

  1. Kasneci, E.; Sessler, K.; Küchemann, S.; Bannert, M.; Dementieva, D.; Fischer, F.; Gasser, U.; Groh, G.; Günnemann, S.; Hüllermeier, E.; et al. ChatGPT for good? On opportunities and challenges of large language models for education. Learn. Individ. Differ. 2023, 103, 102274. [Google Scholar] [CrossRef]
  2. Zhang, X.; Zhang, P.; Shen, Y.; Liu, M.; Wang, Q.; Gašević, D.; Fan, Y. A systematic literature review of empirical research on applying generative artificial intelligence in education. Front. Digit. Educ. 2024, 1, 223–245. [Google Scholar] [CrossRef]
  3. Ouyang, F.; Jiao, P. Artificial intelligence in education: The three paradigms. Comput. Educ. Artif. Intell. 2021, 2, 100020. [Google Scholar] [CrossRef]
  4. Mishra, P.; Koehler, M.J. Technological pedagogical content knowledge: A framework for teacher knowledge. Teach. Coll. Rec. 2006, 108, 1017–1054. [Google Scholar] [CrossRef]
  5. Yorulmaz, A.; Okulu, H.Z.; Muslu-Komurcu, N.; Cokcaliskan, H. Enhancing STEM lesson plans: Pre-service primary teachers’ collaboration with ChatGPT. Educ. Inf. Technol. 2025, 30, 24543–24573. [Google Scholar] [CrossRef]
  6. Chiu, T.K.F.; Chai, C.S. Sustainable curriculum planning for artificial intelligence education: A self-determination theory perspective. Sustainability 2020, 12, 5568. [Google Scholar] [CrossRef]
  7. Lim, C.; Chee, H.-K.; Go, B.; Lim, E.; Min, S. A study on the development of education programs to strengthen AI·digital utilization instructional design competencies of pre-service teachers. J. Korean Assoc. Educ. Inf. Media 2024, 30, 129–153. [Google Scholar] [CrossRef]
  8. Zawacki-Richter, O.; Marín, V.I.; Bond, M.; Gouverneur, F. Systematic review of research on artificial intelligence applications in higher education—Where are the educators? Int. J. Educ. Technol. High. Educ. 2019, 16, 39. [Google Scholar] [CrossRef]
  9. Peikos, G.; Stavrou, D. ChatGPT for science lesson planning: An exploratory study based on pedagogical content knowledge. Educ. Sci. 2025, 15, 338. [Google Scholar] [CrossRef]
  10. Graneheim, U.H.; Lindgren, B.M.; Lundman, B. Methodological challenges in qualitative content analysis: A discussion paper. Nurse Educ. Today 2017, 56, 29–34. [Google Scholar] [CrossRef] [PubMed]
  11. Grossman, P.L. The Making of a Teacher: Teacher Knowledge and Teacher Education; Teachers College Press: New York, NY, USA, 1990. [Google Scholar]
  12. Song, N.; Lee, M.; Noh, T. Characteristics of pedagogical design of pre-service elementary teachers using science teacher’s guides. J. Korean Elem. Sci. Educ. 2024, 43, 504–518. [Google Scholar] [CrossRef]
  13. Ministry of Education of the Republic of Korea. 2022 Revised National Curriculum: Science; Ministry of Education: Sejong, Republic of Korea, 2022.
  14. Liu, X.; Lesniak, K. Progression in children’s understanding of the matter concept from elementary to high school. J. Res. Sci. Teach. 2006, 43, 320–347. [Google Scholar] [CrossRef]
  15. Jung, J.; Chang, J.; Park, J. An Analysis of Different Grade Levels of Elementary School Students’ Reasoning about the Changes of State of Water within a Learning Progression. Asia-Pac. Sci. Educ. 2021, 6, 548–563. [Google Scholar] [CrossRef]
  16. Erickson, H.L. Concept-Based Curriculum and Instruction for the Thinking Classroom; Corwin Press: Thousand Oaks, CA, USA, 2007. [Google Scholar]
  17. Han, J.; Lim, Y.; Ryu, H.; Kim, S.; Wee, S. Development of a concept-based unit design framework for deep learning. J. Curric. Stud. 2025, 43, 25–54. [Google Scholar] [CrossRef]
  18. Novak, J.D. Learning, Creating, and Using Knowledge: Concept Maps as Facilitative Tools in Schools and Corporations; Routledge: New York, NY, USA, 2010. [Google Scholar]
  19. Cuban, L. Oversold and Underused: Computers in the Classroom; Harvard University Press: Cambridge, MA, USA, 2001. [Google Scholar]
  20. Aldemir, T.; Kilinc, S.; Bicer, A.; Grant, P.; Davis, T.; Sweany, N.W. Intelligent-TPACK in practice: Design and evidence from a three-week teacher preparation module. Comput. Educ. Open 2025, 9, 100254. [Google Scholar] [CrossRef]
  21. Mourlam, D.J.; Chesnut, S.R.; Bleecker, H. Exploring preservice teacher self-reported and enacted TPACK after participating in a learning activity types short course. Australas. J. Educ. Technol. 2021, 37, 152–169. [Google Scholar] [CrossRef]
  22. Ramírez-Montoya, M.S.; Andrade-Vargas, L.; Rivera-Rogel, D.; Portuguez-Castro, M. Trends for the future of education programs for professional development. Sustainability 2021, 13, 7244. [Google Scholar] [CrossRef]
Table 1. Analytical categories for scientific concept development.
Table 1. Analytical categories for scientific concept development.
CategoryCodeDefinitionCriterion
Conceptual Structuring (CS)CS1Phenomenon/activity-centeredInquiry and activities present, but no relational structuring of concepts
CS2Relational structuringConcepts organized by comparative criteria
(similarities, differences, or conditions)
CS3Definition/generalization-centeredCondition-inclusive definitions or rules explicitly presented
Generalization (GEN)GEN0Case-levelRestricted to presenting observed facts
GEN1Partial generalizationGeneral statements present, but conditions unspecified
GEN2Condition-inclusive generalizationGeneral statements presented with explicit conditions (“when [condition], [outcome]”)
Conceptual Explicitness (EXP)EXP0Terms absentKey terms not clearly present
EXP1Terms presentTerms appear, but meaning not structured
EXP2Definition/sentence-levelTerms and meanings formulated in explicit sentences or through structured summary activities
Table 2. Analytical categories for AI and digital tool functions.
Table 2. Analytical categories for AI and digital tool functions.
CategoryCodeDefinitionCriterion
Concept presentationAI-PRESPresenting concepts or phenomena visually or verballyImage generation, simulation, phenomenon reproduction,
or concept explanation
Concept relational structuringAI-STRSupporting comparison or
relational organization of concepts
Comparative tables, relation mapping, classification activities
(e.g., using Google Slides to organize and compare group experimental results across indicator types)
Concept refinement and generalizationAI-REFSupporting concept refinement, elaboration, or generalizationDefinitional statement generation, explanation supplementation,
condition-inclusive descriptions, or misconception correction (e.g., using Sim Ground to collect real-time student responses, identify incorrect answers, and address misconceptions during the consolidation stage)
Table 3. Summary of conceptual development patterns in draft lesson plans.
Table 3. Summary of conceptual development patterns in draft lesson plans.
UnitCaseInstructional FocusCSGENEXPRationale
Changes in the State of WaterSW-AArtifact design using changes in stateCS1GEN0–1EXP1No comparative or definitional stage; artifact-centered
SW-BComparative inquiry:
evaporation vs. boiling
CS2GEN1EXP1–2Comparative structure present;
condition-inclusive generalization limited
SW-CBoiling experiment with
expressive activities
CS1GEN0–1EXP1No relational structuring; consolidation via creative output
Acids and BasesAB-AMixing acidic/basic solutions; quiz-based reviewCS1GEN0–1EXP1No condition-inclusive definition;
quiz-only consolidation
AB-BThree-stage indicator classification with board-writing planCS2GEN1–2EXP2Explicit classification statements in board-writing plan
AB-CPrediction–experiment–comparison using Claude (AI)CS2GEN1–2EXP2Structured prediction and comparison with board-writing plan
Table 4. Comparison of conceptual development patterns between draft and revised lesson plans.
Table 4. Comparison of conceptual development patterns between draft and revised lesson plans.
UnitCaseCategoryDraftRevisedChange
Changes in the State of WaterSW-ACSCS1CS2Artifact task coupled with concept articulation via prompt template
GENGEN0–1GEN1Partial generalization via prompt-embedded description of state change
EXPEXP1EXP2Sentence-level concept expression required and structured
SW-BCSCS2CS2Comparative structure maintained; visual evidence strengthened via Claude
GENGEN1GEN1Comparative refinement without condition-inclusive definitional rule
EXPEXP1–2EXP2Group report and comparative summary formalized
SW-CCSCS1CS3Shift to definition-centered structure via concept table and board-writing plan
GENGEN1GEN2Condition-inclusive rule statements added
EXPEXP1EXP2Explicit sentence-level concept formalization achieved
Acids and BasesAB-ACSCS1CS1No substantive change; time reallocation only
GENGEN0–1GEN0–1No substantive change
EXPEXP1EXP1No substantive change
AB-BCSCS2CS2Session condensed; comparative structure retained
GENGEN1–2GEN1–2Maintained
EXPEXP2EXP2Maintained
AB-CCSCS2CS2–CS3Teachable Machine model-building raised structuring level
GENGEN1–2GEN2Classification criteria made explicit through AI model training process
EXPEXP2EXP2Maintained
Table 5. Comparison of AI and digital tool use across units and versions.
Table 5. Comparison of AI and digital tool use across units and versions.
UnitCaseVersionKey ToolsFunctionChange
Changes in the State of WaterSW-ADraftZEP, Padlet, WrtnAI-PRESVisualization retained; prompt template added to mediate concept retrieval
RevisedZEP, Padlet, Gemini (prompt template)AI-PRES + AI-REF (partial)
SW-BDraftPhET, Canva, JavaLab AI-PRES + AI-STRClaude added as comparative evidence; linked to concept refinement task
RevisedPhET (‘dot’ scaffolding), Canva, Claude (microscopic visualization)AI-PRES + AI-STR + AI-REF
SW-CDraftPadlet, OurChild AI, Canva, Suno AIAI-REF + expression-centeredShifted to real-time data-based misconception correction in consolidation stage
RevisedSim Ground, concept summary table, ClaudeAI-REF + AI-STR
Acids and BasesAB-ADraftKahoot, PPTAI-PRESNo substantive change in tool function
RevisedKahoot, PPTAI-PRES
AB-BDraftVivasam, Google SlidesAI-PRES + AI-STRStructure retained; session condensed
RevisedVivasam, Google SlidesAI-PRES + AI-STR
AB-CDraftClaude, PadletAI-PRES + AI-STR (partial)Replaced prediction tool with model-building; classification logic made explicit
RevisedTeachable Machine, Padlet, KahootAI-STR + AI-REF
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Lee, J. How Pre-Service Elementary Teachers Develop Scientific Concepts in AI-Integrated Lesson Designs: Implications for Sustainable Teacher Education. Sustainability 2026, 18, 5211. https://doi.org/10.3390/su18105211

AMA Style

Lee J. How Pre-Service Elementary Teachers Develop Scientific Concepts in AI-Integrated Lesson Designs: Implications for Sustainable Teacher Education. Sustainability. 2026; 18(10):5211. https://doi.org/10.3390/su18105211

Chicago/Turabian Style

Lee, Juyoung. 2026. "How Pre-Service Elementary Teachers Develop Scientific Concepts in AI-Integrated Lesson Designs: Implications for Sustainable Teacher Education" Sustainability 18, no. 10: 5211. https://doi.org/10.3390/su18105211

APA Style

Lee, J. (2026). How Pre-Service Elementary Teachers Develop Scientific Concepts in AI-Integrated Lesson Designs: Implications for Sustainable Teacher Education. Sustainability, 18(10), 5211. https://doi.org/10.3390/su18105211

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop