Next Article in Journal
A Game-Based Eye-Tracking Task for Inclusive Educational Assessment in Children with Autism Spectrum Disorder and Dyslexia: An Exploratory Study
Previous Article in Journal
Design and Expert Validation of the MIGEAS Model: A Data Analytics-Driven Framework for Strategic Management and Functional Integration in Higher Education
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Evaluation of Inductive Coding with LLMs

Department for Mathematics and Science Education, Institute for Geography Education, University of Cologne, Gronewaldstraße 2, 50931 Cologne, Germany
*
Author to whom correspondence should be addressed.
Educ. Sci. 2026, 16(8), 1314; https://doi.org/10.3390/educsci16081314
Submission received: 14 July 2026 / Revised: 7 August 2026 / Accepted: 13 August 2026 / Published: 17 August 2026

Abstract

Large Language Models (LLMs) like ChatGPT are reshaping qualitative research by offering data-driven analysis. Recent studies focus on labeling and classifying qualitative data, yet the generative process behind code system development is underexplored. This study investigates how ChatGPT constructs coding systems from interview data under iterative code engineering. The resulting code system is systematically compared to an inductive coding framework derived by content analysis—quantitatively by SBERT analysis, Jaccard index, and network analysis, and qualitatively. The results of NLP-based analysis show that there are mostly high cosine similarities in the one-to-one mapping between ChatGPT codes and content analysis. The overall Jaccard index is low, indicating limited overlap between the coding units and, consequently, that the two coding systems often did not refer to the same content. However, the mapped codes showed substantial variation: while nearly half exhibited high overlap, others showed no overlap. Network analysis shows that network density is comparable across the two systems, indicating a similar level of interconnectedness among codes despite differences in network size. Finally, iterative prompt engineering could create a stable code system, but ChatGPT lacks in domain-specific coding addressing the research questions adequately, warranting further investigation across larger datasets.

1. Introduction

The integration of generative artificial intelligence (GenAI) into educational settings has sparked significant interest in its transformative potential; not only for teaching and learning (Luis et al., 2025; Stelea et al., 2025; Wang et al., 2025), but increasingly for research practices themselves (Dahal, 2024; Goyanes et al., 2025; Henderson et al., 2026; Schueller et al., 2026). While recent studies emphasize the utility of LLMs like ChatGPT for tasks such as feedback generation, formative assessment, and knowledge summarization (Dai et al., 2023), their potential as tools for qualitative data analysis remains underexplored, particularly in educational research.
Qualitative coding, the interpretative process of assigning meaning to textual data, is a cornerstone of empirical educational research (Braun & Clarke, 2006; Saldaña, 2025). This process is labor-intensive, and transformer-based LLMs like ChatGPT (OpenAI, 2025) have demonstrated promising capabilities for qualitative tasks, including coding, thematic analysis, and synthesis of unstructured data (Christou, 2024; Gao et al., 2025; Goyanes et al., 2025; Kwon et al., 2025; Xu, 2026). Their capacity to generate coding frameworks, apply pre-existing codebooks, and assist in thematic abstraction suggests potential applications in research areas traditionally dominated by human expertise (Xiao et al., 2023). Despite their capacity for pattern recognition and language generation, LLMs like ChatGPT often produce factual inaccuracies and hallucinations, including confidently fabricated information (Dahal, 2024; Henderson et al., 2026), raising ethical concerns (Naeem et al., 2025). Recent efforts to mitigate hallucinations have explored various strategies, including output validation (e.g., Retrieval-Augmented Generation; Ayala & Bechard, 2024), real-time verification frameworks (Kang & Mo, 2024), and input sophistication (Morgan, 2023), primarily through prompt engineering.
Studies show that prompt structure significantly affects model outputs, with better-engineered prompts leading to improved reliability and thematic consistency (Khalid & Witmer, 2026). Prompting formats that have proven particularly relevant include: (1) instruction-based prompting that relies on explicit directives, such as tasks or checklists (Alhindi et al., 2022), (2) role-based prompting that defines a target persona (e.g., “You are a scientific editor”), helping modulate tone, expertise, or evaluative stance, and dismantling hallucinations (Louatouate & Zeriouh, 2025), (3) contextual prompting that embeds background material into the prompt to guide response grounding (Goswami et al., 2024), (4) multi-turn prompting that structures the interaction dialogically across multiple steps, allowing for clarification, refinement, or reflection, and preventing overload (Liu, 2025), (5) example-based prompting that introduces model responses as templates to support output alignment (Ben-David et al., 2022), and (6) chain-of-thought prompting that breaks down tasks into step-by-step sequences, encouraging the model to make intermediate inferences and thereby improving accuracy and interpretability in complex analytical tasks (Kwon et al., 2025). As Gao et al. (2025) demonstrate, LLMs can also support collaborative coding workflows, serving not as replacements but as aids that help human analysts refine and negotiate code (Christou, 2023, 2024).
Given these advancements in prompt engineering and its role in reducing model hallucinations, ensuring contextual grounding, and maintaining systematic coding logic, we propose an inductive coding framework for structuring complex interview data generated by ChatGPT. Each prompt design element (instruction, role, context, iteration, synthesis) translates theoretical insights into stepwise coding procedures. To evaluate the framework, we use both qualitative and quantitative methods. Specifically, we compare the results to a prior qualitative content analysis of the same material, which already included the development of a coding system and the coding process. In addition to qualitative analysis, we apply quantitative methods to extend the analysis by enabling systematic measurement, statistical comparison, and the examination of observable patterns across both approaches.
Research aims: Our empirical case draws on 11 qualitative interviews with geography teachers on stereotyping in the classroom, an interpretively rich topic that requires contextual sensitivity. Based on theoretically derived prompting strategies (instruction, role, context, multi-turn prompting, examples, and chain-of-thought prompting), we developed a prompting framework for inductive coding with ChatGPT. We then systematically compare this code system and the coding of 11 transcripts against those derived from qualitative content analysis. The primary aim is to provide a systematic, structural, and semantic comparison between the two coding systems, thereby identifying areas of convergence and divergence. Explicitly, we compare the outcome of theoretically derived strategies with human-coded approaches by addressing the following Research Questions (RQs):
  • RQ1: Can we map the ChatGPT code system onto the content analysis system based on semantic similarities?
  • RQ2: Do semantically corresponding codes refer to the same content represented by the coding units?
  • RQ3: Are there systematic differences between the two code systems?

2. Materials and Methods

We use a part of the dataset from Dörfel et al. (2025b)consisting of interview data from eleven anonymous semi-structured online interviews with German geography teachers, both in-service (n = 9) and pre-service teachers (n = 2), with varying teaching experience, school context (urban/rural), and regional background on stereotypes in geography class. The interviews were audio-recorded and fully transcribed using noScribe (version 0.4.1) by Dröge (2023), followed by two rounds of manual correction.
The data were collected to answer the two RQs:
  • RQ1: Does stereotyping occur in geography education? If so, how and why?
  • RQ2: What strategies do geography teachers use to address stereotyping in geography class?
Thus, the dataset was analyzed twice: (a) with content analysis (Dörfel et al., 2025a) and (b) with ChatGPT.
Figure 1 gives an overview of the respective procedures, which are then described in detail in the following Sections.

2.1. Content Analysis

The content analysis followed Kuckartz and Rädiker’s (2022) seven-step procedure.
The key steps include:
  • Careful reading, memo writing, and creation of case summaries.
  • Identification of coding units (n = 86) across all transcripts
  • Initial deductive coding using three deductive main codes (examples of stereotypes, reasons for stereotyping, and strategies to deal with stereotypes).
  • Inductive derivation of 20 subcodes
  • Revision of the code structure and complete recoding of all material.
  • Application of analytical strategies aligned with the research questions.
  • Implementation using MAXQDA Analytics Pro Version 24.4.0 (VERBI Software, 2024).
See Dörfel et al. (2025a) for the whole analysis.

2.2. ChatGPT Coding

This research is based on the use of ChatGPT-5.2 11 December 2025 snapshot version (OpenAI, 2025). Any analysis, interpretation, or categorization reflects the interplay between human inquiry and machine-generated responses, meeting the Journal’s AI Policy (MDPI, n.d.). All prompt formats are found in Appendix A. In Phase I, an initial coding system was generated for each interview transcript. Phase II–IV repeated and synthesized outputs to test reliability and to finally have one code system. Phase V tested coding with this code system, considering reliability through systematic iteration. We now describe each phase in more detail.

2.2.1. Phase I: Initial Development

First, we used a prompt to develop a coding framework for each of the 11 transcripts (see Appendix A.1). While creating the prompt, we applied role-based prompting (ChatGPT is to act as a research assistant) (Louatouate & Zeriouh, 2025) instruction-based prompting detailed instructions, such as task, objective, clear rules, output format (Alhindi et al., 2022), and contextual prompting by giving RQ1 and RQ2, guiding the development of the coding framework (Goswami et al., 2024). Finally, one transcript was inserted in each case to avoid overload. Due to the length of the prompts and transcripts, and our aim to replicate an iterative coding process as closely as possible, we sought to minimize potential processing issues. This process was repeated eleven times, resulting in eleven coding systems. Because ChatGPT initially tended to generate overarching definitional codes that addressed general conceptualizations of stereotypes rather than the specific RQs, the prompt was refined. Table 1 shows the output format.

2.2.2. Phase II: Synthesis

The eleven individual AI-generated codebooks were then synthesized into a comprehensive, hierarchical coding system. A stepwise procedure was followed:
  • Considering all coded content from the 11 transcripts,
  • Meaningfully merging content-related similar codes,
  • Creating new, overarching codes where they are thematically justified,
  • Systematically resolving thematic overlaps,
  • And including only codes that were used in phase I.
The prompt can be seen in Appendix A.2. As a result, a first merging code system was developed. Table 2 shows the output format.

2.2.3. Phase III: Repetition

As demonstrated by Zeng et al. (2025), the reliability and reproducibility of LLM outputs, such as those from ChatGPT, can vary significantly. Considering this, we repeated the coding system synthesis 5 times to assess output consistency and stability (Vaidya & Saini, 2026). Each synthesis was conducted in a separate chat session of the project to avoid cross-contamination or influence between outputs. The memory function in the settings was off.

2.2.4. Phase IV: Synthesis

The five synthesized category systems were then merged and clustered into a single final category system by another prompt (Appendix A.3) that included all five code systems. To reduce exceptional cases, the final code system includes only codes that were present in at least three of the five versions of the prior generated code systems.

2.2.5. Phase V: Coding of Transcripts

ChatGPT was then instructed to run coding for each original transcript with its own developed code system (Appendix A.4), using the following main rules:
  • Each transcript was divided into coding units (CU).
  • Using only the predefined codes and definitions
  • Justifying each code assignment based on explicit matches with definitions
  • Returning no matching code if none is applied
As an output, ChatGPT should use a standardized table (Table 3).
Because ChatGPT, as an LLM, produces non-deterministic outputs, we independently ran each transcript five times to assess run-to-run stability. We then evaluated the consistency across runs using the Jaccard Index (J). The Jaccard index was high (J = 0.83) (Greve & Wentura, 1997), indicating strong run-to-run stability. For the final analysis, only codes that were present in at least three of the five independent ratings were retained.

2.3. Data Analysis

We used these datasets to compare the ChatGPT-generated code with pre-existing coding from content analysis and applied quantitative and qualitative methods. First, we assessed how similar the two coding systems are using Natural Language Processing (NLP) and cosine similarity. Secondly, we mapped both systems one-to-one using the Hungarian Algorithm to optimize a global mapping. We then compared the codes assigned by ChatGPT with the content analysis across coding units using the Jaccard Index to assess their overlap. To avoid relying solely on algorithms, we qualitatively evaluated the mapping of both code systems to interpret similarities and divergences. Figure 2 gives an overview of the research design.

2.3.1. Semantic Comparison of Code Systems: SBERT-Based NLP

To address RQ1, whether the ChatGPT-generated coding system can be mapped onto the content analysis system, we first applied a semantic similarity approach based on Sentence-Sentence Bidirectional Encoder Representations from Transformers (SBERT) (Reimers & Gurevych, 2019). This method enables the systematic quantification of conceptual overlap between codes, which is a prerequisite for establishing correspondences across coding systems.
Both systems were transformed into dense, semantically meaningful vector representations using SBERT, an applied NLP technique (Ghojogh & Ghodsi, 2020). This step was necessary to represent textual codes in a comparable numerical form within a shared semantic space, enabling direct comparisons of similarity. For each code, the label and its definition were linked into a single text string to ensure that both the core concept and its contextual specification were captured in the embedding process.
We then employed the pre-trained model paraphrase-multilingual-MiniLM-L12-v2 (Reimers & Gurevych, 2019), as it is optimized for capturing semantic similarity across languages and phrasing variations. This is particularly important given that conceptually equivalent codes may differ in wording between the two systems. Pairwise semantic similarity between all codes was calculated using cosine similarity (Salton et al., 1975), where values closer to 1 indicate stronger conceptual overlap. This provides a continuous measure of alignment between individual codes. A heatmap generated with ggplot2 (Wickham, 2016) provides an overview of similarity distributions across both coding systems.

2.3.2. One-to-One Mapping: Hungarian Algorithm

Building on the cosine similarity matrix, a one-to-one mapping between the two coding systems was computed by maximizing the total semantic similarity across all possible one-to-one matchings. Since the two coding systems contained different numbers of codes (14 ChatGPT-generated codes and 20 content analysis codes), the mapping was formulated as a global optimization problem under exclusivity constraints and solved using the Hungarian algorithm, which guarantees a globally optimal one-to-one assignment (Kuhn, 1955). Instead of selecting the best match for each code independently, the optimization identified the combination of one-to-one mappings that maximized the overall semantic similarity across both coding systems, while ensuring that each ChatGPT-generated code was assigned to at most one content analysis code. Accordingly, only the maximum possible number of one-to-one mappings was retained, while all other codes remained unmapped (Saadaoui, 2026). This approach provides a systematic and reproducible correspondence between two independently developed coding systems without imposing assumptions about higher-level code categories. For subsequent analyses, a cosine similarity threshold of 0.60 was applied to retain only code pairs showing sufficient semantic similarity and to exclude assignments with only weak semantic correspondence, even if they are part of the globally optimal solution.

2.3.3. Overlap of Coding Unit Assignments

To address RQ2, namely whether semantically corresponding codes refer to the same coding units, a one-to-one similarity-based mapping was used to define comparable code pairs between ChatGPT-generated and content-analysis codes. This enables a direct comparison of coding unit assignments at the level of coded segments. For each mapped code pair, binary coding assignments (0/1) were compared across all coding units. The Jaccard Index (Jaccard, 1901) was used to quantify the overlap of coding unit assignment by measuring the proportion of shared coding units in actual code assignments (1/1) relative to all coding units assigned by at least one of the two corresponding codes. In addition, Jaccard values were calculated for each mapped code pair to assess code-specific variation in coding-unit overlap.

2.3.4. Structural Comparison via Network Analysis

To address RQ3, whether systematic structural differences exist between the two coding systems, separate code co-occurrence networks were constructed for the content analysis (CA) and ChatGPT coding systems (Newman et al., 2006) using R (R Core Team, 2026) and the igraph package (Csárdi et al., 2026). Binary coding unit × code matrices were constructed for each coding system. Codes were represented as nodes, while edges indicate their co-occurrence within the same coding segment. Edge weights reflected the number of coding units in which the respective pair of codes co-occurred.
To evaluate the overall connectedness of each network, we calculated network density, representing the proportion of observed code co-occurrences relative to all possible co-occurrences within the network (Bedru et al., 2020). We further calculated the global clustering coefficient, which indicates the extent to which codes connected to the same node also co-occur with one another, thereby forming clusters (Kong et al., 2019). To compare the relative structural importance of codes across systems, node sizes were scaled using eigenvector centrality (EV-Centrality), which captures whether a code is connected to other highly connected codes within the network (Bloch et al., 2023). If a segment was assigned to multiple codes, all pairwise code combinations were represented as edges in the network. To support visual interpretation, the spatial arrangement was determined using the Fruchterman–Reingold algorithm, which positions highly connected nodes closer together, facilitating the identification of structural patterns.
To assess whether the observed network density and global clustering coefficient differed from the values expected under random code assignments while preserving the observed row and column sums of the coding matrix, separate null models were generated for each coding system using the Curveball algorithm (Strona et al., 2014), which is implemented in the vegan package in R (Oksanen et al., 2026). As the underlying coding unit × code matrices form bipartite structures that are projected onto code co-occurrence networks (Catalano & Mantegna, 2026), and the two coding systems were based on different codebooks (content analysis vs. ChatGPT), each network was evaluated relative to its own null model. The algorithm preserves both the number of codes assigned to each coding unit and the overall frequency of each code, while randomizing the underlying binary coding matrix. For each coding system, 10,000 randomized binary coding matrices were generated, from which projected code co-occurrence were reconstructed. The network density and the global clustering coefficient were calculated for each randomized projected network to obtain null distributions. The observed values were evaluated against these distributions by calculating standardized z-scores, 95% null intervals, and two-sided Monte Carlo p-values.

2.3.5. Qualitative Evaluation of Code Systems

To supplement the quantitative analysis, we further evaluate the results qualitatively by the following dimensions:
  • Structural composition and cross-method convergence
We examined whether mapped codes and network clusters remained structurally aligned within the same overarching analytical domains or whether cross-domain overlaps emerged. In addition, we assessed the extent to which semantically similar code pairs identified through SBERT mapping were reproduced as coherent structural units across both coding systems (network analysis), thereby identifying areas of convergence between the approaches.
2.
Granularity patterns and contextualization
We analyzed differences in the level of differentiation and contextual specificity between the two coding systems by examining one-to-one mappings and unmatched codes. Particular attention was given to the extent to which subject-specific and context-dependent aspects were represented in each system.
3.
Explanatory logic and agencies
We examined how both coding systems conceptualized and organized explanations for stereotyping by comparing the thematic structure and causal orientation of mapped codes. This included assessing whether explanations were framed in broader, more general patterns or in more differentiated, multidimensional ways. Furthermore, we evaluated what responsibilities and agencies were distributed across both coding systems.

3. Results

3.1. Development of the ChatGPT Code System and Coding

Following the development and multiple syntheses of the ChatGPT-generated coding system (Phases I–IV), the characteristics of the five synthesized coding systems were examined. Although the five synthesized coding systems showed a high degree of thematic consistency, some structural differences were observed regarding the number and organization of codes. The number of subcodes ranged from 16 to 24. Trial 3 contained the largest number of subcodes (n = 24), whereas Trials 1 and 5 each comprised 16 subcodes, Trial 2 comprised 18, and Trial 4 comprised 20. Except for Trial 4, all synthesized coding systems consisted of three main codes corresponding to RQ1-3. Trial 4 additionally included the higher-order code Perception and Evaluation, which occurred only once and was therefore excluded during the final synthesis.
To derive a robust final coding system (Phase IV), the five synthesized coding systems were merged. Only codes that appeared in at least three of the five synthesized versions were retained in the final coding system. Table 4 presents the final synthesized coding system. Because the interviews were conducted in German, all coding procedures were performed in German. The final coding system was subsequently translated into English using DeepL (8 June 2026).
Following the development of the final coding system, each of the eleven transcripts was coded five times in independent chat sessions using the finalized coding framework (Phase V). To assess the stability of the coding outputs, the pairwise Jaccard index was calculated across the five coding runs for each transcript. Table 5 presents the resulting Jaccard similarity values for measuring run-to-run stability. The overall Jaccard similarity across all transcripts was 0.83, indicating a high run-to-run stability throughout five runs.

3.2. Can We Map the ChatGPT Code System onto the Qualitative Content Analysis Based on Semantic Similarities?

The coding system for content analysis was developed using a deductive–inductive approach guided by two RQs. It comprises three main codes and 20 subcodes (Appendix A.5). The system is structured around the RQs. It thus differentiates into three main codes: Examples of stereotyping, reasons for stereotyping, and strategies in dealing with stereotypes. The three main codes of the ChatGPT code system are: Forms of manifestation, conditions of emergence, and pedagogical strategies, which were then analyzed with 14 subcodes in total. For a mapping of both systems, we chose an NLP-based Similarity Analysis (SBERT) and created (a) a heatmap to show general similarities across all codes and (b) a one-to-one mapping.

3.2.1. NLP-Based Similarity Analysis (SBERT)

In total, 63 code pairings with cosine similarity scores ≥ 0.60 were identified. Overall, the heatmap reveals substantial semantic overlap between the ChatGPT-generated and manually developed coding systems, while also indicating that the strength of this overlap varies across individual codes (Figure 3). In particular, codes such as Professional self-reflection, Didactic reduction and time pressure, Working with data and contextualization, and Perspective shift show high correspondence with codes from the content analysis, with cosine similarity scores exceeding 0.7. In contrast, other codes, most notably Migration and religion and Problem- and deficit focus, exhibit comparatively weaker semantic alignment with the content analysis code system. Beyond individual pairings, a broader structural pattern emerges in the extent of semantic linkage. Certain ChatGPT codes, especially Dialogue-based reflection and intervention and Perspective shift and irritation, demonstrate consistently high similarity across multiple content analysis codes, suggesting broader conceptual coverage.

3.2.2. One-to-One Mapping of ChatGPT Codes to Content Analysis Codes

In the next step, we applied a one-to-one mapping to map all 14 ChatGPT codes to the content analysis codes, in descending order of cosine similarity, using code names and descriptions as the similarity index (Table 6). Overall, the mapping indicates that most ChatGPT codes could be assigned to semantically corresponding content analysis codes with moderate to high similarity, suggesting a substantial structural correspondence between the two coding systems. Only one match out of 14 (Problem- and deficit focus + Country groups) is below the 0.6 threshold and can therefore be considered to have lower semantic overlap. Therefore, the subsequent analyses refer to the remaining 13 mapped code pairs. Since the content analysis identified six more codes than ChatGPT, six codes go unmatched: Subject content and Countries as Examples of stereotyping, Teachers (causes) and Students (causes) as Reasons for stereotyping, and Causal analysis and Student-oriented strategies as Strategies in dealing with stereotypes.

3.3. Do Semantically Corresponding Codes Refer to the Same Content Represented by the Coding Units?

The one-to-one mapping forms the basis for comparing the coding unit assignments of semantically corresponding codes derived from ChatGPT and Content analysis. To quantify the overlap in coding unit assignments, the Jaccard Index was calculated for each pair of mapped codes and across all 13 mapped codes (global Jaccard index) (see Table 7).
The resulting global Jaccard Index is 0.39, indicating limited overall overlap in coding-unit assignments between the semantically mapped ChatGPT and content-analysis codes. The degree of overlap varies substantially across the 13 mapped code pairs (J = 0–0.86). Five mapped code pairs show a comparatively high overlap (J > 0.60), including Dialogue and reflection and Dialogue-based reflection and intervention, Increasing complexity and Making diversity and differentiation visible, Teachers (strategies) and Professional self-reflection, Evidence-based working and Working with data and contextualization, as well as Resources: Textbooks and Material and textbook influence. For these code pairs, both coding approaches assigned semantically corresponding codes to largely the same coding units. In contrast, three mapped code pairs show low to moderate overlap (J = 0.30–0.40), whereas the remaining five code pairs exhibit very low overlap (J < 0.14).

3.4. Are There Systematic Differences Between the Two Code Systems?

To characterize the overall structure of the two coding systems, we calculated the network density, global clustering, and eigenvector centrality. The content analysis network has a density of 0.40 and a global clustering coefficient of 0.56, while the ChatGPT-based network has a density of 0.36 and a global clustering coefficient of 0.61. Compared to their respective null models, both networks do not differ in terms of network density (Content analysis: z = −0.60, p = 0.611; ChatGPT: z = −1.66, p = 0.154) or global clustering (Content analysis: z = 0.34, p = 0.683; ChatGPT: z = 1.19, p = 0.239), suggesting that the observed values for connectivity and clustering did not exceed those expected based solely on the coding structure.
Figure 4 and Figure 5 provide a visual overview of the overall structure of both coding systems: In these network representations, nodes correspond to individual codes, node size reflects eigenvector centrality (larger circles indicate higher centrality, while smaller circles indicate lower centrality), edge thickness indicates the strength of code co-occurrence, and colors represent thematic domains. Both networks exhibit a densely interconnected core surrounded by more peripheral codes, although differences in the distribution of central nodes become apparent. Although the two coding systems showed similar global structural properties, EV-centrality indicates that structural importance is distributed differently across the networks. Whereas the content analysis network distributes centrality across several strongly connected codes, the ChatGPT-based network concentrates centrality in fewer dominant codes. In the content analysis network, Subject content represents the most central node (EV = 1.00), followed by Continents (0.71). In addition, several other codes—including Students (c), Resources: Textbooks, and Teaching methods—also show relatively high centrality values, forming a broader set of highly connected nodes. In the ChatGPT-based network, Spatial and developmental stereotypes emerge as the most central code (EV = 1.00), followed by Problem- and deficit focus (0.91) and Material and textbook influence (0.90).
Across both networks, a similar higher-order pattern emerged. Codes associated with reasons for stereotyping and examples of stereotyping tend to occupy more central positions and are connected to more other codes. In contrast, strategy-related codes generally show lower centrality values and fewer connections, indicating a more peripheral position within the overall network structure.

3.5. Qualitative Evaluation

The findings are reported according to the analytical dimensions.
  • Structural composition
The code system generated by ChatGPT was output in tabular form and matches the requested output format (main code, subcode, description).
It is striking that the mapping from ChatGPT Codes onto the content analysis based on SBERT matches at the main code level. Except for one code mapping (Socio-cultural influences + Migration and religion), all other mappings comply with the same main code level: Examples matched with Forms of manifestations, Reasons for stereotyping with Conditions of emergence of stereotypes, both addressing RQ1, and Strategies in dealing with stereotypes with Pedagogical strategies in dealing with stereotypes, addressing RQ2. The structure of both code systems, oriented on the RQs, seems equivalent and therefore follows a similar analytic logic. Looking more closely at the mapping, it is notable that some mappings seem contradictory: Perspective shift occurs in both code systems, but both codes were mapped to other codes due to greater similarity based on their descriptions. Furthermore, high semantic similarity does not always align with a high Jaccard Index or similar EV-Centrality: While Geography-specific factors moderately co-occur with nine other codes, the semantically mapped ChatGPT code Didactic reduction and time pressure remains structurally peripheral in the ChatGPT network.
  • Granularity and differentiation patterns
The content analysis yielded six additional subcategories compared with ChatGPT, indicating greater differentiation across all three main codes. The six unmapped codes are: Subject content and Countries as Examples for stereotyping; Teachers (causes) and Students (causes) as reasons for stereotyping; Causal analysis; and Student-oriented strategies for Strategies in dealing with stereotypes. That the code Subject content is missing from the ChatGPT code system is thematically striking, since it is the most frequent code and the one with the highest EV-Centrality in the content analysis. Besides this, it is one of the most subject-specific codes and directly addresses part of RQ1: Does stereotyping occur in geography class? And if so, how and why? Finally, the ChatGPT network exhibits a concentration of high EV-centrality values in three codes. In contrast, the content analysis network contains fewer high EV-centrality values overall, with many codes exhibiting similar, moderate EV-centrality levels. Furthermore, some codes in the content analysis are further differentiated into subcodes (Continents, Countries, Country groups), whereas ChatGPT summarizes spatial stereotypes under a single subcode (Spatial and developmental stereotypes).
  • Explanatory logic and agencies
Looking more closely at the explanatory logic of stereotyping, addressing the last part of RQ1, we can identify four mapped reasons for stereotyping. Three reasons of the content analysis remain unmapped: Teachers and students as reasons for stereotyping, and sociocultural influences. Taking all the reasons of the content analysis together, we identified many different actors, like teachers and students as attendees in class, social backgrounds as socialization factors, subject-specific factors like complexity reduction, textbooks as a central medium in class, and general educational policy factors like the time pressure explaining stereotyping in geography class. In contrast, the ChatGPT code system addresses only half of these factors, leaving out two main actors in the classroom. Students, as one of the main actors, remain very limited in addressing Knowledge deficits as a breeding ground, leaving no room for other student-related factors. In contrast, Students (c) have the third-highest EV-Centrality in the content analysis network, indicating structural importance in explaining stereotyping in geography class. By implication, it seems that ChatGPT focused more on structural conditions than on individual actors. In contrast, the content analysis reflects more nuanced reasons, likely leading to a more differentiated analysis and results regarding RQ1.

4. Discussion

4.1. Can We Map the ChatGPT Code System onto the Content Analysis System Based on Semantic Similarities?

The analysis indicates substantial semantic overlap between ChatGPT codes and Content analysis, with moderate-to-high cosine similarities (0.6–0.81), resulting in 13 mappings as determined by the Hungarian algorithm. Generally, the mapping of subcodes matches the corresponding main codes in 11 of 13 pairings, reflecting an orientation throughout the research aims: what stereotypes exist (examples), why they exist (reasons), and how teachers deal with them (strategies). Two pairings (Migration and religion (Manifestations) and Sociocultural influences (Reasons), and Eurocentrism (Conditions of emergence) and Perspective shift (strategies)) do not match equal main codes. This might be a disadvantage of the Hungarian algorithm, since the one-to-one mapping enforces the mapping, and codes might run out of partners. However, both similarities are moderate and above the 0.6 threshold, but the lowest and second-lowest of all mapped codes (0.61 and 0.6).
Another striking result in our data is the consistently high semantic similarity of strategy-related codes, which occupy the six highest ranks (1st–6th). This might be explained by the fact that strategies were less geography-specific than examples or reasons for stereotyping in our dataset, which is supported by looking more closely at the unmapped content analysis codes: First, they primarily represent domain-specific examples and reasons (differentiated spatial stereotypes, such as countries and country groups, Subject content, and geography-specific reasons). Secondly, they represent more actor-specific causes, and thirdly, more differentiated strategy codes.
These findings can be supported by previous research: An overarching pattern can be identified, while domain-specific focus appears more difficult (Lixandru, 2024; Myers et al., 2024). ChatGPT, as a large language model, can identify and emphasize dominant concepts and recurring patterns within large datasets based on probabilistic relationships (Myers et al., 2024). Since the model does not possess an independent contextual understanding of the text, it cannot fully follow interpretative logics that require human reasoning (Huang et al., 2026; Leifheit et al., 2024; Trott et al., 2023), which, in this case, is visible in the domain-specific perspective needed to address the RQs. However, reflecting all results from RQ1, we cannot support the overall conclusion of Morgan (2023) and Turobov et al. (2024) that LLMs in general struggle to generate interpretive themes, as shown by ChatGPT’s structured code system for a sensitive, interpretive topic like stereotyping.

4.2. Do Semantically Corresponding Codes Refer to the Same Content Represented by the Coding Units?

Using the Hungarian algorithm-mapped codes across the eleven transcripts, the Jaccard Index (J = 0.39) was used to quantify the overlap in coding unit assignments between ChatGPT and content analysis. However, the overlap varied highly across the 13 code pairs: five codes showed very low agreement (J < 0.14), whereas three codes showed low-moderate agreement (0.33–0.38) and five codes showed high agreement (J = 0.66–0.89). Comparing the Jaccard Index with semantic similarities further revealed that semantically similar codes did not necessarily refer to the same coding units and vice versa. For example, Teaching methods and Perspective shift and irritation exhibited high semantic similarity (cosine similarity ≈ 0.78), yet their overlap in coding-unit assignments was very low (J = 0.10). In contrast, Dialogue and reflection and Dialogue-based reflection and intervention showed both high semantic similarity (cosine similarity = 0.81) and substantial overlap in coding-unit assignments (J = 0.66).
This indicates that semantically similar code labels do not necessarily originate from the same textual evidence. Rather, the present findings suggest that the two coding systems differ in how qualitative information is structured during coding.
One pattern consistently observed in the present data is a difference in code granularity. Overall, the ChatGPT-generated coding system tended to aggregate conceptually related phenomena into broader code categories, whereas the content analysis differentiated these phenomena into several more specific codes. This can be illustrated by the matched pair Spatial and developmental stereotypes (ChatGPT) and Continents (content analysis). Although both codes overlap conceptually with respect to spatial stereotypes, the ChatGPT code additionally encompasses developmental stereotypes and therefore represents a broader conceptual category. In contrast, the content analysis distinguishes spatial stereotypes into more subcodes (Continents, Countries, and Country groups), thereby distributing coding unit assignments across multiple codes. This refers to almost all ChatGPT codes, which, when reflecting the broader code names, aggregate two parts into one code (e.g., Migration and Religion, Spatial and developmental stereotypes, Perspective shift and irritation, Working with data and contextualization, Social and origin stereotypes, Complexity reduction and time pressure, etc.).
Consequently, coding units assigned to the broader ChatGPT category are only partially represented by the matched content analysis code, resulting in partial overlap despite substantial semantic similarity. The observed differences in code granularity are consistent with previous findings suggesting that LLM-generated coding systems tend to comprise fewer and broader codes than coding systems developed through conventional thematic or qualitative content analysis (Muasher-Kerwin et al., 2026). This, in turn, might affect the applicability of LLMs in inductive coding, but strongly depends on the research aims of each study: While the domain-specific outcome was of high research interest in our original study and therefore got high focus in the development of the code system, identifying overall patterns in larger datasets might be of interest in other studies and thus, LLM coding might be useful. However, critical reflection on the output is essential.
These findings also raise a more fundamental question regarding the emergence of LLM-generated coding systems. In qualitative content analysis following Kuckartz and Rädiker (2022), inductive code systems are assumed to emerge through an iterative process in which coding units are repeatedly examined, compared, differentiated, and progressively abstracted into higher-order conceptual codes. Whether a comparable process underlies the generation of coding systems by LLMs remains unknown, even though we requested a tabular reference of the merged codes during the coding system’s synthesis. Although the resulting ChatGPT coding system appears semantically coherent, the internal process through which individual coding decisions are synthesized into broader conceptual categories cannot be observed. Owing to the black-box nature of large language models (Huang et al., 2026; Leifheit et al., 2024; Trott et al., 2023), it remains unclear whether broader codes emerge through a process analogous to iterative qualitative abstraction or through fundamentally different mechanisms of statistical pattern recognition. Consequently, the present approach cannot determine why semantically similar code systems sometimes exhibit only limited overlap in coding unit assignments. Differences in code granularity are likely to explain part of this pattern, but they do not fully account for the observed discrepancies. Future studies using larger and more diverse datasets should therefore place greater emphasis on reconstructing this process, for example, by systematically documenting intermediate coding steps and anchor examples. Such transparency would not only improve the traceability of LLM-generated code systems but also help identify potential inconsistencies or unsupported abstractions that arise during code generation, particularly given that LLMs may produce hallucinated outputs (Farquhar et al., 2024; Kalai et al., 2026).
A further explanation may lie in differences in trigger conditions. Codes showing high overlap in coding unit assignments may represent concepts with relatively explicit and consistently identifiable coding criteria. In contrast, lower overlap may indicate concepts that depend more strongly on contextual interpretation. This interpretation is consistent with previous findings suggesting that ChatGPT relies more strongly on pattern-based code assignment, whereas manual coding incorporates a more context-sensitive interpretation of textual material (Yue et al., 2025). Together, these findings suggest that observed disagreements reflect not only differences in coding decisions but also differences in how qualitative information is conceptualized and structured during the development of a coding system in both coding procedures: LLM-based and human-derived iterative code systems on sensitive, interpretive material.
Finally, sensitivity to prompting needs to be reflected. Recent studies have shown that prompt design can substantially influence the structure, specificity, and consistency of LLM-generated qualitative analyses (Han et al., 2026); for example, through personality and emotion-driven variations (Ma et al., 2025). Furthermore, as demonstrated by Zeng et al. (2025), the reliability and reproducibility of LLM outputs, such as those from ChatGPT, can vary considerably even under identical conditions. The prompting strategies used in the present study were therefore developed iteratively, guided by established prompting principles and repeatedly refined until stable coding systems were obtained (Alhindi et al., 2022; Ben-David et al., 2022; Goswami et al., 2024; Kwon et al., 2025; Liu, 2025; Louatouate & Zeriouh, 2025). While this supports the methodological robustness of the proposed workflow, alternative prompting strategies may produce different coding structures and should therefore be systematically investigated in future research.

4.3. Are There Systematic Differences Between the Two Code Systems?

Although both networks exhibit a comparable network density and global clustering, the distribution of eigenvector centrality revealed structural differences between the two coding systems. The ChatGPT-based network showed a stronger concentration of EV-centrality in a small number of dominant codes. Descriptively, this pattern indicates that a limited set of overarching categories accounts for a larger share of the observed co-occurrences. In contrast, EV-centrality was distributed across a broader range of codes in the content analysis network, suggesting a more even distribution of structural importance across concepts (Bloch et al., 2023; Newman et al., 2006).
This pattern can be interpreted as a difference in analytic granularity. The higher concentration of central nodes in the GPT system suggests a tendency to aggregate conceptually related phenomena into broader codes, as suggested previously in relation to the code names. Such a pattern is consistent with previous research suggesting that AI-generated coding systems tend to prioritize generalizable patterns over fine-grained conceptual distinctions (Morgan, 2023; Peters & Chin-Yee, 2025). As argued by Christou (2024) and Dahal (2024) GenAI is particularly effective at identifying recurring structures in textual data but may rely on more abstract, less hierarchically differentiated representations than researcher-driven qualitative analysis.
These differences become especially visible in relation to domain-specific content. In the content analysis network, geography-specific factors are integrated into the broader structure and linked to multiple other codes, indicating that subject-specific contextualization is integral to the analytical framework. In contrast, comparable domain-specific concepts are largely absent from the central structure of the ChatGPT-based network. The most closely related concepts (Didactic reduction and time pressure) remain peripheral, supporting the results discussed in RQ1.
Methodologically, these findings demonstrate that similar global network properties do not necessarily imply similar conceptual structures. While network density and global clustering characterize the overall level of connectivity within a network, eigenvector centrality provides complementary information by identifying the concepts that occupy structurally influential positions. Consequently, differences between qualitative coding systems may become apparent primarily at the level of node centrality, even when global network measures remain largely comparable.
Furthermore, the consistently lower centrality of strategy-related codes in both networks suggests that the analytical focus of both systems lies more strongly on describing and explaining the phenomenon than on identifying responses or interventions. Overall, the findings indicate that systematic differences between the two coding systems lie in their conceptual organization, level of differentiation, and integration of domain-specific knowledge.

4.4. Limitations

Limitations of our study include the small dataset and its focus on a specific research topic. Given the interpretive nature of the interviews and the limited sample size, future research should validate both the procedure and the prompts using larger datasets and different research contexts to assess the generalizability of these identified patterns. Prompt sensitivity also warrants further investigation, as alternative prompt formulations may influence the generated coding systems despite iterative prompt refinement and repeated independent coding runs designed to account for the non-deterministic nature of LLM outputs. Although the Hungarian algorithm provides a reproducible one-to-one mapping between the two coding systems, alternative mapping strategies may yield different code correspondences and, consequently, slightly different comparison results. Furthermore, as qualitative content analysis is interpretive, multiple valid coding systems may emerge from the same dataset, particularly when analyzing complex and sensitive phenomena, such as stereotyping. Both coding systems were developed to address the same RQs, with this analytical focus explicitly incorporated into the ChatGPT prompt. Accordingly, the findings should be interpreted as a comparison between two analytical interpretations that identify both similarities and differences, rather than as a comparison against an objective standard.

5. Conclusions

This study aimed to develop an iterative, inductive coding system generated by ChatGPT, using detailed prompts grounded in recognized, proven strategies. This system was then compared, using both quantitative and qualitative methods, with the coding system developed through content analysis to identify similarities and differences and elaborate on striking results. These findings could, in turn, serve as a basis for further evaluation and optimization of the prompts used to develop a coding system with LLMs. Overall, the developed ChatGPT system shows a well-organized structure with many similar granularities of code, without redundancies, through this iterative process. Based on semantic comparison and qualitative evaluation, the ChatGPT code system is divided into three main codes with corresponding subcodes, directly addressing the three RQs. To further validate the semantic mapping, the Jaccard index was calculated at the level of the underlying coding units. The goal was to verify whether semantically mapped code pairs could indeed be traced to the same text segments as their corresponding codes in the content analysis. Our results show partial agreement: For five of the 13 semantically mapped code pairs, a match between the underlying coding units could be demonstrated. At the same time, differences were observed: conceptual differentiation, as evidenced by EV centrality in network analysis, and a lack of domain-specific knowledge to adequately address the RQs. Finally, critical reflection of the output is necessary, since LLMs may hallucinate. The following key learnings may be used to evaluate the prompting framework further while testing it in larger and different research contexts:
  • Key learning I: Methodological robustness emerges through iteration and structured prompting
Five separate ChatGPT coding systems were generated and independently applied to the dataset. Across the repeated coding of each transcript, an average Jaccard Index of 0.83 indicated substantial overlap in coding-unit assignments. These findings are consistent with the proposition that iterative prompting and repeated validation contribute to methodological robustness and thematic convergence (Gao et al., 2025; Kwon et al., 2025). This aligns with the idea of iterative refinement approximating ensemble coding logic (Jiang et al., 2021). Furthermore, these findings support conceptualizing ChatGPT as a probabilistic yet stabilizable analytical collaborator. Accordingly, methodological robustness should not be regarded as an inherent property of LLM-assisted coding but as a characteristic that can be strengthened through structured prompting, iterative refinement, and repeated validation.
  • Key learning II: Differentiation and contextualization differ in the ChatGPT code system
In the content analysis network, geography-specific factors are integrated into the broader structure and linked to multiple other categories, indicating that subject-specific contextualization is integral to the analytical framework. In the ChatGPT-based system, comparable domain-specific elements are largely absent from the central structure. The network analysis further suggests that ChatGPT-generated coding systems tend to prioritize generalizable patterns over fine-grained conceptual distinctions. In contrast, the content analysis network exhibits a broader and more distributed core structure, with several codes showing comparatively high EV-Centrality values. These findings suggest that LLM-assisted coding can efficiently identify overarching thematic structures, while contextualization and conceptual differentiation remain central strengths of content analysis. This might be improved by refining and testing the prompting framework on larger datasets.

Author Contributions

Conceptualization, L.D. & R.A.; methodology, L.D. & R.A.; software, L.D.; validation, L.D.; formal analysis, L.D.; data curation, L.D.; writing—original draft preparation, L.D.; visualization, L.D.; supervision, R.A.; project administration, L.D. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

This study used fully anonymized interview transcripts originally collected by the author for a prior educational research project. No identifying information is present, and no re-identification is possible. As the current analysis involves no interaction with individuals and no personal data, the study does not involve human subjects according to the ICMJE ethical definitions, the Belmont Report, and the Declaration of Helsinki.

Informed Consent Statement

Informed consent was obtained from all subjects involved in the initial interviews.

Data Availability Statement

Data, including R scripts and raw data for all analyses, are available via Zenodo (Dörfel & Ammoneit, 2026): https://doi.org/10.5281/zenodo.21108082.

Acknowledgments

During the preparation of this study, the authors used ChatGPT (GPT-5.2, 11 December 2025 snapshot; OpenAI) to generate qualitative coding outputs that were subsequently analyzed and evaluated as part of the research. All AI-generated outputs were critically reviewed, validated, and interpreted by the authors, who take full responsibility for the content of this publication.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
LLMLarge language model
GenAIGenerative Artificial Intelligence
NLPNatural Language Processing
SBERTSentence Bidirectional Encoder Representations from Transformers
EVEigenvector
RQResearch Question
JJaccard Index

Appendix A

Appendix A.1. Prompt for Developing Code Systems for Each Transcript

Description: Prompt used to generate an inductive coding system for each transcript following the principles of Kuckartz’s qualitative content analysis. The prompt instructed ChatGPT to identify relevant text segments on stereotyping, develop mutually exclusive categories inductively, and provide concise code names and descriptions. This prompt was used eleven times—for each transcript once.
Figure A1. Prompt for developing a code system for each transcript.
Figure A1. Prompt for developing a code system for each transcript.
Education 16 01314 g0a1

Appendix A.2. Prompt for Merging Eleven Code Systems into One Integrated Code System

Description: This appendix includes the prompt to develop an integrated and traceable coding system for stereotyping out of the eleven code systems generated in Appendix A.1. The model was instructed to iteratively compare and merge eleven coding systems derived from individual transcripts, forming a comprehensive, hierarchical, and data-driven code structure. The procedure required stepwise comparison, systematic resolution of thematic overlaps, and the creation of overarching categories where conceptually justified. Only codes that were applied at least once were permitted. No predefined theoretical framework was used; the resulting coding system is fully grounded in the empirical material and serves as the analytical basis for addressing the study’s research questions.
Figure A2. Prompt for merging 11 code systems into one integrated code system.
Figure A2. Prompt for merging 11 code systems into one integrated code system.
Education 16 01314 g0a2

Appendix A.3. Prompt for Merging the Five Merged Code Systems into a Final Code System

Description: The final code system includes only those codes that appear in at least three of the five systems.
Figure A3. Prompt for merging.
Figure A3. Prompt for merging.
Education 16 01314 g0a3

Appendix A.4. Prompt for Coding with Code System

Description: This appendix presents the prompt used to inductively code German-language interview transcripts, following Kuckartz’s approach to qualitative content analysis. It specifies the coder role, the analytical focus on stereotyping, and the alignment with the study’s two research questions. The prompt defines strict reliability requirements by restricting coding to a predefined coding system and prohibiting interpretation beyond code definitions, rewording, or the creation of new codes. In addition, the prompt details the unitization procedure (coding units labeled “T CU1,” “T CU2,” etc.), explicit coding rules (including the possibility of multiple codes or none per unit), and the required documentation of each coding decision through manual-referenced justifications. Finally, it sets out a fixed four-column table format and mandatory formatting constraints, including the requirement to deliver results as a structured Word (.docx) table without additional commentary.
Figure A4. Prompt for coding with the code system.
Figure A4. Prompt for coding with the code system.
Education 16 01314 g0a4

Appendix A.5. Content Analysis—Short Code System (Main- and Subcodes)

Description: Human code system by qualitative content analysis.
Note: The frequencies reported in this table differ from those in the original article on content analysis because, unlike the original study, we only counted the occurrence of a code once per coding unit. In the original article, there was no such restriction, and individual passages within coding units were analyzed separately. For the sake of clarity, however, in this article, all codings of a given code within a coding unit were capped at a maximum value of 1 in order to make them more comparable to ChatGPT’s coding.
Table A1. Code system of content analysis.
Table A1. Code system of content analysis.
Main CodesSubcodesCount (n)
1. Examples of Stereotyping (n = 35)1.1 Continents4
1.2 Country groups3
1.3 Students’ background7
1.4 Countries8
1.5 Subject content13
2. Causes of Stereotyping (n = 30)2.1 Teachers (causes)2
2.2 Educational policy factors2
2.3 Geography-specific factors7
2.4 Students (causes)8
2.5 Sociocultural influences5
2.6 Resources: Textbooks6
3. Strategies for addressing stereotyping (n = 42)3.1 Perspective shift6
3.2 Causal analysis3
3.3 Student-oriented strategies4
3.4 Teachers (strategies)3
3.5 Teaching methods6
3.6 Outside the classroom2
3.7 Dialogue and reflection7
3.8 Increasing complexity5
3.10 Evidence-based working5
Total main codes (n = 3)Total Subcodes (n = 20)Total occurrence (n = 106)

References

  1. Alhindi, T., Chakrabarty, T., Musi, E., & Muresan, S. (2022). Multitask instruction-based prompting for fallacy recognition. In Y. Goldberg, Z. Kozareva, & Y. Zhang (Eds.), Proceedings of the 2022 conference on empirical methods in natural language processing (pp. 8172–8187). Association for Computational Linguistics. [Google Scholar] [CrossRef] [Scilit]
  2. Ayala, O., & Bechard, P. (2024). Reducing hallucination in structured outputs via retrieval-augmented generation. In Proceedings of the 2024 conference of the North American chapter of the association for computational linguistics: Human language technologies (Volume 6: Industry track) (pp. 228–238). Association for Computational Linguistics. [Google Scholar] [CrossRef] [Scilit]
  3. Bedru, H. D., Yu, S., Xiao, X., Zhang, D., Wan, L., Guo, H., & Xia, F. (2020). Big networks: A survey. Computer Science Review, 37, 100247. [Google Scholar] [CrossRef] [Scilit]
  4. Ben-David, E., Oved, N., & Reichart, R. (2022). PADA: Example-based prompt learning for on-the-fly adaptation to unseen domains. Transactions of the Association for Computational Linguistics, 10, 414–433. [Google Scholar] [CrossRef] [Scilit]
  5. Bloch, F., Jackson, M. O., & Tebaldi, P. (2023). Centrality measures in networks. Social Choice and Welfare, 61(2), 413–453. [Google Scholar] [CrossRef] [Scilit]
  6. Braun, V., & Clarke, V. (2006). Using thematic analysis in psychology. Qualitative Research in Psychology, 3(2), 77–101. [Google Scholar] [CrossRef] [Scilit]
  7. Catalano, A., & Mantegna, R. N. (2026). How null-model constraints affect statistical validation in projected bipartite networks. arXiv. [Google Scholar] [CrossRef] [Scilit]
  8. Christou, P. (2023). How to use artificial intelligence (AI) as a resource, methodological and analysis tool in qualitative research? The Qualitative Report, 28(7), 1968–1980. [Google Scholar] [CrossRef] [Scilit]
  9. Christou, P. (2024). Thematic analysis through artificial intelligence (AI). The Qualitative Report, 29(2), 560–576. [Google Scholar] [CrossRef] [Scilit]
  10. Csárdi, G., Nepusz, T., Müller, K., Horvát, S., Traag, V., Zanini, F., & Noom, D. (2026). igraph for R: R interface of the igraph library for graph theory and network analysis [Computer software]. Zenodo. [Google Scholar]
  11. Dahal, N. (2024). How can generative AI (GenAI) enhance or hinder qualitative studies? A critical appraisal from South Asia, Nepal. The Qualitative Report, 29(3), 722–733. [Google Scholar] [CrossRef] [Scilit]
  12. Dai, W., Lin, J., Jin, F., Li, T., Tsai, Y.-S., Gasevic, D., & Chen, G. (2023, July 10–13). Can large language models provide feedback to students? A case study on ChatGPT. 2023 IEEE International Conference on Advanced Learning Technologies (ICALT), Orem, UT, USA. [Google Scholar] [CrossRef] [Scilit]
  13. Dörfel, L., & Ammoneit, R. (2026). Inductive coding with LLMs (Version 1.0) [Dataset]. Zenodo. [Google Scholar] [CrossRef]
  14. Dörfel, L., Ammoneit, R., & Peter, C. (2025a). Stereotyping in German geography classes—Secondary teachers’ challenges and strategies. European Journal of Geography, 16(2), 141–155. [Google Scholar] [CrossRef] [Scilit]
  15. Dörfel, L., Ammoneit, R., & Peter, C. (2025b). Supplementary material for: Stereotyping in German geography classes—Secondary teachers’ challenges and strategies [Supplementary material]. European Journal of Geography. Available online: https://eurogeojournal.eu/index.php/egj/article/view/773/435 (accessed on 10 August 2026).
  16. Dröge, K. (2023). noScribe (Version 0.4.1) [Computer software]. Available online: https://github.com/kaixxx/noScribe (accessed on 10 August 2026).
  17. Farquhar, S., Kossen, J., Kuhn, L., & Gal, Y. (2024). Detecting hallucinations in large language models using semantic entropy. Nature, 630(8017), 625–630. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  18. Gao, J., Shu, Z., Yeo, S. Y., Prakash, A., Huang, C.-M., Dredze, M., & Xiao, Z. (2025). Efficiency with rigor! A trustworthy LLM-powered workflow for qualitative data analysis. arXiv. [Google Scholar] [CrossRef] [Scilit]
  19. Ghojogh, B., & Ghodsi, A. (2020). Attention mechanism, transformers, BERT, and GPT: Tutorial and survey. HAL Open Science. Available online: https://hal.science/hal-04637647v1 (accessed on 10 August 2026).
  20. Goswami, K., Karanam, S., Udhayanan, P., Joseph, K. J., & Srinivasan, B. V. (2024). CoPL: Contextual prompt learning for vision-language understanding. Proceedings of the AAAI Conference on Artificial Intelligence, 38(16), 18090–18098. [Google Scholar] [CrossRef] [Scilit]
  21. Goyanes, M., Lopezosa, C., & Jordá, B. (2025). Thematic analysis of interview data with ChatGPT: Designing and testing a reliable research protocol for qualitative research. Quality & Quantity, 59(6), 5491–5510. [Google Scholar] [CrossRef] [Scilit]
  22. Greve, W., & Wentura, D. (1997). Wissenschaftliche beobachtung: Eine einführung. Beltz. [Google Scholar]
  23. Han, S., Tan, T., Miao, Y., Chen, X., & Sun, N. (2026). Prompting instability: An empirical study of LLM robustness in code vulnerability detection. In M. Liu, X. Yu, C. Xu, & Y. Song (Eds.), AI 2025: Advances in artificial intelligence (Vol. 16370, pp. 233–245). Lecture Notes in Computer Science. Springer Nature. [Google Scholar] [CrossRef] [Scilit]
  24. Henderson, M., Bearman, M., Chung, J., Fawns, T., Buckingham Shum, S., Matthews, K. E., & de Mello Heredia, J. (2026). Comparing generative AI and teacher feedback: Student perceptions of usefulness and trustworthiness. Assessment & Evaluation in Higher Education, 51(5), 863–878. [Google Scholar] [CrossRef] [Scilit]
  25. Huang, S.-Y., Lin, Y.-J., Sheu, S.-J., & Chu, W.-M. (2026). Using large language models to assist qualitative thematic analysis of student reflections on advance care planning education. BMC Medical Education, 26(1), 674. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  26. Jaccard, P. (1901). Étude comparative de la distribution florale dans une portion des Alpes et des Jura. Bulletin de la Société Vaudoise des Sciences Naturelles, 37(142), 547–579. [Google Scholar]
  27. Jiang, J. A., Wade, K., Fiesler, C., & Brubaker, J. R. (2021). Supporting serendipity: Opportunities and challenges for human-AI collaboration in qualitative analysis. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1), 94. [Google Scholar] [CrossRef] [Scilit]
  28. Kalai, A. T., Nachum, O., Vempala, S. S., & Zhang, E. (2026). Evaluating large language models for accuracy incentivizes hallucinations. Nature, 653(8116), 1047–1051. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  29. Kang, T. W., & Mo, Y. (2024). A comprehensive digital twin framework for building environment monitoring with emphasis on real-time data connectivity and predictability. Developments in the Built Environment, 17, 100309. [Google Scholar] [CrossRef] [Scilit]
  30. Khalid, M. T., & Witmer, A.-P. (2026). Prompt engineering for large language model-assisted inductive thematic analysis. Social Science Computer Review, 44(3), 442–465. [Google Scholar] [CrossRef] [Scilit]
  31. Kong, X., Shi, Y., Yu, S., Liu, J., & Xia, F. (2019). Academic social networks: Modeling, analysis, mining and applications. Journal of Network and Computer Applications, 132, 86–103. [Google Scholar] [CrossRef] [Scilit]
  32. Kuckartz, U., & Rädiker, S. (2022). Qualitative inhaltsanalyse: Methoden, praxis, computerunterstützung (5th ed.). Beltz Juventa. [Google Scholar]
  33. Kuhn, H. W. (1955). The Hungarian method for the assignment problem. Naval Research Logistics Quarterly, 2(1–2), 83–97. [Google Scholar] [CrossRef] [Scilit]
  34. Kwon, H., Lu, L., Kang, J., & Mcleod, D. (2025). Leveraging the power of ChatGPT: Evaluating its effectiveness for content analysis and framing research in mass communication. In T. Bui (Ed.), Proceedings of the annual Hawaii international conference on system sciences—Proceedings of the 57th Hawaii international conference on system sciences. Hawaii International Conference on System Sciences. [Google Scholar] [CrossRef] [Scilit]
  35. Leifheit, L., Loefflad, D., Belschner, S., Beuttler, B., Winkelmann, J., Meurers, W. D., & Holz, H. (2024). KI im unterricht. Ludwigsburger Beiträge zur Medienpädagogik, 24, 1–19. [Google Scholar] [CrossRef] [Scilit]
  36. Liu, Z. (2025). A state-update prompting strategy for efficient and robust multi-turn dialogue. arXiv. [Google Scholar] [CrossRef] [Scilit]
  37. Lixandru, I.-D. (2024). The use of artificial intelligence for qualitative data analysis: ChatGPT. Informatica Economica, 28(1), 57–67. [Google Scholar] [CrossRef] [Scilit]
  38. Louatouate, H., & Zeriouh, M. (2025). Role-based prompting technique in generative AI-assisted learning: A student-centered quasi-experimental study. Journal of Computer Science and Technology Studies, 7(2), 130–145. [Google Scholar] [CrossRef] [Scilit]
  39. Luis, S. Y., Reina, D. G., & Marín, S. T. (2025). Towards a retrieval-augmented generation framework for originality evaluation in projects-based learning classrooms. Education Sciences, 15(6), 706. [Google Scholar] [CrossRef] [Scilit]
  40. Ma, W., Yang, Y., Ge, J., Xie, X., & Jiang, L. (2025). Prompt stability in code LLMs: Measuring sensitivity across emotion- and personality-driven variations. arXiv. [Google Scholar] [CrossRef] [Scilit]
  41. MDPI. (n.d.). Instructions for authors. Available online: https://www.mdpi.com/journal/ai/instructions#ethics (accessed on 10 August 2026).
  42. Morgan, D. L. (2023). Exploring the use of artificial intelligence for qualitative data analysis: The case of ChatGPT. International Journal of Qualitative Methods, 22, 16094069231211248. [Google Scholar] [CrossRef] [Scilit]
  43. Muasher-Kerwin, C., Hughes, M. C., & Econie, S. M. (2026). Using large language models to support qualitative health research through coding, theming, quotation selection, and data visualization. Discover Artificial Intelligence, 6(1), 357. [Google Scholar] [CrossRef] [Scilit]
  44. Myers, D., Mohawesh, R., Chellaboina, V. I., Sathvik, A. L., Venkatesh, P., Ho, Y.-H., Henshaw, H., Alhawawreh, M., Berdik, D., & Jararweh, Y. (2024). Foundation and large language models: Fundamentals, challenges, opportunities, and social impacts. Cluster Computing, 27(1), 1–26. [Google Scholar] [CrossRef] [Scilit]
  45. Naeem, M., Smith, T., & Thomas, L. (2025). Thematic analysis and artificial intelligence: A step-by-step process for using ChatGPT in thematic analysis. International Journal of Qualitative Methods, 24, 16094069251333886. [Google Scholar] [CrossRef] [Scilit]
  46. Newman, M., Barabási, A.-L., & Watts, D. J. (2006). The structure and dynamics of networks. Princeton University Press. [Google Scholar]
  47. Oksanen, J., Simpson, G. L., Blanchet, F. G., Kindt, R., Legendre, P., Minchin, P. R., O’Hara, R. B., Solymos, P., Stevens, M. H. H., Szoecs, E., Wagner, H., Bedward, M., Bolker, B., Borcard, D., Carvalho, G., de Caceres, M., Durand, S., Evangelista, H. B. A., Hannigan, G., … Weedon, J. (2026). CRAN: Contributed packages (Version 2.7-5) [Computer software]. CRAN. [Google Scholar]
  48. OpenAI. (2025). ChatGPT (Version 5.2 Dec 11 version) [Computer software]. Available online: https://chatgpt.com (accessed on 11 February 2026).
  49. Peters, U., & Chin-Yee, B. (2025). Generalization bias in large language model summarization of scientific research. Royal Society Open Science, 12(4), 241776. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  50. R Core Team. (2026). R: A language and environment for statistical computing (Version 4.6.1) [Computer software]. R Foundation for Statistical Computing. [Google Scholar]
  51. Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In K. Inui, J. Jiang, V. Ng, & X. Wan (Eds.), Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP) (pp. 3980–3990). Association for Computational Linguistics. [Google Scholar] [CrossRef] [Scilit]
  52. Saadaoui, S. (2026). Semantic fidelity in specialized domains: Advancing language models through adaptive learning, collective reasoning, and consensus evaluation [Doctoral dissertation, University of London]. Available online: https://openaccess.city.ac.uk/id/eprint/37220/ (accessed on 10 August 2026).
  53. Saldaña, J. (2025). The coding manual for qualitative researchers (5th ed.). SAGE. [Google Scholar]
  54. Salton, G., Wong, A., & Yang, C. S. (1975). A vector space model for automatic indexing. Communications of the ACM, 18(11), 613–620. [Google Scholar] [CrossRef] [Scilit]
  55. Schueller, T., Trettin, A., & Huber, S. (2026). Evaluating AI-assisted deductive coding in MAXQDA: A methodological analysis of inputs and outputs. International Journal of Qualitative Methods, 25, 16094069251407046. [Google Scholar] [CrossRef] [Scilit]
  56. Stelea, G. A., Robu, D., & Sandu, F. (2025). AccessiLearnAI: An accessibility-first, AI-powered e-learning platform for inclusive education. Education Sciences, 15(9), 1125. [Google Scholar] [CrossRef] [Scilit]
  57. Strona, G., Nappo, D., Boccacci, F., Fattorini, S., & San-Miguel-Ayanz, J. (2014). A fast and unbiased procedure to randomize ecological binary matrices with fixed row and column totals. Nature Communications, 5, 4114. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  58. Trott, S., Jones, C., Chang, T., Michaelov, J., & Bergen, B. (2023). Do large language models know what humans know? Cognitive Science, 47(7), e13309. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  59. Turobov, A., Coyle, D., & Harding, V. (2024). Using ChatGPT for thematic analysis. arXiv. [Google Scholar] [CrossRef] [Scilit]
  60. Vaidya, S., & Saini, J. R. (2026). Consistency evaluation protocol: A reproducible framework for assessing large language model output repeatability. MethodsX, 17, 104069. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  61. VERBI Software. (2024). MAXQDA (Version 24.4) [Computer software]. VERBI Software. Available online: https://www.maxqda.com/ (accessed on 14 May 2025).
  62. Wang, G., Zhan, Z., & Qin, S. (2025). Synergizing knowledge graphs and LLMs: An intelligent tutoring model for self-directed learning. Education Sciences, 15(9), 1102. [Google Scholar] [CrossRef] [Scilit]
  63. Wickham, H. (2016). ggplot2. Springer International Publishing. [Google Scholar] [CrossRef] [Scilit]
  64. Xiao, Z., Yuan, X., Liao, Q. V., Abdelghani, R., & Oudeyer, P.-Y. (2023). Supporting qualitative analysis with large language models: Combining codebook with GPT-3 for deductive coding. In Companion proceedings of the 28th international conference on intelligent user interfaces (pp. 75–78). ACM. [Google Scholar] [CrossRef] [Scilit]
  65. Xu, W. (2026). Doing thematic analysis in the age of generative AI: Practices, ethics and reflexivity. International Journal of Qualitative Methods, 25, 16094069261425173. [Google Scholar] [CrossRef] [Scilit]
  66. Yue, Y., Liu, D., Lv, Y., Hao, J., & Cui, P. (2025). A practical guide and assessment on using ChatGPT to conduct grounded theory: Tutorial. Journal of Medical Internet Research, 27, e70122. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  67. Zeng, Q., Jin, C., Wang, X., Zheng, Y., & Li, Q. (2025). AIRepr: An analyst-inspector framework for evaluating reproducibility of LLMs in data science. In C. Christodoulopoulos, T. Chakraborty, C. Rose, & V. Peng (Eds.), Findings of the association for computational linguistics: EMNLP 2025 (pp. 10170–10201). Association for Computational Linguistics. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Overview of inductive coding procedures.
Figure 1. Overview of inductive coding procedures.
Education 16 01314 g001
Figure 2. Research design, including research questions and methods.
Figure 2. Research design, including research questions and methods.
Education 16 01314 g002
Figure 3. Heatmap based on cosine similarities.
Figure 3. Heatmap based on cosine similarities.
Education 16 01314 g003
Figure 4. Network of Content Analysis Coding (EV-based) (blue = examples, orange = reasons, green = strategies).
Figure 4. Network of Content Analysis Coding (EV-based) (blue = examples, orange = reasons, green = strategies).
Education 16 01314 g004
Figure 5. Network of ChatGPT Coding (EV-based) (blue = manifestations, orange = conditions of emergence, green = strategies).
Figure 5. Network of ChatGPT Coding (EV-based) (blue = manifestations, orange = conditions of emergence, green = strategies).
Education 16 01314 g005
Table 1. Template for the output of initial development.
Table 1. Template for the output of initial development.
Code NameCode Description
 
Table 2. Template for the merging output.
Table 2. Template for the merging output.
Code with DescriptionMerged Codes Documentation
 
Table 3. Template for the Coding output.
Table 3. Template for the Coding output.
Coding Unit Coded Passage Assigned Code & JustificationNumber of Codes per CU
 
Table 4. Final code system generated by ChatGPT.
Table 4. Final code system generated by ChatGPT.
Main CodeSubcodeDescription
Manifestations Spatial and developmental stereotypesRegions, countries, or continents are assigned sweeping characteristics that portray them as poor, backward, or homogeneous. These representations reduce complex socioeconomic and political structures to simplified narratives. They often appear as objective descriptions but are based on historically developed, socially transmitted patterns of interpretation.
Problem and deficit focusCertain regions or groups are defined primarily through crises, poverty, conflicts, or exploitation. Positive developments, diversity, or everyday normality remain underrepresented. This creates a distorted overall picture that reduces regions or societies to problem situations.
Social and origin stereotypesIndividuals or groups are categorized based on ethnicity, nationality, social background, or ability. These attributions are based on generalized assumptions about abilities, lifestyles, or educational attainment and reproduce social hierarchies.
Migration and religionMigration is reduced to individual motives—particularly flight—or presented in a quantitatively distorted manner. Similarly, religious groups are reduced to a few characteristics, obscuring their internal diversity. Both areas demonstrate how complex oversimplified images replace social processes.
Conditions of emergenceDidactic reduction and time pressureTeaching necessarily requires simplification and structuring. A vast amount of material, exemplary approaches, and limited class time lead to complex relationships being presented in a simplified manner. These structural conditions can foster stereotypical generalizations if differentiation is omitted due to time or structural constraints.
Textbooks, cartoons, images, and curricula convey implicit narratives and perspectives. If these materials are used without reflection, stereotypical representations can be reproduced. This reproduction often occurs unintentionally and is mediated by structure.
Material and textbook influenceTextbooks, cartoons, images, and curricula convey implicit narratives and perspectives. If these materials are used without critical reflection, they can perpetuate stereotypical portrayals. This perpetuation is often unintentional and is structurally embedded.
Knowledge deficits as a breeding groundMissing or fragmentary prior knowledge makes differentiated analyses of complex interrelations more difficult. In such cases, learners—and partly also teachers—resort to simplifying explanatory models. Stereotypes function here as a substitute for differentiated subject knowledge.
EurocentrismEurocentric perspectives implicitly serve as a normative standard for evaluation and comparison. From this perspective, other regions of the world appear deficient or deviant. This hierarchical framing reinforces stereotypical perceptions and shapes global perceptions of inequality.
Pedagogical strategies in dealing with itDiagnostic assessmentAt the beginning of a lesson, existing ideas are systematically identified to make implicit images visible. This diagnostic phase serves as the basis for further didactic work and prevents stereotypical assumptions from persisting uncritically.
Dialogue-based reflection and interventionStereotypical statements are addressed, questioned, and discussed as a class. Teachers intervene when students make sweeping generalizations, ask for clarification, or introduce alternative perspectives. This dialogue gradually deconstructs oversimplified assumptions.
Perspective shift and irritationStereotypical patterns of thinking are disrupted through role-playing, provocative exaggerations, or visual reversals. These teaching methods aim to challenge taken-for-granted assumptions and trigger processes of reflection.
Work with data and contextualizationStereotypical statements are examined using empirical data, statistical information, or in-depth research. By placing them in context, we avoid monocausal explanations and reveal complex interrelationships.
Making diversity and differentiation visibleThe internal diversity of countries, regions, or groups is systematically addressed in order to challenge homogenization. The complexity of social reality is illustrated through multi-perspective analyses.
Professional self-reflectionTeachers reflect on their own mental models and potential stereotypical assumptions. This self-reflection is understood as a prerequisite for professional practice to avoid unintentional reproduction and actively curb the reinforcement of stereotypes.
Table 5. Run-to-run stability for each transcript.
Table 5. Run-to-run stability for each transcript.
TranscriptRun-to-Run Stability (Jaccard)
10.88
20.77
30.79
40.84
50.91
60.91
70.83
80.78
90.86
100.87
110.66
Total0.83
Table 6. One-to-One Mapping based on semantic similarities.
Table 6. One-to-One Mapping based on semantic similarities.
No.ChatGPT CodesContent Analysis CodesCosine Sim.
1Strategies—Dialogue-based reflection and intervention (n = 7)Strategies—Dialogue and reflection (n = 7)0.81
2Strategies—Perspective shift and irritation (n = 5)Strategies—Teaching methods (n = 6)0.78
3Strategies—Making diversity and differentiation visible (n = 5)Strategies—Increasing complexity (n = 5)0.77
4Strategies—Professional self-reflection (n = 4)Strategies—Teachers (s) (n = 3)0.77
5Strategies—Diagnostic assessment (n = 6)Strategies—Outside the classroom (n = 2)0.73
6Strategies—Work with data and contextualization (n = 4)Strategies—Evidence-based learning (n = 5)0.72
7Conditions of emergence—Knowledge deficits as breeding ground (n = 3)Reasons—Education policy factors (n = 2)0.71
8Manifestations—Spatial and development stereotypes (n = 10)Examples—Continents (n = 4)0.69
9Manifestations—Social and origin stereotypes (n = 11)Examples—Students’ background (n = 7)0.66
10Conditions of emergence—Didactic reduction and time pressure (n = 5)Reasons—Geography-specific factors (n = 7)0.65
11Conditions of emergence—Material and textbook influence (n = 7)Reasons—Resources: Textbooks (n = 6)0.63
12Manifestations—Migration and religion (n = 5)Reasons—Sociocultural influences (n = 5)0.61
13Conditions of emergence—Eurocentrism (n = 2)Strategies—Perspective shift (n = 6)0.6
14Manifestations—Problem- and deficit focus (n = 7)Examples—Country groups (n = 3)0.56
n = 81 codesn = 65 codes *
* As the table only reports the frequencies (n) of the mapped codes, the total n is lower (n = 65) than in Appendix A Table A1, which reports the overall n across all 20 subcodes (n = 106).
Table 7. Jaccard Indices for semantically mapped codes.
Table 7. Jaccard Indices for semantically mapped codes.
ChatGPTContent AnalysisJaccard Index
Dialogue-based reflection and interventionDialogue and reflection0.66
Perspective shift and irritationTeaching methods0.1
Making diversity and differentiation visibleIncreasing complexity0.66
Professional self-reflectionTeachers (strategies)0.75
Diagnostic assessmentOutside the classroom0
Working with data and contextualizationEvidence-based working0.8
Knowledge deficits as breeding groundEducational policy factors0
Spatial and developmental stereotypesContinents0.33
Social and origin stereotypesStudents’ background0.38
Complexity reduction and time pressureGeography-specific factors0.33
Material and textbook influenceResources: Textbooks0.86
EurocentrismPerspective shift 0.14
Migration and religionSocio-cultural influences0
Global Jaccard index0.39
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Dörfel, L.; Ammoneit, R. Evaluation of Inductive Coding with LLMs. Educ. Sci. 2026, 16, 1314. https://doi.org/10.3390/educsci16081314

AMA Style

Dörfel L, Ammoneit R. Evaluation of Inductive Coding with LLMs. Education Sciences. 2026; 16(8):1314. https://doi.org/10.3390/educsci16081314

Chicago/Turabian Style

Dörfel, Leoni, and Rieke Ammoneit. 2026. "Evaluation of Inductive Coding with LLMs" Education Sciences 16, no. 8: 1314. https://doi.org/10.3390/educsci16081314

APA Style

Dörfel, L., & Ammoneit, R. (2026). Evaluation of Inductive Coding with LLMs. Education Sciences, 16(8), 1314. https://doi.org/10.3390/educsci16081314

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop