1. Introduction
The integration of generative artificial intelligence (GenAI) into educational settings has sparked significant interest in its transformative potential; not only for teaching and learning (
Luis et al., 2025;
Stelea et al., 2025;
Wang et al., 2025), but increasingly for research practices themselves (
Dahal, 2024;
Goyanes et al., 2025;
Henderson et al., 2026;
Schueller et al., 2026). While recent studies emphasize the utility of LLMs like ChatGPT for tasks such as feedback generation, formative assessment, and knowledge summarization (
Dai et al., 2023), their potential as tools for qualitative data analysis remains underexplored, particularly in educational research.
Qualitative coding, the interpretative process of assigning meaning to textual data, is a cornerstone of empirical educational research (
Braun & Clarke, 2006;
Saldaña, 2025). This process is labor-intensive, and transformer-based LLMs like ChatGPT (
OpenAI, 2025) have demonstrated promising capabilities for qualitative tasks, including coding, thematic analysis, and synthesis of unstructured data (
Christou, 2024;
Gao et al., 2025;
Goyanes et al., 2025;
Kwon et al., 2025;
Xu, 2026). Their capacity to generate coding frameworks, apply pre-existing codebooks, and assist in thematic abstraction suggests potential applications in research areas traditionally dominated by human expertise (
Xiao et al., 2023). Despite their capacity for pattern recognition and language generation, LLMs like ChatGPT often produce factual inaccuracies and hallucinations, including confidently fabricated information (
Dahal, 2024;
Henderson et al., 2026), raising ethical concerns (
Naeem et al., 2025). Recent efforts to mitigate hallucinations have explored various strategies, including output validation (e.g., Retrieval-Augmented Generation;
Ayala & Bechard, 2024), real-time verification frameworks (
Kang & Mo, 2024), and input sophistication (
Morgan, 2023), primarily through prompt engineering.
Studies show that prompt structure significantly affects model outputs, with better-engineered prompts leading to improved reliability and thematic consistency (
Khalid & Witmer, 2026). Prompting formats that have proven particularly relevant include: (1) instruction-based prompting that relies on explicit directives, such as tasks or checklists (
Alhindi et al., 2022), (2) role-based prompting that defines a target persona (e.g., “You are a scientific editor”), helping modulate tone, expertise, or evaluative stance, and dismantling hallucinations (
Louatouate & Zeriouh, 2025), (3) contextual prompting that embeds background material into the prompt to guide response grounding (
Goswami et al., 2024), (4) multi-turn prompting that structures the interaction dialogically across multiple steps, allowing for clarification, refinement, or reflection, and preventing overload (
Liu, 2025), (5) example-based prompting that introduces model responses as templates to support output alignment (
Ben-David et al., 2022), and (6) chain-of-thought prompting that breaks down tasks into step-by-step sequences, encouraging the model to make intermediate inferences and thereby improving accuracy and interpretability in complex analytical tasks (
Kwon et al., 2025). As
Gao et al. (
2025) demonstrate, LLMs can also support collaborative coding workflows, serving not as replacements but as aids that help human analysts refine and negotiate code (
Christou, 2023,
2024).
Given these advancements in prompt engineering and its role in reducing model hallucinations, ensuring contextual grounding, and maintaining systematic coding logic, we propose an inductive coding framework for structuring complex interview data generated by ChatGPT. Each prompt design element (instruction, role, context, iteration, synthesis) translates theoretical insights into stepwise coding procedures. To evaluate the framework, we use both qualitative and quantitative methods. Specifically, we compare the results to a prior qualitative content analysis of the same material, which already included the development of a coding system and the coding process. In addition to qualitative analysis, we apply quantitative methods to extend the analysis by enabling systematic measurement, statistical comparison, and the examination of observable patterns across both approaches.
Research aims: Our empirical case draws on 11 qualitative interviews with geography teachers on stereotyping in the classroom, an interpretively rich topic that requires contextual sensitivity. Based on theoretically derived prompting strategies (instruction, role, context, multi-turn prompting, examples, and chain-of-thought prompting), we developed a prompting framework for inductive coding with ChatGPT. We then systematically compare this code system and the coding of 11 transcripts against those derived from qualitative content analysis. The primary aim is to provide a systematic, structural, and semantic comparison between the two coding systems, thereby identifying areas of convergence and divergence. Explicitly, we compare the outcome of theoretically derived strategies with human-coded approaches by addressing the following Research Questions (RQs):
RQ1: Can we map the ChatGPT code system onto the content analysis system based on semantic similarities?
RQ2: Do semantically corresponding codes refer to the same content represented by the coding units?
RQ3: Are there systematic differences between the two code systems?
3. Results
3.1. Development of the ChatGPT Code System and Coding
Following the development and multiple syntheses of the ChatGPT-generated coding system (Phases I–IV), the characteristics of the five synthesized coding systems were examined. Although the five synthesized coding systems showed a high degree of thematic consistency, some structural differences were observed regarding the number and organization of codes. The number of subcodes ranged from 16 to 24. Trial 3 contained the largest number of subcodes (n = 24), whereas Trials 1 and 5 each comprised 16 subcodes, Trial 2 comprised 18, and Trial 4 comprised 20. Except for Trial 4, all synthesized coding systems consisted of three main codes corresponding to RQ1-3. Trial 4 additionally included the higher-order code Perception and Evaluation, which occurred only once and was therefore excluded during the final synthesis.
To derive a robust final coding system (Phase IV), the five synthesized coding systems were merged. Only codes that appeared in at least three of the five synthesized versions were retained in the final coding system.
Table 4 presents the final synthesized coding system. Because the interviews were conducted in German, all coding procedures were performed in German. The final coding system was subsequently translated into English using DeepL (8 June 2026).
Following the development of the final coding system, each of the eleven transcripts was coded five times in independent chat sessions using the finalized coding framework (Phase V). To assess the stability of the coding outputs, the pairwise Jaccard index was calculated across the five coding runs for each transcript.
Table 5 presents the resulting Jaccard similarity values for measuring run-to-run stability. The overall Jaccard similarity across all transcripts was 0.83, indicating a high run-to-run stability throughout five runs.
3.2. Can We Map the ChatGPT Code System onto the Qualitative Content Analysis Based on Semantic Similarities?
The coding system for content analysis was developed using a deductive–inductive approach guided by two RQs. It comprises three main codes and 20 subcodes (
Appendix A.5). The system is structured around the RQs. It thus differentiates into three main codes: Examples of stereotyping, reasons for stereotyping, and strategies in dealing with stereotypes. The three main codes of the ChatGPT code system are: Forms of manifestation, conditions of emergence, and pedagogical strategies, which were then analyzed with 14 subcodes in total. For a mapping of both systems, we chose an NLP-based Similarity Analysis (SBERT) and created (a) a heatmap to show general similarities across all codes and (b) a one-to-one mapping.
3.2.1. NLP-Based Similarity Analysis (SBERT)
In total, 63 code pairings with cosine similarity scores ≥ 0.60 were identified. Overall, the heatmap reveals substantial semantic overlap between the ChatGPT-generated and manually developed coding systems, while also indicating that the strength of this overlap varies across individual codes (
Figure 3). In particular, codes such as Professional self-reflection, Didactic reduction and time pressure, Working with data and contextualization, and Perspective shift show high correspondence with codes from the content analysis, with cosine similarity scores exceeding 0.7. In contrast, other codes, most notably Migration and religion and Problem- and deficit focus, exhibit comparatively weaker semantic alignment with the content analysis code system. Beyond individual pairings, a broader structural pattern emerges in the extent of semantic linkage. Certain ChatGPT codes, especially Dialogue-based reflection and intervention and Perspective shift and irritation, demonstrate consistently high similarity across multiple content analysis codes, suggesting broader conceptual coverage.
3.2.2. One-to-One Mapping of ChatGPT Codes to Content Analysis Codes
In the next step, we applied a one-to-one mapping to map all 14 ChatGPT codes to the content analysis codes, in descending order of cosine similarity, using code names and descriptions as the similarity index (
Table 6). Overall, the mapping indicates that most ChatGPT codes could be assigned to semantically corresponding content analysis codes with moderate to high similarity, suggesting a substantial structural correspondence between the two coding systems. Only one match out of 14 (Problem- and deficit focus + Country groups) is below the 0.6 threshold and can therefore be considered to have lower semantic overlap. Therefore, the subsequent analyses refer to the remaining 13 mapped code pairs. Since the content analysis identified six more codes than ChatGPT, six codes go unmatched: Subject content and Countries as Examples of stereotyping, Teachers (causes) and Students (causes) as Reasons for stereotyping, and Causal analysis and Student-oriented strategies as Strategies in dealing with stereotypes.
3.3. Do Semantically Corresponding Codes Refer to the Same Content Represented by the Coding Units?
The one-to-one mapping forms the basis for comparing the coding unit assignments of semantically corresponding codes derived from ChatGPT and Content analysis. To quantify the overlap in coding unit assignments, the Jaccard Index was calculated for each pair of mapped codes and across all 13 mapped codes (global Jaccard index) (see
Table 7).
The resulting global Jaccard Index is 0.39, indicating limited overall overlap in coding-unit assignments between the semantically mapped ChatGPT and content-analysis codes. The degree of overlap varies substantially across the 13 mapped code pairs (J = 0–0.86). Five mapped code pairs show a comparatively high overlap (J > 0.60), including Dialogue and reflection and Dialogue-based reflection and intervention, Increasing complexity and Making diversity and differentiation visible, Teachers (strategies) and Professional self-reflection, Evidence-based working and Working with data and contextualization, as well as Resources: Textbooks and Material and textbook influence. For these code pairs, both coding approaches assigned semantically corresponding codes to largely the same coding units. In contrast, three mapped code pairs show low to moderate overlap (J = 0.30–0.40), whereas the remaining five code pairs exhibit very low overlap (J < 0.14).
3.4. Are There Systematic Differences Between the Two Code Systems?
To characterize the overall structure of the two coding systems, we calculated the network density, global clustering, and eigenvector centrality. The content analysis network has a density of 0.40 and a global clustering coefficient of 0.56, while the ChatGPT-based network has a density of 0.36 and a global clustering coefficient of 0.61. Compared to their respective null models, both networks do not differ in terms of network density (Content analysis: z = −0.60, p = 0.611; ChatGPT: z = −1.66, p = 0.154) or global clustering (Content analysis: z = 0.34, p = 0.683; ChatGPT: z = 1.19, p = 0.239), suggesting that the observed values for connectivity and clustering did not exceed those expected based solely on the coding structure.
Figure 4 and
Figure 5 provide a visual overview of the overall structure of both coding systems: In these network representations, nodes correspond to individual codes, node size reflects eigenvector centrality (larger circles indicate higher centrality, while smaller circles indicate lower centrality), edge thickness indicates the strength of code co-occurrence, and colors represent thematic domains. Both networks exhibit a densely interconnected core surrounded by more peripheral codes, although differences in the distribution of central nodes become apparent. Although the two coding systems showed similar global structural properties, EV-centrality indicates that structural importance is distributed differently across the networks. Whereas the content analysis network distributes centrality across several strongly connected codes, the ChatGPT-based network concentrates centrality in fewer dominant codes. In the content analysis network, Subject content represents the most central node (EV = 1.00), followed by Continents (0.71). In addition, several other codes—including Students (c), Resources: Textbooks, and Teaching methods—also show relatively high centrality values, forming a broader set of highly connected nodes. In the ChatGPT-based network, Spatial and developmental stereotypes emerge as the most central code (EV = 1.00), followed by Problem- and deficit focus (0.91) and Material and textbook influence (0.90).
Across both networks, a similar higher-order pattern emerged. Codes associated with reasons for stereotyping and examples of stereotyping tend to occupy more central positions and are connected to more other codes. In contrast, strategy-related codes generally show lower centrality values and fewer connections, indicating a more peripheral position within the overall network structure.
3.5. Qualitative Evaluation
The findings are reported according to the analytical dimensions.
The code system generated by ChatGPT was output in tabular form and matches the requested output format (main code, subcode, description).
It is striking that the mapping from ChatGPT Codes onto the content analysis based on SBERT matches at the main code level. Except for one code mapping (Socio-cultural influences + Migration and religion), all other mappings comply with the same main code level: Examples matched with Forms of manifestations, Reasons for stereotyping with Conditions of emergence of stereotypes, both addressing RQ1, and Strategies in dealing with stereotypes with Pedagogical strategies in dealing with stereotypes, addressing RQ2. The structure of both code systems, oriented on the RQs, seems equivalent and therefore follows a similar analytic logic. Looking more closely at the mapping, it is notable that some mappings seem contradictory: Perspective shift occurs in both code systems, but both codes were mapped to other codes due to greater similarity based on their descriptions. Furthermore, high semantic similarity does not always align with a high Jaccard Index or similar EV-Centrality: While Geography-specific factors moderately co-occur with nine other codes, the semantically mapped ChatGPT code Didactic reduction and time pressure remains structurally peripheral in the ChatGPT network.
The content analysis yielded six additional subcategories compared with ChatGPT, indicating greater differentiation across all three main codes. The six unmapped codes are: Subject content and Countries as Examples for stereotyping; Teachers (causes) and Students (causes) as reasons for stereotyping; Causal analysis; and Student-oriented strategies for Strategies in dealing with stereotypes. That the code Subject content is missing from the ChatGPT code system is thematically striking, since it is the most frequent code and the one with the highest EV-Centrality in the content analysis. Besides this, it is one of the most subject-specific codes and directly addresses part of RQ1: Does stereotyping occur in geography class? And if so, how and why? Finally, the ChatGPT network exhibits a concentration of high EV-centrality values in three codes. In contrast, the content analysis network contains fewer high EV-centrality values overall, with many codes exhibiting similar, moderate EV-centrality levels. Furthermore, some codes in the content analysis are further differentiated into subcodes (Continents, Countries, Country groups), whereas ChatGPT summarizes spatial stereotypes under a single subcode (Spatial and developmental stereotypes).
Looking more closely at the explanatory logic of stereotyping, addressing the last part of RQ1, we can identify four mapped reasons for stereotyping. Three reasons of the content analysis remain unmapped: Teachers and students as reasons for stereotyping, and sociocultural influences. Taking all the reasons of the content analysis together, we identified many different actors, like teachers and students as attendees in class, social backgrounds as socialization factors, subject-specific factors like complexity reduction, textbooks as a central medium in class, and general educational policy factors like the time pressure explaining stereotyping in geography class. In contrast, the ChatGPT code system addresses only half of these factors, leaving out two main actors in the classroom. Students, as one of the main actors, remain very limited in addressing Knowledge deficits as a breeding ground, leaving no room for other student-related factors. In contrast, Students (c) have the third-highest EV-Centrality in the content analysis network, indicating structural importance in explaining stereotyping in geography class. By implication, it seems that ChatGPT focused more on structural conditions than on individual actors. In contrast, the content analysis reflects more nuanced reasons, likely leading to a more differentiated analysis and results regarding RQ1.
4. Discussion
4.1. Can We Map the ChatGPT Code System onto the Content Analysis System Based on Semantic Similarities?
The analysis indicates substantial semantic overlap between ChatGPT codes and Content analysis, with moderate-to-high cosine similarities (0.6–0.81), resulting in 13 mappings as determined by the Hungarian algorithm. Generally, the mapping of subcodes matches the corresponding main codes in 11 of 13 pairings, reflecting an orientation throughout the research aims: what stereotypes exist (examples), why they exist (reasons), and how teachers deal with them (strategies). Two pairings (Migration and religion (Manifestations) and Sociocultural influences (Reasons), and Eurocentrism (Conditions of emergence) and Perspective shift (strategies)) do not match equal main codes. This might be a disadvantage of the Hungarian algorithm, since the one-to-one mapping enforces the mapping, and codes might run out of partners. However, both similarities are moderate and above the 0.6 threshold, but the lowest and second-lowest of all mapped codes (0.61 and 0.6).
Another striking result in our data is the consistently high semantic similarity of strategy-related codes, which occupy the six highest ranks (1st–6th). This might be explained by the fact that strategies were less geography-specific than examples or reasons for stereotyping in our dataset, which is supported by looking more closely at the unmapped content analysis codes: First, they primarily represent domain-specific examples and reasons (differentiated spatial stereotypes, such as countries and country groups, Subject content, and geography-specific reasons). Secondly, they represent more actor-specific causes, and thirdly, more differentiated strategy codes.
These findings can be supported by previous research: An overarching pattern can be identified, while domain-specific focus appears more difficult (
Lixandru, 2024;
Myers et al., 2024). ChatGPT, as a large language model, can identify and emphasize dominant concepts and recurring patterns within large datasets based on probabilistic relationships (
Myers et al., 2024). Since the model does not possess an independent contextual understanding of the text, it cannot fully follow interpretative logics that require human reasoning (
Huang et al., 2026;
Leifheit et al., 2024;
Trott et al., 2023), which, in this case, is visible in the domain-specific perspective needed to address the RQs. However, reflecting all results from RQ1, we cannot support the overall conclusion of
Morgan (
2023) and
Turobov et al. (
2024) that LLMs in general struggle to generate interpretive themes, as shown by ChatGPT’s structured code system for a sensitive, interpretive topic like stereotyping.
4.2. Do Semantically Corresponding Codes Refer to the Same Content Represented by the Coding Units?
Using the Hungarian algorithm-mapped codes across the eleven transcripts, the Jaccard Index (J = 0.39) was used to quantify the overlap in coding unit assignments between ChatGPT and content analysis. However, the overlap varied highly across the 13 code pairs: five codes showed very low agreement (J < 0.14), whereas three codes showed low-moderate agreement (0.33–0.38) and five codes showed high agreement (J = 0.66–0.89). Comparing the Jaccard Index with semantic similarities further revealed that semantically similar codes did not necessarily refer to the same coding units and vice versa. For example, Teaching methods and Perspective shift and irritation exhibited high semantic similarity (cosine similarity ≈ 0.78), yet their overlap in coding-unit assignments was very low (J = 0.10). In contrast, Dialogue and reflection and Dialogue-based reflection and intervention showed both high semantic similarity (cosine similarity = 0.81) and substantial overlap in coding-unit assignments (J = 0.66).
This indicates that semantically similar code labels do not necessarily originate from the same textual evidence. Rather, the present findings suggest that the two coding systems differ in how qualitative information is structured during coding.
One pattern consistently observed in the present data is a difference in code granularity. Overall, the ChatGPT-generated coding system tended to aggregate conceptually related phenomena into broader code categories, whereas the content analysis differentiated these phenomena into several more specific codes. This can be illustrated by the matched pair Spatial and developmental stereotypes (ChatGPT) and Continents (content analysis). Although both codes overlap conceptually with respect to spatial stereotypes, the ChatGPT code additionally encompasses developmental stereotypes and therefore represents a broader conceptual category. In contrast, the content analysis distinguishes spatial stereotypes into more subcodes (Continents, Countries, and Country groups), thereby distributing coding unit assignments across multiple codes. This refers to almost all ChatGPT codes, which, when reflecting the broader code names, aggregate two parts into one code (e.g., Migration and Religion, Spatial and developmental stereotypes, Perspective shift and irritation, Working with data and contextualization, Social and origin stereotypes, Complexity reduction and time pressure, etc.).
Consequently, coding units assigned to the broader ChatGPT category are only partially represented by the matched content analysis code, resulting in partial overlap despite substantial semantic similarity. The observed differences in code granularity are consistent with previous findings suggesting that LLM-generated coding systems tend to comprise fewer and broader codes than coding systems developed through conventional thematic or qualitative content analysis (
Muasher-Kerwin et al., 2026). This, in turn, might affect the applicability of LLMs in inductive coding, but strongly depends on the research aims of each study: While the domain-specific outcome was of high research interest in our original study and therefore got high focus in the development of the code system, identifying overall patterns in larger datasets might be of interest in other studies and thus, LLM coding might be useful. However, critical reflection on the output is essential.
These findings also raise a more fundamental question regarding the emergence of LLM-generated coding systems. In qualitative content analysis following
Kuckartz and Rädiker (
2022), inductive code systems are assumed to emerge through an iterative process in which coding units are repeatedly examined, compared, differentiated, and progressively abstracted into higher-order conceptual codes. Whether a comparable process underlies the generation of coding systems by LLMs remains unknown, even though we requested a tabular reference of the merged codes during the coding system’s synthesis. Although the resulting ChatGPT coding system appears semantically coherent, the internal process through which individual coding decisions are synthesized into broader conceptual categories cannot be observed. Owing to the black-box nature of large language models (
Huang et al., 2026;
Leifheit et al., 2024;
Trott et al., 2023), it remains unclear whether broader codes emerge through a process analogous to iterative qualitative abstraction or through fundamentally different mechanisms of statistical pattern recognition. Consequently, the present approach cannot determine why semantically similar code systems sometimes exhibit only limited overlap in coding unit assignments. Differences in code granularity are likely to explain part of this pattern, but they do not fully account for the observed discrepancies. Future studies using larger and more diverse datasets should therefore place greater emphasis on reconstructing this process, for example, by systematically documenting intermediate coding steps and anchor examples. Such transparency would not only improve the traceability of LLM-generated code systems but also help identify potential inconsistencies or unsupported abstractions that arise during code generation, particularly given that LLMs may produce hallucinated outputs (
Farquhar et al., 2024;
Kalai et al., 2026).
A further explanation may lie in differences in trigger conditions. Codes showing high overlap in coding unit assignments may represent concepts with relatively explicit and consistently identifiable coding criteria. In contrast, lower overlap may indicate concepts that depend more strongly on contextual interpretation. This interpretation is consistent with previous findings suggesting that ChatGPT relies more strongly on pattern-based code assignment, whereas manual coding incorporates a more context-sensitive interpretation of textual material (
Yue et al., 2025). Together, these findings suggest that observed disagreements reflect not only differences in coding decisions but also differences in how qualitative information is conceptualized and structured during the development of a coding system in both coding procedures: LLM-based and human-derived iterative code systems on sensitive, interpretive material.
Finally, sensitivity to prompting needs to be reflected. Recent studies have shown that prompt design can substantially influence the structure, specificity, and consistency of LLM-generated qualitative analyses (
Han et al., 2026); for example, through personality and emotion-driven variations (
Ma et al., 2025). Furthermore, as demonstrated by
Zeng et al. (
2025), the reliability and reproducibility of LLM outputs, such as those from ChatGPT, can vary considerably even under identical conditions. The prompting strategies used in the present study were therefore developed iteratively, guided by established prompting principles and repeatedly refined until stable coding systems were obtained (
Alhindi et al., 2022;
Ben-David et al., 2022;
Goswami et al., 2024;
Kwon et al., 2025;
Liu, 2025;
Louatouate & Zeriouh, 2025). While this supports the methodological robustness of the proposed workflow, alternative prompting strategies may produce different coding structures and should therefore be systematically investigated in future research.
4.3. Are There Systematic Differences Between the Two Code Systems?
Although both networks exhibit a comparable network density and global clustering, the distribution of eigenvector centrality revealed structural differences between the two coding systems. The ChatGPT-based network showed a stronger concentration of EV-centrality in a small number of dominant codes. Descriptively, this pattern indicates that a limited set of overarching categories accounts for a larger share of the observed co-occurrences. In contrast, EV-centrality was distributed across a broader range of codes in the content analysis network, suggesting a more even distribution of structural importance across concepts (
Bloch et al., 2023;
Newman et al., 2006).
This pattern can be interpreted as a difference in analytic granularity. The higher concentration of central nodes in the GPT system suggests a tendency to aggregate conceptually related phenomena into broader codes, as suggested previously in relation to the code names. Such a pattern is consistent with previous research suggesting that AI-generated coding systems tend to prioritize generalizable patterns over fine-grained conceptual distinctions (
Morgan, 2023;
Peters & Chin-Yee, 2025). As argued by
Christou (
2024) and
Dahal (
2024) GenAI is particularly effective at identifying recurring structures in textual data but may rely on more abstract, less hierarchically differentiated representations than researcher-driven qualitative analysis.
These differences become especially visible in relation to domain-specific content. In the content analysis network, geography-specific factors are integrated into the broader structure and linked to multiple other codes, indicating that subject-specific contextualization is integral to the analytical framework. In contrast, comparable domain-specific concepts are largely absent from the central structure of the ChatGPT-based network. The most closely related concepts (Didactic reduction and time pressure) remain peripheral, supporting the results discussed in RQ1.
Methodologically, these findings demonstrate that similar global network properties do not necessarily imply similar conceptual structures. While network density and global clustering characterize the overall level of connectivity within a network, eigenvector centrality provides complementary information by identifying the concepts that occupy structurally influential positions. Consequently, differences between qualitative coding systems may become apparent primarily at the level of node centrality, even when global network measures remain largely comparable.
Furthermore, the consistently lower centrality of strategy-related codes in both networks suggests that the analytical focus of both systems lies more strongly on describing and explaining the phenomenon than on identifying responses or interventions. Overall, the findings indicate that systematic differences between the two coding systems lie in their conceptual organization, level of differentiation, and integration of domain-specific knowledge.
4.4. Limitations
Limitations of our study include the small dataset and its focus on a specific research topic. Given the interpretive nature of the interviews and the limited sample size, future research should validate both the procedure and the prompts using larger datasets and different research contexts to assess the generalizability of these identified patterns. Prompt sensitivity also warrants further investigation, as alternative prompt formulations may influence the generated coding systems despite iterative prompt refinement and repeated independent coding runs designed to account for the non-deterministic nature of LLM outputs. Although the Hungarian algorithm provides a reproducible one-to-one mapping between the two coding systems, alternative mapping strategies may yield different code correspondences and, consequently, slightly different comparison results. Furthermore, as qualitative content analysis is interpretive, multiple valid coding systems may emerge from the same dataset, particularly when analyzing complex and sensitive phenomena, such as stereotyping. Both coding systems were developed to address the same RQs, with this analytical focus explicitly incorporated into the ChatGPT prompt. Accordingly, the findings should be interpreted as a comparison between two analytical interpretations that identify both similarities and differences, rather than as a comparison against an objective standard.
5. Conclusions
This study aimed to develop an iterative, inductive coding system generated by ChatGPT, using detailed prompts grounded in recognized, proven strategies. This system was then compared, using both quantitative and qualitative methods, with the coding system developed through content analysis to identify similarities and differences and elaborate on striking results. These findings could, in turn, serve as a basis for further evaluation and optimization of the prompts used to develop a coding system with LLMs. Overall, the developed ChatGPT system shows a well-organized structure with many similar granularities of code, without redundancies, through this iterative process. Based on semantic comparison and qualitative evaluation, the ChatGPT code system is divided into three main codes with corresponding subcodes, directly addressing the three RQs. To further validate the semantic mapping, the Jaccard index was calculated at the level of the underlying coding units. The goal was to verify whether semantically mapped code pairs could indeed be traced to the same text segments as their corresponding codes in the content analysis. Our results show partial agreement: For five of the 13 semantically mapped code pairs, a match between the underlying coding units could be demonstrated. At the same time, differences were observed: conceptual differentiation, as evidenced by EV centrality in network analysis, and a lack of domain-specific knowledge to adequately address the RQs. Finally, critical reflection of the output is necessary, since LLMs may hallucinate. The following key learnings may be used to evaluate the prompting framework further while testing it in larger and different research contexts:
Five separate ChatGPT coding systems were generated and independently applied to the dataset. Across the repeated coding of each transcript, an average Jaccard Index of 0.83 indicated substantial overlap in coding-unit assignments. These findings are consistent with the proposition that iterative prompting and repeated validation contribute to methodological robustness and thematic convergence (
Gao et al., 2025;
Kwon et al., 2025). This aligns with the idea of iterative refinement approximating ensemble coding logic (
Jiang et al., 2021). Furthermore, these findings support conceptualizing ChatGPT as a probabilistic yet stabilizable analytical collaborator. Accordingly, methodological robustness should not be regarded as an inherent property of LLM-assisted coding but as a characteristic that can be strengthened through structured prompting, iterative refinement, and repeated validation.
In the content analysis network, geography-specific factors are integrated into the broader structure and linked to multiple other categories, indicating that subject-specific contextualization is integral to the analytical framework. In the ChatGPT-based system, comparable domain-specific elements are largely absent from the central structure. The network analysis further suggests that ChatGPT-generated coding systems tend to prioritize generalizable patterns over fine-grained conceptual distinctions. In contrast, the content analysis network exhibits a broader and more distributed core structure, with several codes showing comparatively high EV-Centrality values. These findings suggest that LLM-assisted coding can efficiently identify overarching thematic structures, while contextualization and conceptual differentiation remain central strengths of content analysis. This might be improved by refining and testing the prompting framework on larger datasets.