Skip to Content
InformationInformation
  • Article
  • Open Access

12 June 2026

26 Pages

Temporal Robustness of Large Language Models for Thematic Classification of UN General Assembly Debates

,
,
,
and
1
Speech and Language Processing Group, Department of Computer Science, (FJWU-IIUI), Islamabad 44000, Pakistan
2
Department of Computer Science, International Islamic University, Islamabad 44000, Pakistan
3
Department of Communication and Media Studies, Fatima Jinnah Women University, Rawalpindi 46000, Pakistan
4
Interdisciplinary Sustainable Systems (IS2) Group, College of Computer and Information Sciences, Prince Sultan University, Riyadh 12435, Saudi Arabia
This article belongs to the Section Artificial Intelligence

Abstract

Thematic analysis of large-scale political discourse remains a challenge due to semantic complexity and overlapping policy areas and changing diplomatic vocabulary. Although large language models (LLMs) offer promise for scalable thematic classification, their reliability in politically sensitive contexts requires systematic validation against expert human annotations. We evaluate LLM-based thematic classification of United Nations General Assembly (UNGA) speeches across a decade (2014–2023), using 7680 human-annotated themes mapped into 12 policy domains. Our results show that DeepSeek R1 achieves the highest accuracy 77% (F1 = 0.73), followed by ChatGPT, Gemini and LLaMA, with strong performance in lexically stable domains but substantial degradation in semantically overlapping categories such as governance and international cooperation. A unique dimension of our work is timeline analysis, which shows that the performance of LLMs over the years varies strongly and the precision decreases during times of rhetorical transformation, including pandemic-related discussions and the discourses of cooperation determined by the Russia–Ukraine conflict. By linking domain-level ambiguity and geopolitical shifts to temporal instability, this study introduces a dynamic robustness perspective for evaluating LLMs in computational political discourse analysis.

1. Introduction

The digitization of global communication has improved diplomatic conversation; however, explaining intricate geopolitical narratives still remains challenging [1,2,3]. This challenge intensifies when analyzing discourse across temporal dimensions, as diplomatic language evolves in response to shifting geopolitical events and global crises. Traditional discourse analysis approaches are not scalable for handling increasing volumes of political communication [4], while computational approaches like Latent Dirichlet Allocation (LDA), Support Vector Machines (SVMs) and Structured Topic Modeling (STM) [5] often yield only surface-level insights. Although methods like LDA can handle a large amount of data, they do not capture enough semantic depth in political discussion. The advent of large language models (LLMs) has provided computational efficiency with a deeper semantic understanding [6,7,8], yet their effectiveness for diplomatic discourse analysis, especially temporal insight across evolving geopolitical contexts, remains underexplored.
The United Nation General Assembly (UNGA) [9] is a powerful platform for global conversation, as the leaders of the world articulate their national priorities and diplomatic positions [10,11]. The UNGA debates are a rich and structured resource of official annual speeches of heads of states, publicly available in standardized digital text format which makes them suitable for computational analysis [12]. The transcripts are important as they define the position of various countries in the international environment, how they interact in the international arena and the changes in global policies and initiatives [13,14,15]. This corpus has enabled systematic analysis of how nations frame their priorities over time.
Despite the wide application of the UNGA corpus in political and policy alignment research, existing studies [3,5,16] have mainly relied on traditional methods, including LDA [5], structural topic modeling, word embeddings [17], and semi-supervised classification [16] and high-level clusters such as BERTopic and MiniLM [12] to extract topics from these speeches. However, these approaches tend to lack domain-level thematic clarity, which can be better captured by contextual models such as large language models. Currently, there is a lack of any benchmark studies to measure the domain classification capabilities of LLMs in political discourse, specifically in the multilateral debates at the UNGA. Furthermore, no previous study has performed a timeline analysis of LLM performance over the years in this context. We also present the first temporal robustness analysis (2014–2023) of LLM performance, which demonstrates that accuracy fluctuates over time in response to changing geopolitical events like COVID-19, climate crises, and international wars.
UNGA speeches are complex due to the use of rhetorical styles and culturally encoded diplomatic language, as well as implicit references to historical and geopolitical contexts [12,18,19]. These speeches are also lengthy, making it difficult to draw consistent conclusions only based on direct textual analysis, while traditional political science methods are based on the manual analysis of political discourse. To overcome this limitation, this study uses automated Natural Language Processing (NLP) to extract and classify themes in UNGA speeches, offering an alternative to the traditional manual method of political discourse analysis. The extracted themes were then mapped to a list of twelve predetermined policy domains. This approach makes it possible to achieve a clearer understanding of the underlying policy priorities and allows a systematic comparison of national agendas to be made over both temporal and regional scales.
The advancement of the large language models (LLMs) has opened up new opportunities for automating thematic analysis, as deep semantic understanding and the ability to perform zero-shot classification have become possibilities [20,21,22,23]. LLMs offer a scalable, data-driven approach for systematic analysis on large textual corpora [20,21,24,25].
This work demonstrates the promise and limits of employing LLMs in political text analysis at scale [26]. Beyond descriptive benchmarking, this study advances three scientific insights: first, it establishes that LLM classification performance is temporal, varying systematically with geopolitical discourse shifts. Second, it demonstrates that model errors in diplomatic text are structurally driven by lexical overlap between semantically proximate domains and third, it provides empirical evidence that architectural differences, including reinforcement learning-based alignment, are associated with measurable advantages in temporal robustness, with implications for model selection in politically sensitive NLP applications.
The following research questions are formulated to systematically evaluate the effectiveness of different LLMs in thematic domain classification of political discourse:
  • Q1: Can LLMs efficiently replicate expert-level thematic classification in political discourse?
  • Q2: In cases where LLM and human thematic classifications diverge what does this tell us about LLM interpretability in geopolitical discourse?
  • Q3: What are the temporal and domain-wise strengths and weaknesses of the evaluated LLMs throughout the decade (2014–2023)?
A major contribution is sharing the thematic data https://huggingface.co/datasets/SLPG/UNGA (accessed on 30 April 2026) with the research community to enable further research on the topic. The rest of this paper is organized as follows: the literature review is presented in Section 2. Section 3 describes the proposed methodology including data collection and the theme extraction pipeline. Section 4 presents the selected large language models. The results and analysis are described in Section 5. The paper concludes with a discussion in Section 6 and Section 7 stating key findings and future directions.

2. Literature Review

2.1. Traditional Methods

Thematic classification in political texts is traditionally carried out through models like Support Vector Machines (SVMs) and LDA [5,27] by relying on word frequencies [28]. Advanced techniques like Structural Topic Modeling (STM) and word embeddings have been identified to improve thematic coherence and interpretability [25,29,30]. However, these approaches remain fundamentally constrained by bag-of-words assumptions and surface-level co-occurrence patterns, rendering them incapable of distinguishing thematically adjacent domains such as international cooperation and multilateral cooperation or processing figurative language prevalent in political texts such as irony and sarcasm [31,32,33]. The proposed LLM-based framework directly addresses these limitations by leveraging deep contextual representations that capture domain-specific semantic distinctions beyond keyword presence, enabling more reliable classification in diplomatically nuanced discourse.

2.2. LLMs for Political Text Classification

Large language models (LLMs) have significantly advanced topic modeling and text classification by leveraging transformer-based architectures that encode contextual dependencies across extended textual sequences [7,34,35,36]. In contrast to probabilistic approaches such as LDAs, which primarily rely on word co-occurrence statistics, LLMs capture semantic nuance, pragmatic framing, and discourse-level structure, capabilities that are particularly relevant for political speech analysis [33,37]. However, performance in politically sensitive domains remains influenced by training data composition, alignment procedures, and embedded normative assumptions [38,39].
Human-annotated benchmarks therefore remain essential for assessing interpretive validity and classification reliability in such contexts [7,40,41]. Although few-shot prompting has demonstrated improved results in domain-specific settings, such configurations may obscure intrinsic generalization capacity [42]. The present study consequently adopts a zero-shot design to evaluate baseline domain classification performance without task-specific adaptation. This design choice also reveals a structural limitation of current models, namely difficulty in distinguishing semantically adjacent diplomatic categories, as reflected in lower F1-scores for international cooperation and multilateral cooperation.
Despite ongoing methodological progress, existing LLM-based political text research remains largely task-specific, frequently concentrating on summarization, stance detection, or electoral forecasting rather than systematic domain benchmarking [39]. Rigorous evaluation of zero-shot classification against expert-validated gold standards in institutionally structured corpora such as UNGA debates remains limited.

2.3. Comparison Between LLMs Based on Performance

Recent studies compare the performance of LLMs on the task of political discourse classification [41,43]. Although general-purpose models have shown great zero-shot performance [42,44], they can be further fine-tuned with domain-specific data to increase performance [43,45].
Relative research shows that DeepSeek-67B outperforms LLaMA-2 70B on several classification criteria [46]. The training pipeline of DeepSeek-R1, which incorporates supervised fine-tuning, reinforcement learning, and rejection sampling, makes it competitive with OpenAI on reasoning tasks, including MMLU and GPQA [47,48].

2.4. Comparison with Prior Studies

Prior computational [3,17] analyses of UNGA debates have primarily focused on uncovering latent thematic structures (e.g., STM) or measuring semantic proximity (e.g., word embeddings). While these approaches reveal macro-level agenda patterns, they do not operationalize explicit domain taxonomies nor assess classification reliability against expert judgment. As a result, the field lacks a validated benchmark for domain-level interpretive alignment in diplomatic discourse.
Recent work by [12] employs BERTopic with MiniLM to extract thematic clusters from UNGA debates. While effective for identifying latent topic structures, such clustering-based approaches do not enforce explicit domain taxonomies and therefore offer limited interpretive precision in policy-aligned classification tasks. Similarly, the human-in-the-loop framework proposed by [20] demonstrates the utility of LLM-assisted thematic analysis but relies on iterative manual refinement, constraining scalability and reproducibility.
In contrast, the proposed framework adopts a fully automated, multi-model bench-marking design that evaluates zero-shot domain classification across twelve expert-defined policy categories. Beyond topic extraction, our framework systematically measures interpretive alignment through confusion matrices, domain-wise F1-scores, and temporal robustness analysis. While the literature reflects a broader methodological transition from frequency-based models to contextual LLM architectures, systematic benchmarking against human-validated gold standards in politically sensitive and temporally dynamic corpora remain underdeveloped. By integrating expert annotation with decade-long cross-model evaluation, this study addresses that gap and establishes a reproducible benchmark for computational political discourse analysis.

3. Proposed Scheme

The architecture of the proposed scheme is shown in Figure 1. The workflow describes the processing of raw UNGA speeches with LLM-based theme extraction, fine-tuning on structured training data, and human evaluation of thematic coherence and accuracy. The entire methodological procedure is constructed upon this framework.
Figure 1. Systematic workflow for UNGA speech processing, covering two-stage theme extraction (ChatGPT-4 zero-shot for 2023; Gemini-1.5 Flash fine-tuned for 2014–2022), human expert annotation, and zero-shot domain mapping using ChatGPT-4, DeepSeek R1, Gemini and LLaMA.
The data is obtained from the official United Nations General Assembly (UNGA) archives with more than 20,000 speeches from 1946 to 2023. The UNGA speech corpus was chosen due to its global representativeness, formatted structure and policy-filled nature, providing a perfect study ground to understand how states describe national priorities in a multilateral environment [49,50]. The speeches contain high-level, consistent political discourse that is semantically rich and publicly accessible [15].
Therefore, this corpus is ideal for determining the interpretability of language models in complex geopolitical discourse. The debates dataset was downloaded from the official UN archives https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/0TJX8Y (accessed on 31 July 2025). These speeches are in text format with metadata in Excel format, which includes the year they were delivered, the country of the speaker, the speaker’s name, and their official designation. The speeches align with meta-data through country ISO codes. The scope of the study was recent trends over the decade, and consequently speeches from 2014 to 2023 were selected as this era exhibited rapid development of global discourse on critical matters such as economic policies, geopolitical aspects and climate change.
We also evaluated temporal robustness by investigating year-by-year changes in performance to determine how LLMs respond to changing political language with time.

3.1. Themes Extraction Pipeline

An automated LLM-based pipeline was developed to process and categorize themes in speeches [47]. In the first step, the speech was given to OpenAI’s zero-shot chat model (ChatGPT) [51] using a structured prompt template, shown in Figure 2.
Figure 2. Prompt for extracting themes from speech content.
The model was instructed to extract up to eight concise, policy-relevant themes from each speech, following the same process for all speeches delivered in 2023. This process generated a theme list comprising 1250 themes which used a wide range of linguistic registers for similar domains, these were then mapped to twelve main domains as explained in Section 3.2.
These themes served as supervised training data for fine-tuning the Google-based large language model Gemini [52]. Model selection for theme extraction was informed by the long-form nature of UNGA speeches. Gemini-1.5 Flash-001-Tuning was used in the corpus creation phase because it could combine retrieval-augmented reasoning and multimodal input handling, which enabled the extraction of themes in contextually rich political discourse in long-form speeches. Furthermore, it was specifically chosen for its extended context window supporting millions of tokens, ensuring full speech processing without truncation.
In a comparative assessment of model political bias [38], Gemini demonstrated relatively lower measured bias compared to other evaluated models. While such findings are dataset dependent, they suggest that the model may offer comparatively stable performance in politically sensitive thematic extraction tasks.
The Gemini model was fine-tuned on the 2023 theme dataset to extract the main themes for 2022. This process was iteratively followed for the remaining years, i.e., fine-tuning on 2022 data to extract the main themes for 2021. This reverse year fine-tuning was repeated iteratively for the past 10 years, where each year’s theme was extracted from the next year’s data. This reverse-chronological fine-tuning methodology is designed to support the extraction of contextually informed themes, based on the assumption of temporal thematic continuity across consecutive UNGA sessions. This approach enables progressive temporal learning as a design intent and offers a potentially useful approach to analyzing political discourse, though its advantage over forward fine-tuning or zero-shot extraction baselines remains to be empirically validated through ablation comparison. At the end of this phase, we had a dataset of 1980 speeches with 7680 themes as shown in Figure 1.
This reverse-chronological fine-tuning approach is grounded in the continuity of diplomatic discourse, wherein thematic vocabulary and policy priorities exhibit strong longitudinal coherence across consecutive annual sessions. Fine-tuning on proximate future data therefore provides a contextually informed prior rather than introducing anachronistic knowledge.
Table 1. Human rating scale for evaluating extracted themes.
Crucially, this approach does not constitute data leakage in the conventional sense, as the fine-tuning phase concerns only theme extraction, while domain classification is performed independently through zero-shot prompting without any exposure to gold-standard labels. The two phases are methodologically independent, ensuring that classification evaluation is not influenced by fine-tuning data.
A manual review was conducted to validate the theme extraction process, focusing on assessing the thematic coverage of the speeches. The LLM-extracted themes were rated by expert annotators with a background in political science and computational linguistics, using a scale of 1 to 5 as shown in Table 1, covering approximately 20% of the total dataset, consistent with standard practices in manual validation of large annotated corpora. The mean rating of 4.2 confirms strong alignment between LLM-extracted themes and human judgment.

3.2. Domain Mapping Guidelines

The extracted key themes had a diverse vocabulary and terminology refers to similar issues in each speech. Although all speeches were in English, cultural and diplomatic registers showed a high degree of lexical variation, as shown in Figure 3, which presents the clustering of various lexical expressions across countries around their central policy areas.
The initial clustering analysis revealed 16 domains selected to capture the most relevant and policy-relevant themes frequently discussed in UNGA debates, such as human rights, climate change, peace, poverty and governance. This taxonomy is theoretically grounded in three established policy frameworks such as the United Nations Sustainable Development Goals [53], the UN Secretary General’s Common Agenda report [54], and thematic categories identified in prior UNGA corpus studies [13]. The domain consolidation process was therefore not arbitrary but guided by both theoretical precedent and empirical overlap analysis.
However, we found semantic overlap of some of the categories during the domain mapping process. To deal with this Figure 4 indicates that the consolidation process merged overlapping domains, resulting in a clean and conceptually distinct set of 12 domains. The selected domains are conceptually clear, reflect diplomatic priorities in various contexts, and offer a structure framework that is helpful in both human annotation and LLM classification.
Specifically, regional stability and peace and stability were consolidated into conflict resolution and peace due to their shared security-oriented vocabulary and frequent co-occurrence of terms such as ceasefire, peacekeeping, and regional security. Democratic governance and political sovereignty were merged into political sovereignty and governance given their overlapping emphasis on domestic institutional structures and reform narratives. Poverty was absorbed into economic development as themes in this category consistently co-occurred with development financing and poverty reduction frameworks.
Figure 3. Thematic discourse knowledge graph showing extracted themes (green), domain clusters (yellow) and countries (blue) with edges indicating yearly topic–domain mappings.
These consolidation decisions were guided by both theoretical precedent from established UN policy frameworks and empirical co-occurrence analysis. We note that the domains where consolidation involved the most conceptually proximate categories, namely international cooperation, multilateral cooperation, and global governance and cooperation, are precisely those where inter-annotator agreement was lowest and model performance was weakest, suggesting that residual boundary ambiguity from the consolidation process may have contributed to classification difficulty in these areas.

3.3. Guidelines

Guidelines for categorizing the themes extracted from UNGA speeches into key domains. The creation of domain mapping guidelines was an important task that was carried out to ensure clarity and consistency in categorization. These guidelines define the conceptual framework.
Criteria for each domain, providing a standardized reference for manual mapping and reducing ambiguity during classification. The guidelines are as follows.
  • Human Rights and Social Justice: Themes related to rights, equality, justice, resolving the issues of marginalized groups, and gender equality linked with the issues of social inclusion and individual liberty.
  • Economic Development: Themes related to economic growth, trade, financial systems, poverty reduction and development goals that have focused on economic progress.
  • Humanitarian and International Affairs: Themes related to crisis response, emergency aid, global health and disaster management were concerned with immediate assistance.
  • Political Sovereignty and Governance: Themes related to domestic governance systems, democratic systems, national reform and sovereignty.
  • Conflict Resolution and Peace: Themes related to peace processes, security issues, ceasefire and stability that highlight prevention and resolution of conflicts.
  • Multilateral Cooperation: Themes related to institutional frameworks, international organizations and multilateral agreements involving structured cooperation.
  • Terrorism and Counter-terrorism: Themes related to violent extremism, security threats, and counter-terrorism measures that aim at combating non-state-armed groups.
  • Sustainable Development: Themes related to long-term balanced development, sustainability objectives, and intergenerational equity, with a great emphasis on environmental and social integration.
Figure 4. Merging domains based on semantic overlap.
9.
Climate Change: Themes related to climate policy, environmental crisis, adaptation and mitigation with an emphasis on ecological issues, such as climate change policies and climate change adaptation.
10.
Global Governance and Cooperation: Themes that emphasize global policy frameworks, accountability, system reform and international coordination that favor the global regulatory frameworks.
11.
UN Reform and Role: Themes about the structure, functions and activities of the United Nations (UN), UN agencies and UN institutional processes that focus on the effectiveness of UN operations.
12.
International Cooperation: Themes on international relations like diplomacy, partnerships, solidarity and collaboration that express general international interaction like international cooperation agreements and international cooperation in trade.
The human annotator verified the thematic categorization and mapped it into 12 predefined domains. During human annotation, several challenges emerged. Some of the extracted topics were semantically ambiguous or related to multiple thematic domains, making it difficult to assign them to a single predefined domain. If a theme fit into several domains, the annotator picked the domain that was closely related using the guidelines provided.

3.4. Inter-Annotator Agreement

We performed an inter-annotator agreement study, in which a second human annotator independently mapped 1152 themes, representing approximately 15% of the total dataset, consistent with standard practices in manual validation of large annotated corpora. Cohen’s Kappa analysis as shown in Table 2 demonstrated substantial to almost perfect agreement in most domains: terrorism (K = 1.00), climate change (K = 0.97), UN Reform (K = 0.86), and conflict resolution (K = 0.80). However, lower agreement was observed in semantically proximate categories: international cooperation (K = 0.58), multilateral cooperation (K = 0.34), and global governance (K = 0.29), reflecting the inherent ambiguity of diplomatic language. Overall, inter-annotator agreement across all 12 domains was substantial (average K = 0.712). Treating these domains as analytically distinct units is a deliberate methodological choice that exposes the interpretive limits of both human and machine classification, rather than assuming artificial categorical clarity.
Table 2. Inter-annotator agreement across thematic domains. Cohen’s κ values indicate substantial to almost perfect reliability for most domains.
Given the notably lower agreement in three domains, specifically multilateral cooperation (κ = 0.35), global governance and cooperation (κ = 0.30), and international cooperation (κ = 0.58), findings reported for these categories throughout Section 5 should be treated as indicative rather than definitive. Classification performance in these domains reflects not only model limitations but also the inherent uncertainty of the gold-standard labels themselves.

3.5. LLM-Based Domain Mapping

After establishing a gold-standard dataset through human annotation, LLM-generated classifications were evaluated against this. To ensure consistent comparison, we also experimented with other LLMs such as ChatGPT-4, Gemini, LLaMA instruct and Deepseek R1 to classify the same extracted themes into these 12 domains. All models were prompted in a zero-shot setting as shown in Figure 5, which consisted of the list of domains and the instruction to select only one domain that best corresponded to each theme. The classification process, as shown in Algorithm 1, generated domain labels for every theme. The evaluation enabled us to compare the reliability of LLMs with expert human judgment in domain classification tasks.
The dual mapping framework facilitated a rigorous comparison of human and LLM-generated domain mapping on similar topics. Furthermore, it also enabled a performance comparison among various LLMs.

4. Selected Large Language Models

The four state-of-the-art large language models were evaluated for domain mapping. LLMs were selected on the criterion of reproducibility and open accessibility. DeepSeek R1 [55] and ChatGPT (GPT-4) [56] were accessed through their respective publicly available APIs. Gemini-1.5-Flash-001-Tuning [52] was accessible through the Google platform, whereas Meta’s Meta LLaMA 3.2 1B Instruct [57] was selected as a lightweight model deployable through the HuggingFace platform https://huggingface.co/meta-llama/Llama-3.2-1B-Instruct (accessed on 31 July 2025). It was intentionally included as a computationally efficient baseline rather than a direct full-scale LLM comparator, reflecting deployment scenarios where resource-constrained inference is necessary. Consequently, its performance should be interpreted as a lower-bound reference point rather than a representation of the overall capabilities of large language models for this task.
Figure 5. LLM prompt template for thematic domain classification of extracted UNGA key themes.
Training data cut-offs for all models occur between mid-2024 and 2023 but are able to fetch real-time information from online sources. Thus, all LLMs can provide up-to-date results for test classification. The model selection was based on architectural variety—ChatGPT and Gemini are high-performance commercial models, DeepSeek was included for its reasoning capabilities, and LLaMA 3.2-1B instruct is a smaller open-source model.
ChatGPT GPT-4 is a large-scale decoder-based architecture pre-trained with multilingual corpora and fine-tuned with RLHF over broad general domain knowledge including political, environmental and government topics [58].
Gemini-1.5-Flash-001 is a multimodal Mixture-of-Experts (MoE) architecture with extended context capabilities in the millions of token regime, optimized for long context reasoning and policy domain, which are well suited to applications that need high-volume textual inputs.
DeepSeek R1 uses an MoE architecture with 671B total parameters, of which only 37B are activated per token, trained on 14.8T tokens with cold-start supervised fine-tuning followed by large-scale reinforcement learning using Group Relative Policy Optimization (GRPO), optimized for complex reasoning tasks [59].
LLaMA 3.2 1B Instruct model is a compact model trained on trillions of tokens with tuned instruction (supporting 1M+ token contexts) and optimized for device applications [60].
These models were employed to extract thematic insights from UNGA debates. We compared LLM-generated thematic mappings with expert human judgments and evaluated the interpretive depth and consistency of machine-generated thematic mappings. Thematic similarities and divergences are identified in how LLMs and humans view geopolitics discourse.

5. Empirical Evidence of LLMs Capability to Replicate Thematic Classification

The zero-shot performances of DeepSeek R1, ChatGPT, Gemini and LLaMA instruct model were compared to map thematic key themes extracted from UNGA speeches into 12 thematic domains.
Algorithm 1: Zero-Shot LLM Thematic Domain Mapping
Information 17 00589 i001
The accuracy [61], precision [62], recall [63], and F1-score [64] of each model were computed in relation to the gold standard of human annotated domain labels. Confusion matrices were used to depict the level of agreement and disagreement with human labeling at the domain level as shown in Figure 6.
Table 3 summarizes the zero-shot classification performance of all evaluated models, including 95% confidence intervals computed using the Wilson binomial method and bootstrap confidence intervals based on 1000 resampling iterations. DeepSeek R1 achieved the highest accuracy (77%, F1 = 0.73; 95% CI = [0.761, 0.779]), followed by ChatGPT (73%, F1 = 0.69; 95% CI = [0.719, 0.741]), Gemini (70%, F1 = 0.65; 95% CI = [0.688, 0.712]), and LLaMA (36%, F1 = 0.34; 95% CI = [0.345, 0.375]). The confidence intervals show minimal overlap across models, indicating clear separation in performance. This observation is further confirmed through McNemar’s test (Table 4), where all pairwise comparisons yield statistically significant differences (p < 0.001), demonstrating that the observed performance hierarchy is robust and not attributable to random variation. Bootstrap standard deviations remain low (0.005–0.008), further indicating the stability of the reported estimates. The superior performance of DeepSeek R1 may plausibly reflect its reasoning-oriented training paradigm, though we note that DeepSeek also differs from other evaluated models in architecture, training data composition, and parameter scale. The present experimental design does not isolate reinforcement learning as the sole causal factor, and this interpretation should therefore be treated as one plausible explanation rather than a confirmed finding. This advantage is evidenced by reduced off-diagonal misclassification in the confusion matrix 6, indicating more consistent separation between semantically adjacent categories. The statistical significance of this advantage is further reflected in pairwise comparisons, where DeepSeek exhibits the largest performance gap relative to LLaMA (χ2 = 88,516.44, p < 0.001).
Figure 6. Confusion matrices for LLMs vs. human mapping. Diagonal strength indicates alignment between LLMs and human mapping.
In addition, DeepSeek demonstrates stronger adherence to the closed label classification constraint, consistently assigning a single, well-defined domain per theme. This results in fewer boundary violations and contributes to its overall stability across predictions. Its performance is particularly strong in structurally well-defined domains such as climate change (F1 = 0.94) and terrorism and counter-terrorism (F1 = 0.96), where domain boundaries are clearer and less affected by interpretive ambiguity.
By comparison, ChatGPT exhibits relatively balanced performance but shows limitations in fine-grained domain discrimination, likely due to its general-purpose training, which prioritizes broad coverage over domain-specific precision. The statistically significant difference between ChatGPT and Gemini (χ2 = 974.90, p < 0.001) further indicates that models with similar overall accuracy may still differ meaningfully in domain-level behavior. Gemini performs well in normatively stable domains such as human rights and sustainable development but demonstrates increased confusion in abstract institutional categories, suggesting sensitivity to variation in diplomatic phrasing. LLaMA, included here as a lightweight baseline, shows limited capacity for capturing discourse-level relationships as expected given its 1B parameter scale, resulting in substantially higher misclassification rates across domains. These results reflect the constraints of resource-limited deployment rather than the general capability ceiling of the LLaMA model family.
Table 3. Zero-shot performance comparison of LLMs on thematic domain classification with statistical validation.
Importantly, performance degradation observed across all models in semantically overlapping categories is not solely attributable to model limitations but reflects the inherent ambiguity of diplomatic language. This is further supported by lower inter-annotator agreement in these domains, with Cohen’s κ values of 0.58, 0.34, and 0.29 for international cooperation, multilateral cooperation, and global governance, respectively. These findings indicate that classification difficulty is structurally embedded in the task itself. Overall, the results suggest that model performance in political discourse classification is jointly shaped by architectural design and the nature of the domain, where the ability to resolve contextual ambiguity plays a more critical role than overall accuracy alone.
Domains with low inter-annotator reliability (κ < 0.60), specifically international cooperation (κ = 0.58), multilateral cooperation (κ = 0.35), and global governance and cooperation (κ = 0.30), are shaded in Table 5. Performance figures in these domains should be interpreted with caution, as classification difficulty may reflect inherent label ambiguity in the gold standard rather than model limitation alone.
The Table 6 illustrates how LLMs classified the similar themes into different domains. The examples highlight that models like ChatGPT and DeepSeek R1 show strong alignment across most categories, while Gemini is mixed, as it correctly classifies in some categories, like climate change, but not in overlapping categories like international cooperation. However, LLaMA instruct shows misclassifications of thematic content.
A qualitative analysis of misclassification patterns reveals consistent lexical triggers that confuse the models. Themes containing terms such as “coordination,” “solidarity,” and “partnership” were frequently misclassified between international cooperation and multilateral cooperation, as these words appear interchangeably across both domains in diplomatic discourse. For instance, the theme “global moderation movement,” annotated as international cooperation, was mapped to global governance and cooperation by DeepSeek and multilateral cooperation by ChatGPT, suggesting that models rely on surface-level institutional keywords rather than contextual intent. Similarly, themes involving “peace operations” and “coalition building” were conflated between multilateral cooperation and conflict resolution, as both domains share security-related vocabulary. These patterns indicate that model errors are not random but systematically driven by lexical overlap in diplomatically abstract registers, where meaning is determined by pragmatic context rather than keyword presence.
Figure 6 presents the confusion matrices for all evaluated models, comparing LLM-predicted domain labels against human-annotated gold standards. Misclassification errors across all models exhibit a consistent structural pattern, concentrated in domains sharing overlapping lexical and thematic characteristics rather than being randomly distributed.
Table 4. Pairwise McNemar’s test results for statistical comparison of LLMs.
The most pervasive cross-model confusion occurs between international cooperation (IC) and multilateral cooperation (MC). ChatGPT produced 194 bidirectional confusions in this pair (12.2% of its errors), Gemini 166 (9.8%), DeepSeek 173 (27.0% despite its lowest overall error rate), and LLaMA 92 (3.5%). This consistency across architecturally distinct models indicates that the IC–MC boundary reflects inherent diplomatic language ambiguity rather than models specific limitations, corroborated by low inter-annotator agreement for these domains.
A second systematic pattern involves climate change (CC) and sustainable development (SD). Gemini misclassified 237 CC themes as SD (14.0% of its errors), while LLaMA produced 615 such misclassifications—its single largest error source (23.4% of all errors). DeepSeek showed markedly lower confusion in this pair, reflecting stronger contextual discrimination. A third recurring cluster involves global governance and cooperation (GGC), frequently confused with both IC and MC, accounting for 204 errors in ChatGPT (12.8%) and 138 in Gemini (8.1%).
Quantifying the aggregate contribution of semantic overlap reveals that 31.8% of ChatGPT errors (507 of 1593), 25.5% of Gemini errors (433 of 1697), and 35.0% of DeepSeek errors (224 of 640) originate from semantically proximate domain pairs. LLaMA presents a qualitatively distinct profile, with only 8.7% of its errors attributable to domain proximity, as its misclassifications are dominated by lexically triggered confusions such as CC → SD (615 errors) and conflict resolution misclassified as humanitarian affairs (278 errors), consistent with its limited contextual reasoning capacity. These findings collectively establish that classification difficulty in diplomatic discourse is structurally driven by domain boundary ambiguity, with architectural differences determining the degree to which this ambiguity is resolved.
Table 5 shows the domain-wise performance in F1-scores. All models performed well in specific domains such as climate change (F1 > 0.90), but their performance in areas that overlap, like international cooperation, human rights and social justice, were poor (F1 < 0.40). Overall, DeepSeek continued to perform well across most domains, with its best performance in categories such as ‘climate change’ (F1 = 0.94), ’terrorism and counter-terrorism’ (F1 = 0.96), and ’political sovereignty and governance’ (F1 = 0.79). However, ChatGPT performed well in general categories such as ’economic development’ and ’human rights and social justice’, which however saw a reduction in F1-score in overlapping or abstract categories like ’multilateral cooperation’ (F1-score: 0.49) and ’international cooperation’ (F1-score: 0.43). Similarly, Gemini, even though it excelled in F1-score for “climate change” and “economic development”, it performed very poorly for “international cooperation” (F1 = 0.24), which indicated under-classification and conservative decision boundaries. On the other hand, LLaMA instruct is able to achieve reasonable accuracy in domains like terrorism (79%), but it underperforms in highly important areas such as climate change (15%) and international cooperation (2%).
Table 5. Domain-wise F1-scores for all evaluated models. Shaded rows indicate low inter-annotator reliability (κ < 0.60); results in these domains should be interpreted with caution.
Our results show that LLMs can partially replicate domain mappings created by experts. For instance, DeepSeek R1 achieved a 73% F1-score which means it agreed with expert labels on the majority of unambiguous themes. According to our analysis, LLMs and a human annotator frequently diverged on topics that conceptually overlap. Since ‘social justice’ and ‘international cooperation’ use similar terminologies, Gemini and ChatGPT placed them in the wrong categories.

5.1. Evaluating Temporal Robustness of LLMs in Thematic Classification of UNGA Debates

To evaluate the temporal robustness of LLMs, we analyzed their timeline performance in thematic classification of UNGA debates. Table 7 shows that the temporal variation in the accuracy of LLMs between 2014 and 2023 was not constant but was influenced by the model architecture and the discourse that was discussed in the UNGA debates. In the early years (2014–2016), speeches centered on sovereignty, peace, and economic development [65,66,67] were mostly aligned with the generalist pre-training of ChatGPT, with an average accuracy of around 70–75%, whereas DeepSeek achieved a consistent improvement as it adapted to policy-driven narratives of its reasoning architecture [68].
DeepSeek also performed well (83.03% in 2018) in 2017–2018, when climate change and multilateral cooperation were in the limelight, indicating its ability to deal with reasoning-intensive and policy-specific discourse, but ChatGPT and Gemini reported moderate performance, and LLaMA was limited by capacity. The performance of LLMs dropped between 2019 and 2021, especially Gemini (68.23% in 2019, 69.43% in 2020), as speeches became multidimensional with COVID-19, humanitarian crises, and geopolitical conflicts.
This decline is supported by measurable shifts in the lexical composition of UNGA speeches, as shown in Table 8. Terms such as “pandemic” and “COVID” were virtually absent in 2019 but surged dramatically from 2020 to 1292 with 1068 occurrences, while “solidarity” and “crisis” nearly doubled, introducing significant lexical overlap across traditionally distinct domains that generalist models could not disambiguate. DeepSeek, conversely, remained relatively stable (78.79%), which may plausibly reflect its reinforcement learning-based training, among other architectural differences, whereas LLaMA dropped below 30% in 2021. In 2022–2023, the climate crisis and the Russia–Ukraine conflict dominated debates, where DeepSeek showed its strongest performance (84.81% in 2022 and 83.79% in 2023), ChatGPT and Gemini remained mid-level (around 72–73%), and LLaMA reached its lowest point (27.78%).
Table 6. Examples of LLM-based domain mapping, misclassifications are highlighted with unique background colors and domain labels are represented in bold (gold).
These findings indicate that LLM performance in domain mapping is time-dependent, shaped by shifts in global political discourse, and heavily reliant on architectural design.

5.2. Domain-Wise Temporal Analysis of LLMs Performance

To assess the statistical significance of temporal performance variation, independent samples t-tests were conducted comparing theme-level classification accuracy across three discourse periods: pre-COVID (2017–2019), COVID-period (2020–2021), and post-COVID (2022–2023). The results confirm statistically significant performance declines during the COVID period for all four models. ChatGPT exhibited the most robust decline, with mean accuracy falling from 77.94% to 71.69% (t = 4.065, p < 0.001), Gemini from 73.86% to 70.07% (t = 2.278, p = 0.023), and LLaMA from 38.70% to 33.08% (t = 2.047, p = 0.041). DeepSeek presents a contrasting pattern, with accuracy increasing significantly from 88.10% to 90.91% during the COVID period (t = −2.506, p = 0.012), consistent with, though not conclusively attributable to, its reinforcement learning-based training design, given that architectural and data differences across models were not experimentally controlled. In the post-COVID period, no model exhibited statistically significant change, indicating performance stabilization across all architectures.
The domain-wise temporal analysis further indicates that model performance is jointly shaped by discourse dynamics and architectural biases (Figure 7). DeepSeek demonstrated consistently strong performance in lexically stable domains such as climate change, terrorism, and UN Reform, achieving 100% accuracy in multiple years (2014, 2015, 2020, and 2023). However, its accuracy declined sharply to 25.64% in human rights in 2021, coinciding with the rapid rhetorical transformation introduced by COVID-19 and refugee crisis discourse, where rights-based language was reframed around health emergencies and humanitarian migration—contexts that diverge substantially from the stable policy vocabulary on which the model performs reliably.
ChatGPT performed strongly in humanitarian and conflict-related domains, scoring 95.56% in conflict resolution in 2020 and 100% in terrorism across multiple years, consistent with its broad pre-training on humanitarian law, treaty texts, and global security discourse [69]. However, it demonstrated persistent weakness in international and multilateral cooperation, with accuracy dropping to 23.81% in 2023, coinciding with the fragmented diplomatic rhetoric of the Russia–Ukraine conflict period, where terms such as coordination, solidarity, and partnership were deployed interchangeably across domains.
Table 7. The timeline analysis (2014–2023) of LLM accuracies shows the pattern of performance during a decade of UNGA debates.
This pattern reflects a structural bias in which the model performs reliably in domains with explicit humanitarian framing but struggles where meaning is determined by pragmatic and situational context rather than stable lexical markers.
Gemini exhibited high and stable performance in normatively codified domains, reaching 98.08% in human rights in year 2015 and 100% in sustainable development between 2021 and 2023, reflecting strong alignment with rights-based and developmental discourse characterized by consistent terminology such as fundamental freedoms, equality, and sustainable development goals. Conversely, it consistently underperformed in international cooperation, with accuracy frequently falling below 40% during 2019–2021, when overlapping crises involving climate change, refugee migration, and the COVID-19 pandemic obscured thematic boundaries in cooperation-related discourse through the use of abstract diplomatic registers such as shared responsibility and collective frameworks that lack stable lexical anchoring.
LLaMA produced the weakest and most unstable results across all domains, with accuracy approaching zero in international cooperation and governance-related categories, consistent with its limited parameter scale and constrained capacity for discourse-level abstraction and long-context reasoning. Its isolated strong performance in lexically restricted domains such as terrorism (95.23% in 2015, 100% in 2021 and 2023) and sustainable development confirms reliance on superficial keyword recognition rather than deeper semantic understanding, a limitation that became particularly evident during 2019–2021 when overlapping crises required nuanced thematic interpretation beyond its representational capacity.
Collectively, these findings demonstrate that temporal performance instability in LLM-based diplomatic discourse classification is statistically verifiable, discourse-driven, and systematically modulated by model architecture. Performance degradation is concentrated in periods of rhetorical disruption and domains characterized by abstract institutional language, while recovery and stability are determined by the degree to which model architectures support contextual reasoning beyond surface-level lexical pattern matching.

6. Discussion

The results indicate the potential and limitations of LLM-based thematic classification in political discourse. Although LLMs can effectively handle large volumes of text, human interpretation remains essential for ambiguous instances, particularly where diplomatic language is context-dependent and pragmatically nuanced.
Table 8. Lexical frequency shifts in UNGA speeches (2019–2021).
The findings go beyond performance comparison to offer theoretically grounded insights: the systematic concentration of misclassification errors in semantically overlapping domains suggests that current LLMs lack the pragmatic reasoning capacity required for policy-level text analysis, while the observed correlation between geopolitical disruption and classification degradation introduces a novel temporal robustness dimension to LLM evaluation methodology, advancing knowledge beyond conventional static benchmarking.
This study compares LLM-labeled outputs with human-annotated domain judgment and finds that LLMs have the potential to go beyond being generative tools and are increasingly competent enough to perform semantic tasks traditionally equated to expert judgment. Model comparisons reveal architectural advantages and weaknesses. DeepSeek R1 outperforms in technical domains, which may plausibly reflect its reinforcement learning approach, though architectural differences, training data composition, and parameter scale are confounding factors that the current design cannot disentangle. ChatGPT performs steadily across domains due to its large training diversity, Gemini’s diverse results suggest it has fewer diplomatic sources, and LLaMA instruct small scale (1B parameters), included explicitly as a lightweight baseline, performs as expected for its parameter scale. Its results are not intended to characterize the LLaMA model family broadly, and future work should incorporate a larger LLaMA variant, such as LLaMA 3.1 70B, to enable a more informative comparison at equivalent scale.
The error analysis identified three limitations in LLM performance. Firstly, there was systematic misinterpretation of politically sensitive terminologies, models often misclassifying sovereignty concerns under human rights instead of political sovereignty. Secondly, the models handled semantically overlapping domains inconsistently, especially having problems identifying the difference between international cooperation and multilateral cooperation (with F1-scores as low as 0.41 and 0.24 for some of the models). This confusion reflects the models’ incapacity to capture subtle policy distinctions that are obvious to human experts. Thirdly, all models also indicated a drop in performance when presented with complicated, multidimensional themes which covered multiple policy areas at once, suggesting contextual understanding is adversely affected when themes crossed domain boundaries. The typology of systematic errors (semantic overlap, political sensitivity, multidimensional themes) is a new contribution since previous studies on UNGA did not examine model errors beyond accuracy scores.
This study adds to the growing literature on computational political discourse analysis by providing benchmark metrics for evaluating the performance of LLM diplomatic text classification. The reliability of gold-standard annotations is confirmed by our inter-annotator agreement analysis (0.68–1.00 across nine domains) but moderate agreement in overlapping categories is due to the inherent semantic complexity of diplomatic discourse.
Figure 7. Temporal analysis of LLMs on UNGA debates (2014–2023) showing domain-specific stability and variation.
The gold-standard annotation created will be a useful resource for future research in computational analysis of international diplomatic discourse. While the present evaluation design introduces an inherent degree of methodological circularity, in that the themes used for classification were initially extracted by ChatGPT and subsequently by fine-tuned Gemini, it is important to note that this circularity is partially mitigated by the human validation layer embedded in the pipeline. Specifically, all LLM-extracted themes were independently mapped to domains by human expert annotators with a background in political science and computational linguistics, ensuring that the final gold-standard labels reflect expert judgment rather than raw LLM output. Furthermore, the 12-domain taxonomy was defined a priori based on established UN policy frameworks, independent of any LLM-generated content. The evaluation therefore measures LLM classification performance against a human-validated, theoretically grounded schema rather than a purely LLM-mediated construct.
However, future work should explore pipelines initiated from human-generated or theory-driven theme sets to establish a fully independent gold standard. To partially address this concern within the current study, we conducted a supplementary validation using a subset of 150 themes independently generated by the expert political science annotator, without reference to any LLM output. Domain classification performance of all four models on this human-generated subset was compared against their performance on the LLM-derived subset. DeepSeek R1 achieved 74.3% accuracy on the human-generated subset versus 77.0% on the LLM-derived subset; ChatGPT scored 71.2% versus 73.0%; and Gemini scored 68.4% versus 70.0%; and LLaMA scored 33.8% versus 36.0%. The marginal performance difference (approximately 2–3 percentage points across all models) suggests that while LLM-extraction bias cannot be ruled out entirely, it does not substantially inflate classification results. Critically, the relative ranking of models and the pattern of domain-wise strengths and weaknesses remain consistent across both subsets, indicating that the reported findings are not an artifact of circular label construction. Nevertheless, we acknowledge that a fully independent, expert-generated gold standard remains the ideal benchmark, and we encourage future work in this direction.

Comparative Analysis

The comparative analysis of this study reveals the performance of various LLMs in formal policy areas that align with the literature. The finding shows that DeepSeek R1 outperformed other models, especially in structured domains like terrorism and counter-terrorism (F1 = 0.96), climate change (0.94), and UN Reform (0.89). This aligns with [60], representing the finding that DeepSeek has a higher classification accuracy due to (MoE) architecture which enhances domain specific performance.
In contrast, ChatGPT demonstrated robust generalist performance with an accuracy of 0.90 on climate change and F1-scores ranging between 0.77 and 0.84 on human rights, economic development, and conflict resolution. Such results indicate its consistent performance as observed by [21,67], which shows 0.86–0.89 accuracy on climate-related classifications. However, LLaMA performed worse in all domains, with the lowest scores of 0.15 on climate change and 0.02 on international cooperation. These findings are in line with the result of [60] that LLaMA-3.1 had an average classification error 0.511 (49% accuracy), compared to DeepSeek with only a 0.286 error rate. Compared with the 72% accuracy of the existing study [34,70] for the classification of political content, ChatGPT performed better in our domain-specific classification of UNGA discourse, attaining a higher accuracy (0.73), indicating its increased reliability in structured political domains. Furthermore, Gemini showed almost the same performance on most of the domains as ChatGPT (e.g., 0.84 on climate change, 0.75 on human rights). The precision of Gemini is dataset-sensitive [39,69,71], whereas GPT-4 demonstrates better generalization, especially on the tasks involving reasoning [33,43]. Despite that, ChatGPT-4 models are always outperformed in terms of coherence and relevance by LLaMA and Gemini [26,70,72,73]. A domain-wise comparison in the existing study [16], used a semi-supervised approach through seed word dictionaries, Naive Bayes and LDA classification techniques to analyze UNGA debates. The LLM-based method yields superior performance in all domains with a wider thematic scope, as shown in Table 9.
Table 9. Domain-wise F1-scores comparison with existing study classification techniques: Naive Bayes, LDA vs. our LLMs approach.

7. Conclusions

This paper proposes a structured evaluation of large language models (LLMs) for thematic domain classification of the United Nations General Assembly speeches. We have adopted a two-step procedure: we extracted themes with fine-tuned LLMs and then assigned them to pre-defined domains through a zero-shot classification. The outputs were compared against expert human annotations on twelve diplomatic categories. We compared the interpretive alignment and domain-wise consistency of human and machine mappings, in addition to the standard classification metrics of accuracy and F1-score.
The findings showed that full-scale LLMs, specifically DeepSeek R1 and ChatGPT, can perform relatively well in achieving expert-level domain classification, particularly in lexically stable categories such as climate change, where domain boundaries are clearly defined. However, performance declined notably in semantically overlapping domains, highlighting the interpretive limits of current models in diplomatically nuanced discourse. The lightweight LLaMA 3.2 1B baseline, included as a resource-constrained reference point, confirms that parameter scale meaningfully constrains classification performance in complex diplomatic text, and its results should not be taken as reflective of broader LLM capability for this task. The results imply that although LLMs exhibit significant ability to reproduce expert-level thematic classification in structured political language, this capability remains contingent on domain clarity and terminological stability. The point of divergence between human and LLM classification is always attributable to semantic overlap and lack of pragmatic reasoning, which is most pronounced in the domain of international cooperation and multilateral cooperation. Moreover, classification performance proves sensitive to geopolitical disruptions, declining most notably during 2019–2021 coinciding with COVID-19 discourse, while DeepSeek R1 recovered strongly during the Russia–Ukraine conflict period in 2022–2023 (84.81%), a pattern consistent with its reinforcement learning-based design, though this interpretation remains correlational given that multiple architectural factors differ across the evaluated models. Future work will investigate domain-specific fine-tuning to enhance classification performance in overlapping policy areas and to increase the predictability of models in multifaceted diplomatic situations. Applications to other international forums would also support the generalizability of this approach. A methodological limitation of the current pipeline is the absence of an ablation study comparing reverse chronological fine-tuning against forward fine-tuning and zero-shot extraction baselines. Future work should empirically validate whether the temporal directionality of fine-tuning meaningfully improves theme extraction quality, particularly during periods of rhetorical disruption such as the COVID-19 pandemic years (2020–2021).

Author Contributions

F.M.: Conceptualization, Methodology, Software, Validation, Investigation, Data Curation, Writing—Original Draft Preparation, Visualization. S.A.R.: Conceptualization, Validation, Resources, Writing—Review & Editing, Supervision, Project Administration. S.I.N.: Conceptualization, Writing—Review & Editing, Formal Analysis, Data Curation. M.G.A.M.: Formal Analysis, Writing—Review & Editing. M.I.: Formal Analysis, Resources, Data Curation, Project Administration. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Data Availability Statement

The data supporting the findings of this study are publicly available. The UN General Assembly (UNGA) speech dataset can be accessed via the Harvard Dataverse (https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/0TJX8Y) (accessed on 28 April 2026). Further details are available from the corresponding author upon reasonable request.

Acknowledgments

The authors would like to acknowledge the support of Prince Sultan University for paying the Article Processing Charges (APC) of this publication. Most of the work was conducted while the authors were affiliated with the Department of Computer Science, Fatima Jinnah Women University, Rawalpindi, Pakistan.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Wojciechowska, J.; Sypniewski, M.; Śmigielska, M.; Kamiński, I.; Wiśnios, E.; Schreiber, H.; Pieliński, B. Deep Dive into the Language of International Relations: NLP-based Analysis of UNESCO’s Summary Records. arXiv 2023, arXiv:2307.16573. [Google Scholar]
  2. Gielens, E.; Sowula, J.; Leifeld, P. Goodbye human annotators? Content analysis of social policy debates using ChatGPT. J. Soc. Policy 2025, 62, 385–404. [Google Scholar] [CrossRef] [Scilit]
  3. Baturo, A.; Dasandi, N. What drives the international development agenda? An NLP analysis of the united nations general debate 1970–2016. In Proceedings of the 2017 International Conference on the Frontiers and Advances in Data Science (FADS); IEEE: New York, NY, USA, 2017; pp. 171–176. [Google Scholar]
  4. Sawicki, J.; Ganzha, M.; Paprzycki, M. The state of the art of natural language processing—A systematic automated review of NLP literature using NLP techniques. Data Intell. 2023, 5, 707–749. [Google Scholar] [CrossRef] [Scilit]
  5. Gurciullo, S.; Mikhaylov, S. Topology analysis of international networks based on debates in the United Nations. arXiv 2017, arXiv:1707.09491. [Google Scholar] [CrossRef] [Scilit]
  6. Sousa, S.; Jantscher, M.; Kröll, M.; Kern, R. Large language models for electronic health record de-identification in english and german. Information 2025, 16, 112. [Google Scholar] [CrossRef] [Scilit]
  7. Khan, A.H.; Kegalle, H.; D’Silva, R.; Watt, N.; Whelan-Shamy, D.; Ghahremanlou, L.; Magee, L. Automating Thematic Analysis: How LLMs Analyse Controversial Topics. arXiv 2024, arXiv:2405.06919. [Google Scholar] [CrossRef] [Scilit]
  8. Herbst, R.; Stille, D.; Gwilliam, G.F.; Von Konrat, M.; Campbell, T.; Gaswick, W.; Grewe, F.; Hansen, K.; Iacobelli, F.; Jellema, K.; et al. Unlocking the past: The potential of large language models to revolutionize transcription of natural history collections. Data Intell. 2025, 7, 237–264. [Google Scholar] [CrossRef] [Scilit]
  9. Voeten, E. Data and analyses of voting in the United Nations: General Assembly. In Routledge Handbook of International Organization; Routledge: Abingdon, UK, 2013; pp. 54–66. [Google Scholar]
  10. Mesquita, R.; Pires, A. What are UN General Assembly resolutions for? Four views on parliamentary diplomacy. Int. Stud. Rev. 2023, 27, viac058. [Google Scholar] [CrossRef] [Scilit]
  11. Mesquita, R.; Pires, A. The references of the nations: Introducing a corpus of United Nations General Assembly resolutions since 1946 and their citation network. J. Peace Res. 2024, 69, 1279–1291. [Google Scholar] [CrossRef] [Scilit]
  12. Grzyb, M.; Krzyziński, M.; Sobieski, B.; Spytek, M.; Pieliński, B.; Dan, D.; Wróblewska, A. Mining United Nations General Assembly Debates. arXiv 2024, arXiv:2406.13553. [Google Scholar] [CrossRef] [Scilit]
  13. Jankin, S.; Baturo, A.; Dasandi, N. Words to unite nations: The complete United Nations General Debate Corpus, 1946–present. J. Peace Res. 2024, 69, 1339–1351. [Google Scholar] [CrossRef] [Scilit]
  14. Liang, Y.; Yang, L.; Wang, C.; Xia, C.; Meng, R.; Xu, X.; Wang, H.; Payani, A.; Shu, K. Benchmarking llms for political science: A united nations perspective. arXiv 2025, arXiv:2502.14122. [Google Scholar] [CrossRef] [Scilit]
  15. Olivier, C.; Álvarez, L.M. The United Nations General Assembly. In Las Respuestas a la Corrupción Desde el Derecho Internacional de Los Derechos Humanos; Tirant lo Blanc: Valencia, Spain, 2025; p. 71. [Google Scholar]
  16. Watanabe, K.; Zhou, Y. Theory-driven analysis of large corpora: Semisupervised topic classification of the UN speeches. Soc. Sci. Comput. Rev. 2022, 46, 346–366. [Google Scholar] [CrossRef] [Scilit]
  17. Gurciullo, S.; Mikhaylov, S.J. Detecting policy preferences and dynamics in the un general debate with neural word embeddings. In Proceedings of the 2017 International Conference on the Frontiers and Advances in Data Science (FADS); IEEE: New York, NY, USA, 2017; pp. 74–79. [Google Scholar]
  18. Baturo, A.; Gray, J. Leaders in the United Nations General Assembly: Revitalization or Politicization? Rev. Int. Organ. 2024, 19, 721–752. [Google Scholar] [CrossRef] [Scilit]
  19. Hosli, M.O.; Kantorowicz, J. The European Union in the United Nations: An Analysis of General Assembly Debates. In Proceedings of the 14th Annual Conference on the Political Economy of International Organization (PEIO14), Online, 7–9 July 2022. [Google Scholar]
  20. Dai, S.C.; Xiong, A.; Ku, L.W. LLM-in-the-loop: Leveraging large language model for thematic analysis. arXiv 2023, arXiv:2310.15100. [Google Scholar]
  21. Mumtaz, S.; Sanam, N.; Mumtaz, F. AI-Powered security classification of generator-matrix s-box for post-quantum lightweight cryptography. Telecommun. Syst. 2026, 89, 20. [Google Scholar] [CrossRef] [Scilit]
  22. De Paoli, S. Performing an inductive thematic analysis of semi-structured interviews with a large language model: An exploration and provocation on the limits of the approach. Soc. Sci. Comput. Rev. 2024, 49, 997–1019. [Google Scholar] [CrossRef] [Scilit]
  23. Drápal, J.; Westermann, H.; Savelka, J. Using Large Language Models to Support Thematic Analysis in Empirical Legal Studies. In Proceedings of the JURIX, Maastricht, The Netherlands, 18–20 December 2023; pp. 197–206. [Google Scholar]
  24. Gallegos, I.O.; Rossi, R.A.; Barrow, J.; Tanjim, M.M.; Kim, S.; Dernoncourt, F.; Yu, T.; Zhang, R.; Ahmed, N.K. Bias and fairness in large language models: A survey. Comput. Linguist. 2024, 57, 1097–1179. [Google Scholar] [CrossRef] [Scilit]
  25. Jansen, J.A.; Manukyan, A.; Al Khoury, N.; Akalin, A. Leveraging large language models for data analysis automation. PLoS ONE 2025, 20, e0317084. [Google Scholar] [CrossRef] [Scilit]
  26. Ahmed, R.; Rauf, S.A.; Latif, S. Leveraging large language models and prompt settings for context-aware financial sentiment analysis. In Proceedings of the 2024 5th International Conference on Advancements in Computational Sciences (ICACS); IEEE: New York, NY, USA, 2024; pp. 1–9. [Google Scholar]
  27. Jelodar, H.; Wang, Y.; Yuan, C.; Feng, X.; Jiang, X.; Li, Y.; Zhao, L. Latent Dirichlet allocation (LDA) and topic modeling: Models, applications, a survey. Multimed. Tools Appl. 2019, 78, 15169–15211. [Google Scholar] [CrossRef] [Scilit]
  28. Dicle, M.F.; Dicle, B. Content analysis: Frequency distribution of words. Stata J. 2018, 18, 379–386. [Google Scholar]
  29. Turco, L.R. Speaking Volumes: Introducing the UNGA Speech Corpus. Int. Stud. Q. 2024, 26, sqae001. [Google Scholar] [CrossRef] [Scilit]
  30. Khan, S.; Ahmed, F.; Mubeen, M. A text-mining research based on LDA topic modelling: A corpus-based analysis of Pakistan’s UN assembly speeches (1970–2018). Int. J. Humanit. Arts Comput. 2022, 16, 214–229. [Google Scholar] [CrossRef] [Scilit]
  31. Curley, C.; Siapera, E.; Carthy, J. Covid-19 protesters and the far right on telegram: Coconspirators or accidental bedfellows? Soc. Media+ Soc. 2022, 8, 20563051221129187. [Google Scholar] [CrossRef] [Scilit]
  32. Rauf, S.A.; Yvon, F. Translating scientific abstracts in the bio-medical domain with structure-aware models. Comput. Speech Lang. 2024, 87, 101623. [Google Scholar] [CrossRef] [Scilit]
  33. Rizinski, M.; Jankov, A.; Sankaradas, V.; Pinsky, E.; Mishkovski, I.; Trajanov, D. Comparative analysis of NLP-based models for company classification. Information 2024, 15, 77. [Google Scholar] [CrossRef] [Scilit]
  34. Orellana, S.; Bisgin, H. Using natural language processing to analyze political party manifestos from New Zealand. Information 2023, 14, 152. [Google Scholar] [CrossRef] [Scilit]
  35. Peña, A.; Morales, A.; Fierrez, J.; Serna, I.; Ortega-Garcia, J.; Puente, I.; Cordova, J.; Cordova, G. Leveraging large language models for topic classification in the domain of public affairs. In Proceedings of the International Conference on Document Analysis and Recognition; Springer: Berlin/Heidelberg, Germany, 2023; pp. 20–33. [Google Scholar]
  36. Gujral, P.; Awaldhi, K.; Jain, N.; Bhandula, B.; Chakraborty, A. Can LLMs Help Predict Elections?(Counter) Evidence from the World’s Largest Democracy. arXiv 2024, arXiv:2405.07828. [Google Scholar]
  37. Gupta, R. Comparative Analysis of DeepSeek R1, ChatGPT, Gemini, Alibaba, and LLaMA: Performance, Reasoning Capabilities, and Political Bias. Authorea Prepr. 2025. [Google Scholar] [CrossRef] [Scilit]
  38. Afzal, T.; Abdul Rauf, S.; Malik, M.G.A.; Imran, M. Fine-tuning QurSim on monolingual and multilingual models for semantic search. Information 2025, 16, 84. [Google Scholar] [CrossRef] [Scilit]
  39. Mollo, D.C.; Millière, R. The vector grounding problem. arXiv 2023, arXiv:2304.01481. [Google Scholar]
  40. Kim, E.; Suk, J.; Kim, S.; Muennighoff, N.; Kim, D.; Oh, A. Llm-As-An-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation. arXiv 2024, arXiv:2412.10424. [Google Scholar]
  41. Ruifan, L.; Zhiyu, W.; Yuantao, F.; Shuqin, Y.; Guangwei, Z. Enhanced Prompt Learning for Few-shot Text Classification Method. Acta Sci. Nat. Univ. Pekin. 2024, 67, 1–12. [Google Scholar]
  42. Le Mens, G.; Gallego, A. Positioning political texts with Large Language Models by asking and averaging. Political Anal. 2024, 37, 274–282. [Google Scholar]
  43. Heseltine, M.; Clemm von Hohenberg, B. Large language models as a substitute for human experts in annotating political text. Res. Politics 2024, 11, 20531680241236239. [Google Scholar] [CrossRef] [Scilit]
  44. Bi, X.; Chen, D.; Chen, G.; Chen, S.; Dai, D.; Deng, C.; Ding, H.; Dong, K.; Du, Q.; Fu, Z.; et al. Deepseek llm: Scaling open-source language models with longtermism. arXiv 2024, arXiv:2401.02954. [Google Scholar]
  45. Guo, D.; Yang, D.; Zhang, H.; Song, J.; Zhang, R.; Xu, R.; Zhu, Q.; Ma, S.; Wang, P.; Bi, X.; et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv 2025, arXiv:2501.12948. [Google Scholar]
  46. Fjelstul, J.; Hug, S.; Kilby, C. Decision-making in the United Nations General Assembly: A comprehensive database of resolution-related decisions. Rev. Int. Organ. 2025, 271, 3. [Google Scholar]
  47. Assembly, U.G. Declaration on Principles of International Law Concerning Friendly Relations and Cooperation Among States in Accordance with the Charter of the United Nations; United Nations General Assembly: New York, NY, USA, 2025. [Google Scholar]
  48. Roumeliotis, K.I.; Tselikas, N.D. Chatgpt and open-ai models: A preliminary review. Future Internet 2023, 15, 192. [Google Scholar] [CrossRef] [Scilit]
  49. Islam, R.; Ahmed, I. Gemini-the most powerful LLM: Myth or Truth. In Proceedings of the 2024 5th Information Communication Technologies Conference (ICTC); IEEE: New York, NY, USA, 2024; pp. 303–308. [Google Scholar]
  50. United Nations. Transforming Our World: The 2030 Agenda for Sustainable Development. 2015. Available online: https://sdgs.un.org/2030agenda (accessed on 31 July 2025).
  51. United Nations. Our Common Agenda: Report of the Secretary-General, 2021. Available online: https://digitallibrary.un.org/record/3939309?v=pdf (accessed on 31 July 2025).
  52. Joshi, S. A Comprehensive Review of DeepSeek: Performance, Architecture and Capabilities. Preprints 2025. [Google Scholar] [CrossRef] [Scilit]
  53. Biswas, S.S. Role of chat gpt in public health. Ann. Biomed. Eng. 2023, 58, 868–869. [Google Scholar] [CrossRef] [Scilit]
  54. Xin, Q.; Nan, Q. Enhancing inference accuracy of llama llm using reversely computed dynamic temporary weights. Authorea Prepr. 2024. [Google Scholar] [CrossRef] [Scilit]
  55. Li, L.H.; Nugroho, A.C. Evaluating Large Language Models: Challenges, Limitations, and Future Directions. In Proceedings of the 2025 Joint International Conference on Digital Arts, Media and Technology with ECTI Northern Section Conference on Electrical, Electronics, Computer and Telecommunications Engineering (ECTI DAMT & NCON); IEEE: New York, NY, USA, 2025; pp. 29–34. [Google Scholar]
  56. Mercer, S.; Spillard, S.; Martin, D.P. Brief analysis of DeepSeek R1 and its implications for Generative AI. arXiv 2025, arXiv:2502.02523. [Google Scholar] [CrossRef] [Scilit]
  57. Rahman, A.; Mahir, S.H.; Tashrif, M.T.A.; Aishi, A.A.; Karim, M.A.; Kundu, D.; Debnath, T.; Moududi, M.A.A.; Eidmum, M. Comparative analysis based on deepseek, chatgpt, and google gemini: Features, techniques, performance, future prospects. arXiv 2025, arXiv:2503.04783. [Google Scholar] [CrossRef] [Scilit]
  58. Naidu, G.; Zuva, T.; Sibanda, E.M. A review of evaluation metrics in machine learning algorithms. In Proceedings of the Computer Science On-Line Conference; Springer: Berlin/Heidelberg, Germany, 2023; pp. 15–25. [Google Scholar]
  59. Michaud, E.J.; Liu, Z.; Tegmark, M. Precision machine learning. Entropy 2023, 27, 175. [Google Scholar] [CrossRef] [Scilit]
  60. Müller, P.; Brummel, M.; Braun, A. Spatial recall index for machine learning algorithms. In Proceedings of the London Imaging Meeting; Society for Imaging Science and Technology: Springfield, VA, USA, 2021; Volume 2, pp. 58–62. [Google Scholar]
  61. Hinojosa Lee, M.C.; Braet, J.; Springael, J. Performance metrics for multilabel emotion classification: Comparing micro, macro, and weighted f1-scores. Appl. Sci. 2024, 14, 9863. [Google Scholar] [CrossRef] [Scilit]
  62. Sertyesilisik, E.; Sertyesilisik, B. Impacts of the War in Ukraine on Global Sustainable Development and Trade. In International Trade, Economic Crisis and the Sustainable Development Goals; Emerald Publishing Limited: Leeds, UK, 2024; pp. 231–241. [Google Scholar]
  63. Dadhich, M.; Bhati, S.; Bhaskar, A.A.; Kumar, S.R.K.; Gupta, V. Climate Change Initiatives of G20: Analysis of Global Governance and Sustainability. In Community Resilience and Climate Change Challenges: Pursuit of Sustainable Development Goals (SDGs); IGI Global Scientific Publishing: Hershey, PA, USA, 2025; pp. 1–12. [Google Scholar]
  64. Leal-Arcas, R.; Alsheikh, A.; Maqtoush, J.; Althunayan, N.; Dal, H.; Berto, L.A.; Aldughaither, N.; Bashiti, A.; Afaneh, B. Trade, Geopolitics, and Environment. J. Infrastruct. Policy Dev. 2024, 8, 5114. [Google Scholar] [CrossRef] [Scilit]
  65. Liu, A.; Feng, B.; Wang, B.; Wang, B.; Liu, B.; Zhao, C.; Dengr, C.; Ruan, C.; Dai, D.; Guo, D.; et al. Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model. arXiv 2024, arXiv:2405.04434. [Google Scholar]
  66. Zhou, C.; Li, Q.; Li, C.; Yu, J.; Liu, Y.; Wang, G.; Zhang, K.; Ji, C.; Yan, Q.; He, L.; et al. A comprehensive survey on pretrained foundation models: A history from bert to chatgpt. Int. J. Mach. Learn. Cybern. 2024, 16, 9851–9915. [Google Scholar] [CrossRef] [Scilit]
  67. Trajanov, D.; Lazarev, G.; Chitkushev, L.; Vodenska, I. Comparing the performance of Chat-GPT and state-of-the-art climate NLP models on climate-related text classification tasks. In Proceedings of the 4th International Conference on Environmental Design (ICED2023), Athens, Greece, 20–22 October 2023. [Google Scholar]
  68. Kuznetsova, E.; Makhortykh, M.; Vziatysheva, V.; Stolze, M.; Baghumyan, A.; Urman, A. In generative AI we trust: Can chatbots effectively verify political information? J. Comput. Soc. Sci. 2025, 8, 15. [Google Scholar] [CrossRef] [Scilit]
  69. Yang, L.; Xu, S.; Sellergren, A.; Kohlberger, T.; Zhou, Y.; Ktena, I.; Kiraly, A.; Ahmed, F.; Hormozdiari, F.; Jaroensri, T.; et al. Advancing multimodal medical capabilities of Gemini. arXiv 2024, arXiv:2405.03162. [Google Scholar] [CrossRef] [Scilit]
  70. Bogireddy, S.R.; Dasari, N. Comparative analysis of ChatGPT-4 and LLaMA: Performance evaluation on text summarization, data analysis, and question answering. In Proceedings of the 2024 15th International Conference on Computing Communication and Networking Technologies (ICCCNT); IEEE: New York, NY, USA, 2024; pp. 1–7. [Google Scholar]
  71. Akter, S.N.; Yu, Z.; Muhamed, A.; Ou, T.; Bäuerle, A.; Cabrera, Á.A.; Dholakia, K.; Xiong, C.; Neubig, G. An in-depth look at Gemini’s language abilities. arXiv 2023, arXiv:2312.11444. [Google Scholar]
  72. Lamichhane, B. Evaluation of chatgpt for nlp-based mental health applications. arXiv 2023, arXiv:2303.15727. [Google Scholar]
  73. Carlà, M.M.; Gambini, G.; Baldascino, A.; Boselli, F.; Giannuzzi, F.; Margollicci, F.; Rizzo, S. Large language models as assistance for glaucoma surgical cases: A ChatGPT vs. Google Gemini comparison. Graefe’s Arch. Clin. Exp. Ophthalmol. 2024, 262, 2945–2959. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.