Next Article in Journal
FedX: Privacy-Preserving Explainable Federated Ensemble Intrusion Detection System for Edge-Enabled Internet of Vehicles
Previous Article in Journal
A Hybrid PoS–PoW Blockchain Framework for Secure Cyber Threat Intelligence Sharing: Design, Implementation, and Evaluation
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Teaching AI to Decode Vaccine Hesitancy Narratives: A Few-Shot Learning and Topic Modeling Approach

by
Md Enamul Kabir
1,
Shakhawat H. Tanim
2,
Deanna D. Sellnow
1,*,
Geneva Lei P. Luteria
1 and
Lior Rennert
2
1
Social Media Listening Center, Department of Communication, Clemson University, Clemson, SC 29634, USA
2
Center for Public Health Modeling and Response, Department of Public Health Sciences, Clemson University, Clemson, SC 29634, USA
*
Author to whom correspondence should be addressed.
Big Data Cogn. Comput. 2026, 10(5), 159; https://doi.org/10.3390/bdcc10050159
Submission received: 30 November 2025 / Revised: 15 April 2026 / Accepted: 5 May 2026 / Published: 16 May 2026

Abstract

Vaccine hesitancy—which can be defined as a delay in acceptance or the refusal to get vaccinated—has substantially increased over the past decade. This study introduces a computational and qualitative approach designed to efficiently classify stance and uncover narratives in social media discourse without relying on extensive manual annotation. Using 298,356 COVID-19 vaccine-related X posts geolocated to South Carolina (June 2021–May 2022), zero-shot and few-shot learning with instruction-tuned large language models (Mistral-7B, Meta-Llama-3.1, and DeepSeek-7B) was applied for stance detection while Latent Dirichlet Allocation (LDA) was used for topic modeling. The topic modeling identified five dominant themes in vaccine hesitant conversations: skepticism of vaccine efficacy, comparative framing, scientific justification, disapproval of regulations, and distrust. Temporal analysis revealed that skepticism peaked during late 2021, coinciding with booster campaigns and mandate debates. These findings suggest that vaccine hesitancy is influenced through complex rhetorical strategies rather than misinformation alone. These underlying narratives often frame skepticism as rational and evidence-based, using scientific language and statistical reasoning to challenge the effectiveness of vaccines.

1. Introduction

According to the World Health Organization [1], five years into the post-pandemic era, vaccine hesitancy remains a major global health debate [2]. Risk communication scholars argue that, as a global risk issue, vaccine hesitancy may be effectively addressed through glocal risk management (i.e., think globally, act locally). Robertson coined the term glocalization to account for the fact that “the local is essentially included within the global” rather than as separate from it [3] (p. 35). In other words, “globalization involves the linking of localities” and the concept of glocalization very transparently acknowledges this fact [4]. This paper focuses on social media communication among South Carolinians regarding vaccine skepticism as a local microcosm reflecting the larger vaccine hesitancy global health concern.
Existing studies confirm a number of factors influencing South Carolinians’ perceptions regarding vaccine hesitancy. Foremost among them is a growing distrust in vaccine safety spurred by misinformation [5]. A recent study at the Medical University of South Carolina, for example, found that low confidence in science and diminished collective responsibility were the primary drivers of hesitancy in the South Carolina region [6]. Furthermore, research highlights that rural and minority communities in the South face structural barriers such as limited healthcare access and socioeconomic constraints, amplifying hesitancy to receive vaccines [7,8]. Studies also emphasize that vaccine hesitancy is not merely due to a knowledge gap but a complex interplay of cognitive biases, risk perception, and trust in institutions, making localized behavioral research essential for effective interventions [6,9]. Clearly, a major factor contributing to vaccine hesitancy is attributed to what Goldenberg [10] describes as a “crisis of trust.” Scientific institutions are no longer perceived as neutral arbiters of truth but as politically compromised actors. These trends threaten decades of progress in eliminating vaccine-preventable diseases [11,12].
It is critical to address such issues through contextually sensitive strategies to prevent future outbreaks and reduce morbidity and mortality from preventable illnesses. However, efficient tools to monitor these large-scale social media conversations remain scarce since the traditional method of training an algorithm to detect vaccine hesitancy requires time-consuming data annotation. In addition, few studies have explored the deeper narratives that drive the hesitancy. This study addresses this gap by introducing a hybrid computational approach, combining “zero-shot and few-shot learning” with large language models (LLMs) and Latent Dirichlet Allocation (LDA) topic modeling for capturing the narratives in the discourse. The selection of this approach was driven by the following rationale. First, few-shot and zero-shot learning offer flexibility by allowing dynamic adjustment of rules and examples without retraining the model, enabling rapid adaptation to evolving research needs. Second, social media language is fluid and context-dependent; few-shot learning accommodates nuanced interpretations—such as sarcasm or quoted text—through curated examples embedded in the prompt. Finally, with few-shot classification rules and examples directly in the prompt, the process becomes transparent and easily replicable by other researchers, strengthening methodological rigor [13]. The research was guided by three objectives: (1) to use LLM/AI to efficiently detect vaccine hesitancy in big data, (2) to capture the underlying narratives in vaccine hesitancy discourse, and (3) to explore the temporal patterns of vaccine hesitancy narratives.

1.1. Vaccine Hesitancy Detection Strategies

The complexity in vaccine hesitancy detection arises from the informal, fast-changing nature of conversation on social media that contains sarcasm, misinformation, mixed signals, and grammatically ill-structured content. Yet, researchers rely heavily on such social media data, as it offers the raw nature of rhetorics people use to justify their hesitancy. Sasse et al. [14] demonstrated that social media data can serve as a reliable proxy for predicting hesitancy trends. Furthermore, Li et al. [15] and Limaye et al. [16] showed that theory-driven interventions delivered through social platforms can improve vaccine confidence. Past studies captured common features of vaccine hesitancy-related posts. For example, emotionally charged, factually accurate but negatively framed content is often preferred over outright falsehoods [17]. Taubert et al. [18] further emphasized that conspiracy narratives and distrust in institutions compound hesitancy beyond simple misinformation. Neff et al. [19] also addressed the limited understanding of platform-specific features and the dynamics of misinformation, which highlights the need for more nuanced strategies for vaccine hesitancy detection.
Classification of text as supportive, opposing, or neutral toward a specific issue has emerged as a key method for analyzing vaccine discourse. Glandt et al. [20] introduced COVID-19 stance datasets and demonstrated baseline performance using transformer-based models. Balaji et al. [21] applied Siamese BERT networks for stance classification on vaccine-related tweets. Wang et al. [22] proposed a comprehensive stance detection dataset (VSDQ) and leveraged chain-of-thought reasoning for improved classification. Barberia et al. [23] reviewed stance detection studies, cautioning against measurement bias and emphasizing the need for robust annotation protocols.

1.2. Few-Shot and Zero-Shot Learning for Stance Detection

Few-shot learning addresses the scarcity of labeled data in stance detection. Brown et al. [24] introduced in-context learning as a paradigm for adapting large models with minimal examples, inspiring subsequent work. Recent frameworks such as MLSD [25] employ metric learning for cross-target stance detection, achieving significant improvements in macro-F1 scores. Khiabani and Zubiaga [26] proposed multimodal approaches combining textual and social network features, outperforming text-only baselines in few-shot settings. Liu et al. [27] introduced target-aware contrastive learning and consistency regularization to enhance few-shot stance detection. In recent days, LLMs such as GPT, Llama, and Mistral have revolutionized such text classification tasks. Kostina et al. [28] found that LLMs outperform traditional models in complex classification tasks, though inference costs remain high. Dos Santos et al. [29] reported that Llama3 achieved state-of-the-art precision and F1 scores across multiple domains. Recent studies concluded that few-shot learning is practical for quick adaptation, considering that parameter-efficient fine-tuning remains essential for achieving state-of-the-art accuracy while balancing computational costs [22,30].

1.3. Topic Modeling for Capturing Narratives

Computational approaches, particularly topic modeling, have emerged as effective tools for uncovering latent themes in large-scale text data. Latent Dirichlet Allocation (LDA) has been widely applied to social media datasets to identify recurring frames such as concerns about side effects, autonomy in health decisions, and skepticism toward pharmaceutical companies [31,32]. Studies on YouTube, Reddit, and X confirm that LDA can reveal narratives such as fear of side effects and distrust in government messaging [32,33,34]. These studies demonstrate the potential of topic modeling to reveal thematic structures that go beyond traditional analysis.
However, most previous works remained limited to quantitative outputs, relying on coherence scores and keyword lists to interpret topics. Sievert and Shirley [35] cautioned that such metrics alone may not guarantee meaningful interpretation. Algorithmically derived topics often lack contextual clarity, which can be mitigated by qualitative validation to ensure that computational findings align with real-world narratives. The present study advances this methodological frontier by integrating LDA-based topic modeling with qualitative thematic analysis, which combines each topic cluster as a separate document. This approach allows us to move beyond surface-level categorization and engage deeply with the rhetorical dimensions of vaccine discourse. By examining representative posts within each topic cluster, the present study offers an interpretive landscape that enhances validity and supports actionable insights for vaccine discourse. Interpretability is critical for stance detection and health communication research. Past studies advocated for an explainable AI framework where human-annotated rationales can further enhance transparency and trust in automated systems [36,37]. Thus, the dual-layered design employed in this study contributes to methodological hybridity in computational communication [38,39]. Alongside introducing this efficient methodological framework, the study also sets forth research questions designed to guide the in-depth qualitative exploration of the narratives underlying vaccine hesitancy discourse.
RQ1: Which topics dominated vaccine hesitancy discussions in South Carolina?
RQ2: How did the vaccine hesitancy discourse shift over time?

2. Materials and Methods

To combat the expensive and time-consuming data annotation challenges, this study employed zero-shot and few-shot machine learning and an LDA topic model approach to capture vaccine hesitancy discussions towards COVID-19 vaccines.

2.1. Data Collection

Initially, COVID-19 vaccine-related tweets posted between 2020 and 2025 were explored. The data was filtered using a Boolean query designed to capture language specifically associated with vaccine hesitancy. A geo-fence was applied in the data pull, which narrowed the dataset to posts generated from only South Carolina. The dataset contains publicly available X posts and was therefore determined to be exempt from IRB review. Geolocation metadata was used only to filter posts originating from South Carolina, and no attempts at user identification were made. All data were stored in a secure, access-restricted drive, and identifiable information was removed during preprocessing. Data were analyzed only in aggregate to ensure user privacy.
The specific publishing time of the posts (narrowed down to the second) was also extracted to examine the temporal shift in conversation. A primary timeline analysis showed that the vaccine hesitancy debate peaked around mid-2021 to mid -2022 (Figure 1). Thus, to capture the peaked conversation period, we shortened our final dataset to June 2021 through May 2022, which contained 298,356 posts. Even though the keywords selected in the query were leaning toward vaccine hesitancy, there is a possibility that some posts that are not vaccine hesitant might have slipped through those shared similar keywords. Thus, the dataset was further segmented to isolate posts expressing vaccine hesitancy.

2.2. Few-Shot and Zero-Shot Learning

Few-shot learning is a natural language processing (NLP) technique and a branch of machine learning where a model is given a small number of labeled examples within the prompt before making a classification [24,40]. Unlike traditional supervised learning, which requires training on large, labeled datasets, few-shot learning leverages the model’s pre-trained knowledge and adapts it to the task with only a handful of examples embedded directly into the query [41,42].
The zero-shot learning method, however, requires no task-specific training data and instead relies on natural language hypothesis templates to assign classifications [43]. The classification process employed a two-stage approach to improve accuracy on noisy social media text. In the first stage, tweets were evaluated for relevance to vaccination using the hypothesis template “This tweet is {}” with candidate labels “about vaccination” and “not about vaccination.” Only tweets classified as relevant proceeded to the second stage, where stance classification was performed using the hypothesis template “The stance of this tweet toward vaccination is {}” with candidate labels “pro-vaccine,” “vaccine hesitant,” and “irrelevant.”
X posts which were identified as non-vaccine-related were automatically assigned the “irrelevant” label in stage one. A zero-shot classification approach was implemented using the DeBERTa-v3-large model fine-tuned for zero-shot tasks [44]. The zero-shot classification pipeline was implemented using the Transformers library [45].

2.3. Instruction-Tuned Large Language Models

Three instruction-tuned causal language models were deployed to perform few-shot in-context learning classification: Mistral-7B-Instruct-v0.1 [46], Meta-Llama-3.1-8B-Instruct [47], and DeepSeek-LLM-7B [48]. These models received identical classification instructions through carefully structured prompts that included system instructions, few-shot examples, and the target tweet. Few-shot prompting has been shown to effectively guide large language models to perform classification tasks without fine-tuning [24,49].

2.3.1. Prompt Design and Iterative Refinement

The few-shot learning method requires a prompt that guides the LLM and a few examples from each class. The initial prompt was a minimal set of examples and label definitions. However, a known limitation of few-shot learning with instruction-tuned large language models (LLMs) is their sensitivity to the exact wording, ordering, and framing of prompts. This challenge was evident in our early prompt iterations. As prior work shows, loosely structured prompts often lead to inconsistent outputs, particularly on short, noisy, and sarcastic social media text [50,51]. Min et al. [50] specifically caution that small variations in prompt wording can substantially change model predictions because LLMs rely heavily on lexical cues embedded in the instructions.
To reduce this prompt sensitivity, we adopted an iterative refinement process that incorporated explicit decision rules and balanced few-shot examples. This design choice follows recommendations from prior research, which finds that structured instructions—with clear definitions, rule hierarchies, and diverse examples—substantially improve consistency [50,51].

2.3.2. Explicit Classification Rules

A structured set of decision rules addressed common ambiguities (e.g., distinguishing between quoted content vs. the author’s stance, handling sarcastic posts, interpreting mentions of “unvaccinated” without assuming hesitancy).

2.3.3. Balanced Few-Shot Examples

The prompt included a mix of clear, borderline, and tricky cases to demonstrate how each category should be applied. Examples covered headlines, sarcastic remarks, anecdotal vaccine experiences, and policy discussions.
The prompt template consisted of a task description defining the three classification categories, followed by six labeled example tweets demonstrating the desired classification behavior. The examples included two pro-vaccine instances (e.g., “I believe in vaccines. They save lives. I got all my doses as soon as I could” labeled as Pro-vaccine), two vaccine-hesitant instances (e.g., “I’m not sure about this vaccine. It came out too fast” labeled as Vaccine hesitant), and two irrelevant instances (e.g., “Breaking: CDC updates travel guidelines for international travelers” labeled as Irrelevant). Each target tweet was then presented to the model following this few-shot example format to generate its classification. A deterministic decoding was applied to ensure that identical prompts produce identical outputs, an important factor for reproducibility in computational social science [39].
For the Mistral-7B and DeepSeek-LLM models, prompts were provided as continuous text sequences. The Llama-3.1 model utilized its native chat template format with explicit role-based message structure to align with its training methodology [52]. All models were implemented using the transformers library [45]. Models were configured to generate brief classification labels (maximum 8 tokens). The combined prompt and tweet content was limited to 2048 tokens, accommodating the full text of all tweets in the dataset.
Classified labels were extracted from model outputs using rule-based pattern matching to identify the three classification categories. The extraction process accommodated minor variations in formatting, capitalization, and spacing. Ambiguous outputs that did not clearly match any category were conservatively assigned to the “irrelevant” category. All labels were standardized to consistent canonical forms (pro-vaccine, vaccine hesitant, or irrelevant) prior to evaluation to enable accurate comparison across models.

2.4. Model Validation

To establish ground-truth labels for model validation, 1000 tweets were randomly sampled using a fixed random seed to ensure reproducibility. The researcher manually classified each tweet in the validation sample into one of three mutually exclusive categories: pro-vaccine (expressing support for vaccination), vaccine-hesitant (expressing doubts, concerns, or opposition to vaccination), or irrelevant (not substantively related to vaccine stance). This manual classification process served as the benchmark against which all automated classification methods were evaluated. (see Appendix A)

2.4.1. Performance Metrics

Model performance was assessed by comparing predicted classifications against the manual annotations by two independent researchers for the 1000-tweet validation sample. For each classification model and prompt variation, we computed multiple performance metrics to provide a comprehensive evaluation. We calculated class-specific accuracy for each of the three categories. For a given class c, class-specific accuracy was computed as:
Class   Accuracy c   =   Number   of   correctly   predicted   instances   of   class   c Total   number   of   true   instances   of   class   c   with   valid   predictions
This metric assessed each model’s ability to correctly identify instances of each stance category independently, revealing potential biases toward or against specific classes. To summarize performance across all three classes, we computed macro-averaged accuracy as the arithmetic mean of the three class-specific accuracies:
Macro - Average   Accuracy =   A c c u r a c y Pro - vaccine   + A c c u r a c y Vaccine   hesitant   +   A c c u r a c y Irrelevant 3
This metric provided equal weight to each class regardless of prevalence in the validation sample, offering a balanced assessment of model performance across all stance categories [53].

2.4.2. Comparison Procedure

Predicted labels for the validation sample were extracted from each model’s output file and aligned with manual classifications using unique message identifiers (UniversalMessageID). All models were evaluated on an identical set of validation instances to enable direct comparison. The evaluation framework allowed for a systematic comparison of zero-shot and instruction-tuned models, including Mistral-7B, DeepSeek-LLM, and Llama-3.1, where Llama-3.1 prediction for vaccine hesitancy produced the highest agreement with human annotators (93%). Intercoder reliability between human annotators and Llama-3.1 was also acceptable (Krippendorff’s α = 0.82 Cohen’s κ = 0.82). Thus, Llama-3.1 was chosen for classification of all 298,356 posts. The agreement for pro-vaccine and irrelevant posts was, respectively, 55%, and 35%. It is important to note that a boolean query was already applied to retrieve posts explicitly related to vaccine hesitancy, substantially narrowing the dataset to content of interest. As the query was designed to predominantly capture vaccine-hesitant posts, any pro-vaccine or irrelevant posts retrieved were incidental and not representative of broader pro-vaccine discourse; accordingly, these posts were excluded from the analysis. The subsequent application of the LLM was therefore to further refine and clean this pre-filtered dataset. The LLM was used primarily as an additional approach for corpus refinement and exclusion, not as a multi-class prediction model. At this stage of the analysis, posts which were not vaccine-hesitant functioned analytically as noise to be excluded, rather than as substantive outcome categories of interest. The model classified 179,876 posts as vaccine hesitant, and subsequent analyses were carried out for a deeper dive into the posts.

2.5. Topic Model

A Latent Dirichlet Allocation (LDA) model was employed for topic modeling. LDA was selected because it was better aligned with the study’s qualitative objectives. While neural topic models perform well in most cases, it was cautioned to generate latent spaces that were less directly interpretable, as their likelihood-driven objectives did not consistently support human-readable topic structures [54]. LDA, by contrast, allowed topics to be validated through coherence metrics and manual inspection, which matched the study’s interpretability-focused workflow [55]. All processing and modeling steps were conducted in Python 3.12 using pandas for data handling, NLTK for text preprocessing and lemmatization [56], and gensim for topic modeling [57] as described below:

2.5.1. Text Cleaning

Before modeling, the text was cleaned to reduce noise. Each post was converted to lowercase. URLs, user mentions (e.g., @username), non-alphabetic characters (numbers, punctuation, and special symbols), and hashtag symbols were removed. This type of normalization is standard in topic modeling for social media data because posts often contain platform artifacts (links, handles, and hashtags) that do not reflect semantic content and can distort inferred topics [38,58].

2.5.2. Tokenization, Stopword Removal, and Lemmatization

After basic cleaning, each post was tokenized into individual word tokens using nltk.word_tokenize() [56]. We then removed stopwords to reduce high-frequency, low-information words. The stopword list combined (a) standard English stopwords from NLTK, (b) built-in English stopwords from scikit-learn, and (c) a custom domain-informed list including conversational fillers (e.g., “yeah,” “lol,” “okay,” “really”), function-style verbs (“got,” “get”), and extremely generic or ubiquitous vaccine terms (e.g., “vaccine,” “COVID,” “COVID-19,” “vax”). Removing such domain-generic tokens prevents the model from letting one obvious keyword dominate every topic and forces separation based on more specific frames [38]. Very short tokens (fewer than 3 characters) were also removed. The remaining tokens were lemmatized using WordNet lemmatization to reduce inflected forms to their base lemma (e.g., “hurting,” “hurt,” “hurts” → “hurt”), which helps merge semantically identical variants into a single feature and stabilizes topic discovery [56].

2.5.3. Collocation Modeling (Bigrams)

To better capture multi-word expressions that behave like single concepts (e.g., “side effect,” “blood clot,” “class action”), a bigram model was trained on the tokenized corpus using gensim. This identifies statistically frequent word pairs and merges them into single tokens such as side_effect or blood_clot [57]. The trained bigram phraser was then applied to each document so the final token list for each post included both unigrams and learned bigrams. Including bigrams improves topic interpretability in short, informal social text because many core ideas are expressed as short noun phrases rather than long explanations.

2.5.4. Dictionary Construction

A gensim dictionary was constructed for mapping each unique token in tokens_bigram [57]. To reduce noise from extremely rare or overly generic words, we applied frequency-based filtering where terms that appeared in fewer than 5 documents and more than 50% of documents were removed. This step limits the vocabulary to words that are informative but not so rare that they represent idiosyncratic one-offs, and not so common that they are effectively background language. Each post was then converted into a standard bag-of-words (BoW) representation. This BoW corpus is the required input format for Latent Dirichlet Allocation [38].
Finally, a Latent Dirichlet Allocation (LDA) model was trained [38] on the full cleaned corpus using gensim.models. LdaModel [57]. LDA is an unsupervised generative model that assumes each document is a mixture of latent topics, and each topic is a probability distribution over words. After training, each learned topic was inspected by examining the top-weighted words (i.e., the most probable terms in that topic).

2.6. Model Quality Assessment

2.6.1. Topic Coherence

We quantified topic interpretability using the coherence metric [55]. Topic coherence measures the semantic consistency of the top terms in each topic by checking how often those terms co-occur together across the corpus. Higher coherence indicates that a topic’s top words tend to appear in similar documents and therefore likely reflect a meaningful shared concept, whereas lower coherence suggests that the topic is mixing unrelated terms [55]. The best possible score generated was 0.34 (see Figure 2). Topic models on noisy and informal language may lead to coherence scores in the 0.30–0.40 range and can still yield theoretically interpretable topics when paired with human qualitative validation and narrative labeling [35,55].

2.6.2. Qualitative Topic Screening

Initially, 10 topics were generated. However, some topics were removed due to low coherence scores and low interpretability. For each topic, the representative posts were extracted, reviewed qualitatively, and assigned thematic labels based on recurring narratives. Two criteria were used to refine the topics: (1) Low interpretability. Topics dominated by generic filler terms or mixed themes that did not reflect a coherent narrative frame. (2) Redundancy/noise. Topics whose top terms and sampled posts did not meaningfully differ from other topics.
Based on the manual review of representative posts for each topic (i.e., reading actual messages most strongly associated with that topic), five topics were removed from further analysis. For example, one of the topics primarily consists of retweeted links to articles or videos rather than original user-generated posts, limiting its interpretive value. In all, these topics were judged to be lexically diffuse and/or semantically noisy rather than representing a distinct public narrative. This qualitative vetting step is common in recent LDA-based works on messy user-generated text. Researchers often focus on more interpretable thematic groupings rather than assume all topics are valid “ground truth” [35,38].

2.6.3. Topic Export for Human Interpretation

To support qualitative coding and narrative labeling, we exported one text packet per topic. For each of the topics in the refined model, we gathered all posts whose dominant topic was that topic and wrote them into a separate .txt file. This “packet per topic” approach allowed a human reviewer to read real posts in their original language, identify the shared stance or storyline, and then summarize that topic. This step anchors the machine-generated clusters in interpretable rhetoric and satisfies the validity in research [35,38].

3. Results

To answer the first research question, RQ1: Which topics dominated vaccine hesitancy discussions in South Carolina, among the 179,876 vaccine-hesitant posts classified by Llama 3.1, the LDA topic model revealed five major topics in which skepticism of vaccine efficacy dominated the discourse (n = 73,122), far surpassing all other themes. Comparative framing (n = 8554) and scientific justification (n = 7823) appeared moderately, while disapproval of regulations (n = 5012) and distrust (n = 4438) were less frequent (Figure 3). Overall, skepticism toward vaccine effectiveness emerged as the central focus of online vaccine-hesitancy discussions. Further exploratory analysis shows word clouds for each dominant topic in the discussion (Figure 4).

3.1. Skepticism of Vaccine Efficacy

3.1.1. Fear of Side Effects

Being one of the dominant discourses under the skepticism towards vaccine efficacy, this subtheme captures how South Carolina users utilizes personal and proximate experiences to challenge effectiveness of COVID-19 vaccine. Rather than engaging with population-level statistics or institutional guidance, users frequently drew on anecdotal, embodied evidence such as stories of illness, breakthrough infections, injury, or death within their personal networks, to assert that vaccines were either ineffective or yielded harmful side effects. For example, one user described learning that two fully vaccinated individuals contracted COVID-19, concluding that vaccine-induced immunity was a “lie” and framing institutional messaging as intentionally deceptive. Another post referenced the death of vaccinated individuals and a claimed vaccine injury to the user’s son, emphasizing that FDA approval “means absolutely nothing.”
“Well I finally know someone with Covid! Well 2! Fully vaccinated from early last year, was up in NY caught Covid! Hopefully it will be milder they are old! More lies people felt immune from vaccine! Will they ever tell the truth!”
“There have been fully vaccinated people that have died from Covid. Do an ounce of research. The vaccine that injured my son was fully FDA approved. That means absolutely nothing.”
However, understanding vaccine hesitancy requires acknowledging that lived experience often carries greater weight than scientific reasoning for individuals making vaccine-related decisions. As prior research has shown, fear of adverse side effects can strongly influence vaccine attitudes even in the presence of scientific reassurance [59]. Especially since COVID-19 pandemic, this phenomenon has been increasingly taking the place as the major factor behind vaccine hesitancy [60,61]. Within this context, posts describing side effects or perceived vaccine failure should not be dismissed solely as misinformation, but understood as expressions of subjective risk calculation and distrust, shaped by personal experience rather than institutional authority.

3.1.2. Counter-Narratives of Transmission

Posts under the subtopic added a counter-narrative of how the virus is being transmitted and spread. It rather attacked those who had received the vaccine and blamed them for spreading COVID-19 and ultimately making the pandemic worse in the process. Examples include: “The fully vaccinated are getting COVID, spreading it, and dying. Enjoy the next year or so, there is no reversing what you have done to yourself.” “So it would appear that the vaccinated are the more likely spreaders of COVID now. Go figure.”

3.1.3. Skepticism About Reporting Accuracy

Many users pointed out that information about vaccination rarely showed the full story. Instead, posts forwarded that primarily the media did not describe the side effects of those who were vaccinated. Examples include: “Anyone else notice that the government-controlled media is reporting about deaths of unvaccinated individuals at fever pitch? Why don’t they ever do stories about the thousands of vaccinated folks who died from Covid, or about the ones who died from taking the Jab?”
“Why is that @Twitter doesn’t report on people who did fall ill after getting a COVID vaccine? Their one-sided reporting on all issues shows that they are just one more provider of useless corporate propaganda.”
“Just use your common sense. Always wash your hands, cover your cough in your elbow, keep your distance from people. I myself am unvaccinated as well, I have had Covid but survived!! Imagine that, something they don’t like to talk about, SURVIVAL RATES AMONG UNVACCINATED”

3.2. Distrust

Users reported a general distrust of entities (e.g., Pfizer, FDA) through naming fears of collusion and corruption. Examples include: “Those refusing vaccines because Big Pharma is unilaterally evil and never to be trusted because of their profit incentive are turning to ineffective COVID treatments manufactured by.. Big Pharma? Ivermectin sales have increased fourfold since 2020. Thats millions in profit.”
The narratives in this topic center on deep skepticism toward major health and regulatory institutions, portraying them as corrupt and profit-driven. These narratives collectively construct a distrustful discourse that undermines confidence in vaccine safety and public health messaging. For example, “Remember when leftists used to pretend to be concerned with big business and corruption, nepotism, revolving door between government and big pharma. This former FDA chief is still on the board of Pfizer, making a lot of money pitching “safe & effective” Covid vaccines on MSM. Pic.twitter.com/QiAbRY9U2l”. Another user maintained, “After the FDA’s approval, Many people will now ask 2 questions: 1. If they approved this, what else has the FDA approved that shouldn’t be approved? 2. What hasn’t the FDA approved that should be approved? The FDA’s blatant corruption today will end up waking up millions.”

3.3. Disapproval of Regulations

Posts under “Disapproval of Regulations” narrate a rejection of policies and regulations (e.g., mask mandates, vaccine requirements) imposed by institutions (e.g., government, schools) to restrict the spread of COVID-19. Users portrayed them as irrational, inconsistent, and politically motivated rather than science-based. Users highlight contradictions, such as rising cases despite strict mandates, to argue that regulations were ineffective and oppressive. They express frustration over excessive control and mock the shifting guidance of authorities. These posts frame the mandates as intrusive overreach that violated common sense and individual freedom. Examples include: “Despite a covid vaccine mandate, an indoor mask mandate, and requiring school kids to eat lunch outside in the cold, New York just set an all time high for covid cases today.”
“The CDC ended masking on Friday, Congress ended masks on Sunday, New York & NYC ended masking & covid vaccine mandates this week. Now California, Washington, & Oregon end school masks today. Joe Biden speaks tomorrow. The “science” was all bullshit.”

3.4. Scientific Justification

This topic reveals an unprecedented perspective in vaccine hesitancy narratives. Posts in the “Scientific Justification” topic predominantly center around a deliberate effort to reclaim scientific authority in defense of vaccine skepticism. Rather than rejecting science outright, many users sought to justify their decision behind the hesitancy by selectively citing studies, preprints, or media outlets that promoted natural immunity and antibodies over vaccines. This rhetorical move transforms vaccine hesitancy into a stance of rational dissent, where they are not “anti-science,” but rather correcting what they view as scientific bias or suppression of inconvenient evidence. Examples include: “Studies spell otherwise: I’ve been meaning to do this for a while. I collated 15 studies (there’s about 60 in total) demonstrating why natural immunity from prior infection is more durable, longer lasting, and more effective. CDC’s position is simply indefensible theblaze.com/op-ed/horowitz…:
“A new systematic review confirms that covid recovered patients have immunity as complete or better vs. reinfection than the vaccines do. There is no scientific justification for vaccine passports/mandates. medrxiv.org/content/10.110…”
The rhetoric in this topic also references figures like Joe Rogan serving as proof that recovery and “natural immunity” work, reinforcing a belief that bodily resilience outperforms manufactured protection. For example, a user wrote, “Much to the disappointment of the coronabros, Joe Rogan recovered from covid in a couple of days and is completely fine. Plus, he now has natural immunity, which studies show is far more protective than vaccinated immunity: outkick.com/joe-rogan-dead…”.

3.5. Comparative Framing

Posts aligned with “Comparative Framing” rely on straightforward contrasts to evaluate vaccine effectiveness. Rather than rejecting numerical evidence outright, users selectively compare case counts across time or place to question expected outcomes of vaccination. These comparisons center on perceived inconsistencies, such as increases in infections following vaccine rollout. Within this broader framing, two recurring forms of comparison emerge: temporal comparisons, which contrast pandemic conditions before and after widespread vaccination, and geographic comparisons, which juxtapose case trends across different locations to express apparent contradictions.

3.5.1. Temporal Comparison

Posts described that despite a rise in vaccinations, the amount of individuals infected with COVID-19 had risen. Examples include: “Why are #COVID deaths approx 10 times greater this summer than this time last year when so many have been vaccinated?” Another user maintained, “Cornell University has five times more Covid cases now than it did this time last year. The school is 95% vaccinated.”

3.5.2. Geographic Comparison

Users also commented on the uptick of illness by comparing rates to other states and countries. Examples include, “Vermont is one of the most vaccinated states in the country and is about to hit an all time high in covid cases. Look forward to Joe Biden and the media telling me how this is all Ron Desantis’s fault.”
“Why is Biden, @CDCgov & US media pushing/coercing every American to get vaccinated at the exact moment the vaxx experiment in Isreal is crashing/burning? Israel, the most vaccinated nation in the world, now has more COVID-19 infections per capita than any country in the world.”
“RT @sternbergh can’t stop staring in disbelief at this map of current global covid hot spots—the U.S., the wealthiest country in the world, where vaccines are freely available, has the highest per capita case rate right now of any country except Mongolia pic.”
To answer the research question RQ2: How did the vaccine hesitancy discourse shift over time, a timeline analysis was performed, which portrays the trends of major vaccine-hesitancy discourses on social media (Figure 5).
Among the topics, skepticism of vaccine efficacy remained the most dominant theme, reaching peaks exceeding 600 posts per day during late 2021, coinciding with heightened public debate surrounding vaccine booster campaigns. The remaining topics exhibited relatively stable and lower volumes, generally below 100 posts per day. Minor spikes in comparative framing and disapproval of regulations were observed towards the end of 2021. The temporal patterns illustrate that doubts about vaccine effectiveness dominated online vaccine-hesitancy conversations, while regulatory and moral arguments appeared more episodic and event-driven. Overall, the timeline analysis revealed that vaccine hesitancy in South Carolina was not a uniform phenomenon but a collection of interrelated rhetorical strands varying with contextual triggers.

4. Discussion

This study introduces a hybrid approach that combines few-shot learning with structured prompts and topic modeling to analyze vaccine hesitancy discourse at scale. Unlike traditional supervised learning, which requires extensive labeled datasets, the present study leverages instruction-tuned large language models and minimal examples to classify stance efficiently. The transition from loosely defined prompts to a rules-driven, few-shot, structured-output system significantly reduced misclassification caused by keyword bias. This refinement addresses two persistent challenges in computational communication research: (1) scaling annotation without sacrificing accuracy and (2) ensuring interpretability for replication. Building on these methodological improvements, the analysis of the topic model reveals how users in South Carolina strategically employed various narratives—some of them acting as parallel counter-narratives rooted in pseudo-science and logical reasoning (albeit fallacious)—to support their vaccine hesitancy decisions.
Among the most prevalent topics, “scientific justification” revealed that users actively engaged with scientific language, cited studies, and employed scientific findings to justify their hesitancy, thereby constructing a resistance of knowledge. This finding challenges a long-standing assumption that vaccine hesitancy stems primarily from ignorance, misinformation, or lack of scientific literacy [6,62]. Instead, these narratives illustrate how individuals are reframing scientific justification to legitimize their hesitancy, positioning themselves as informed, evidence-seeking citizens rather than ignorant skeptics. Thus, vaccine hesitancy is not merely a deficit of knowledge but a contest over who gets to define what counts as legitimate science. Such rhetoric exemplifies what [10] warned of, that is, a tendency to view scientific institutions not as neutral arbiters of truth but as politically compromised actors.
The “comparative framing” topic is particularly revealing, illustrating how users employed temporal and geographic comparisons to justify their hesitancy. Through these contrasts, they crafted a rhetoric of contradiction—arguing that if high vaccination rates coincide with persistent infections, then the promises of vaccine efficacy must have failed. This form of rhetorical empiricism, disguised as statistical reasoning, relies on surface-level data rather than expert interpretation. Similar practices were uncovered in Rosenberg et al.’s [63] study, which showed the use of a rhetoric of data-driven decision-making to advocate in support of unorthodox conclusions.
The temporal analysis also produced valuable insight regarding the evolution of these narratives over time. More specifically, the online conversations served as both reflections and amplifiers of societal anxieties surrounding the vaccine debate over time. First, the period of highest activity aligns with mandated enforcement of vaccines in workplaces, schools, and federal agencies in late 2021. In fact, the “skepticism of the vaccine efficacy” topic reached its apex during this phase. It is not surprising that major concerns by vaccine efficacy skeptics emerged as a major topical theme given the amplification of vaccine misinformation distributed online over time [64]. The “distrust” topic also rose moderately over time, fueled by misinformation surrounding pharmaceutical profits and the politicization of health guidance. Public conversations often revolved around narratives of betrayal (e.g., “they said it was one shot”) in such crises [65]. Finally, the “scientific justification” topic gained traction over time as skeptics increasingly adopted the language of science, citing selective studies or misinterpreting official data to construct legitimacy around hesitancy.
Interestingly, the volume of discourse appearing across all topics declined after the Omicron surge in early 2022. This decline likely reflects pandemic fatigue, which is well-documented in the health risk and crisis communication research [66] and shifting media attention toward post-pandemic crisis recovery, as well as other issues such as inflation and geopolitical tensions (e.g., the war in Ukraine).

Limitation and Future Recommendation

Several limitations of this study merit discussion. The analysis relied solely on geotagged posts from X (formerly Twitter) originating in South Carolina during 2021–2022, which represent a subset of global social media content. While this aligns with the research scope, it does not represent conversations about the COVID-19 vaccine on other platforms such as Facebook, Reddit, Thread, etc. This may also underrepresent certain demographics and overrepresent others, limiting generalizability beyond the subset of users who share location data because not many users refrain from disclosing location on their X accounts. This study focused exclusively on text-based tweets, ignoring images, videos, and memes that often carry strong rhetorical weight in vaccine discourse. This omission limits the ability to capture visual narratives and multimodal framing strategies prevalent on social platforms. Also, X was the sole data source for this study due to geolocation availability. However, vaccine hesitancy discourse varies across platforms (e.g., TikTok, Facebook, Reddit), so findings may not generalize to other ecosystems with different affordances and audience demographics.
The transferability of these findings to other geographic contexts should be considered limited but informative. Although several of the narrative patterns identified in this study—such as skepticism of vaccine efficacy, distrust in institutions, and the use of comparative or experiential reasoning—have been documented across broader national and international research on vaccine hesitancy [14,18], their specific expression is shaped by local sociopolitical environments and demographic characteristics. Because the dataset was restricted to posts geolocated to South Carolina, the themes observed here may reflect contextual factors unique to that setting, including state-level political discourse, regional health-communication practices, and community trust structures. Thus, while the overarching rhetorical patterns identified in this study may appear in other regions, variation in tone, intensity, and prominence should be expected when applied to different geographic, cultural, or policy contexts.
This study used a few-shot approach for the model’s learning, which depends heavily on prompt design quality and example selection. Though the iterative refinement process improved consistency, the method remains vulnerable to prompt sensitivity. Because LLMs interpret prompts holistically, even minor textual modifications, such as reordering examples, changing label descriptions, or altering connective phrases can yield alternative outputs. This introduces reproducibility concerns, especially since prompts were manually curated. We attempted to mitigate this by using deterministic decoding, balanced examples, and fixed prompt templates across all models; however, prompt sensitivity cannot be fully removed under current LLM architectures. Also, only three instruction-tuned LLMs (Mistral, Llama, and DeepSeek) were tested due to availability issues, and zero-shot classification relied on a single DeBERTa model. This narrow selection limits conclusions about generalizability across architectures.
Future research can build on this methodological framework introduced in this study, which combines few-shot and zero-shot learning with structured decision rules to address annotation burden and improve interpretability. While this approach demonstrated strong performance, the zero-shot classification pipeline did not achieve comparable accuracy, suggesting an opportunity for improvement. Enhancing zero-shot methods could make stance detection even more efficient by eliminating the need for curated examples, which would be particularly valuable for real-time monitoring during health crises [24,36]. Additionally, future work should explore multimodal extensions by incorporating images, videos, and memes, as visual rhetoric often amplifies misinformation [64]. Comparative studies across platforms (e.g., X, TikTok, Reddit) could reveal how platform-specific affordances shape discourse, improving the robustness of stance detection models [19].

5. Conclusions

The present study carves a deeper landscape in vaccine hesitancy research in many ways. Specifically, this study reveals that vaccine hesitancy is not simply rooted in misinformation but is a dynamic, context-driven discourse shaped by trust, identity, and competing claims to scientific knowledge. Through the integration of few-shot learning and topic modeling, this study uncovers how users construct vaccine skepticism, weaving narratives that range from personal experiences to selective interpretations of scientific evidence as a means of legitimizing their position. These findings offer a critical insight for health communication: engaging hesitancy requires more than fact-checking; it demands rebuilding trust and engaging with the narratives that people form to make sense of risk. As digital platforms continue to amplify polarized rhetoric, scholars and practitioners must move beyond reactive responses toward proactive strategies that address the deeper crisis of trust in health institutions. Ultimately, the challenge is not to persuade vaccine skeptics but to rebuild the bridge between scientific evidence and public conversation because when trust erodes, facts fail to persuade.

Author Contributions

Conceptualization, M.E.K., S.H.T. and D.D.S.; Methodology, M.E.K., S.H.T. and L.R.; Validation, M.E.K., S.H.T. and G.L.P.L.; Formal Analysis, M.E.K. and S.H.T.; Data Curation, M.E.K. and S.H.T.; Writing—Original Draft Preparation, M.E.K., S.H.T., D.D.S. and G.L.P.L.; Writing—Review and Editing, D.D.S., M.E.K. and L.R.; Visualization, M.E.K.; Supervision. L.R. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The original data presented in the study are openly available in Open Science Framework: https://osf.io/8faet/overview?view_only=ea7547c998c74e598ac26082220ae325 (accessed on 29 November 2025).

Conflicts of Interest

The authors declare no conflicts of interest.

Appendix A

“““
You are a helpful vaccine stance classifier. Classify each tweet as:
- Pro-vaccine
- Vaccine hesitant
- Irrelevant
Decision rules (apply in order):
Don’t keyword-match blindly: many posts include anti-vax terms due to how this dataset was collected.
“Unvaccinated” rule: mere mentions are not Hesitant.
- Blaming/criticizing the unvaccinated or encouraging vaccination → Pro-Vaccine.
- Personal doubt/uncertainty about getting vaccinated → Vaccine Hesitant.
Tweet: “I believe in vaccines. They save lives. I got all my doses as soon as I could.”
Label: Pro-vaccine
Tweet: “I’m not sure about this vaccine. It came out too fast.”
Label: Vaccine hesitant
Tweet: “Breaking: CDC updates travel guidelines for international travelers.”
Label: Irrelevant
Tweet: “I’m unvaccinated and still healthy. No way I’m taking that shot.”
Label: Vaccine hesitant
Tweet: “Get vaccinated”
Label: Pro-vaccine
Tweet: “Study: Vaccinated individuals have a lower risk of hospitalization.”
Label: Pro-vaccine
Tweet: “France announces new rules for international travel amid COVID surge.”
Label: Irrelevant
”””

References

  1. Increases in Vaccine-Preventable Disease Outbreaks Threaten Years of Progress, Warn WHO, UNICEF, Gavi. Available online: https://www.who.int/news/item/24-04-2025-increases-in-vaccine-preventable-disease-outbreaks-threaten-years-of-progress--warn-who--unicef--gavi (accessed on 24 April 2025).
  2. Wilson, S.L.; Wiysonge, C. Social Media and Vaccine Hesitancy. BMJ Glob. Health 2020, 5, e004206. [Google Scholar] [CrossRef] [Scilit]
  3. Robertson, R. Glocalization: Time–Space and Homogeneity–Heterogeneity. In Global Modernities; Featherstone, M., Lash, S., Robertson, R., Eds.; Sage: London, UK, 1995; pp. 25–54. [Google Scholar]
  4. Roudometof, V.; Dessì, U. Culture and Glocalization: An Introduction. In Handbook of Culture and Glocalization; Roudometof, V., Dessi, U., Eds.; Edward Elgar: Cheltenham, UK, 2022; pp. 1–26. [Google Scholar]
  5. Rossi, M.M.; Parisi, M.A.; Cartmell, K.B.; McFall, D. Understanding COVID-19 Vaccine Hesitancy in the Hispanic Adult Population of South Carolina: A Complex Mixed-Method Design Evaluation Study. BMC Public Health 2023, 23, 2359. [Google Scholar] [CrossRef] [Scilit]
  6. Rancher, C.; Moreland, A.D.; Smith, D.W.; Cornelison, V.; Schmidt, M.G.; Boyle, J.; Dayton, J.; Kilpatrick, D.G. Using the 5C Model to Understand COVID-19 Vaccine Hesitancy Across a National and South Carolina Sample. J. Psychiatr. Res. 2023, 160, 180–186. [Google Scholar] [CrossRef] [Scilit]
  7. Kanyangarara, M.; Vora, S.; Seck, F.; Dhankhode, N.; Mundagowa, P.T. COVID-19 Vaccination Coverage and Associated Factors Among Underserved Communities in South Carolina: Results from a Cross-Sectional Study. J. Racial Ethn. Health Disparities, 2025; in press. [CrossRef] [Scilit]
  8. Richman, A.R.; Schwartz, A.J.; Maness, S.B.; Sanchez, L.; Torres, E. Exploring Vaccine Hesitancy, Structural Barriers, and Trust in Vaccine Information Among Populations Living in the Rural Southern United States. Vaccines 2025, 13, 699. [Google Scholar] [CrossRef] [Scilit]
  9. Hornsey, M.J.; Harris, E.A.; Fielding, K.S. The Psychological Roots of Anti-Vaccination Attitudes: A 24-Nation Investigation. Health Psychol. 2023, 37, 307–315. [Google Scholar] [CrossRef] [Scilit]
  10. Goldenberg, M.J. Vaccines, Beliefs and Autarchy: The History and Philosophy of Vaccine Skepticism; Routledge: New York, NY, USA, 2021. [Google Scholar]
  11. South Carolina’s Measles Outbreak Illustrates Danger of Vaccine Misinformation. Available online: https://truthout.org/articles/south-carolinas-measles-outbreak-illustrates-danger-of-vaccine-misinformation/ (accessed on 26 November 2025).
  12. South Carolina’s Vaccination Rate Too Low as Measles Surges. Available online: https://www.thestate.com/opinion/article307032491.html (accessed on 27 May 2025).
  13. Zhang, H.; Zhang, X.; Huang, H.; Yu, L. Prompt-Based Meta-Learning for Few-Shot Text Classification. In Proceedings of EMNLP 2022; Association for Computational Linguistics: Stroudsburg, PA, USA, 2022; pp. 1342–1357. [Google Scholar] [CrossRef] [Scilit]
  14. Sasse, K.; Mahabir, R.; Gkountouna, O.; Crooks, A.; Croitoru, A. Understanding the Determinants of Vaccine Hesitancy in the United States: A Comparison of Social Surveys and Social Media. PLoS ONE 2024, 19, e0301488. [Google Scholar] [CrossRef] [Scilit]
  15. Li, L.; Wood, C.E.; Kostkova, P. Vaccine Hesitancy and Behavior Change Theory-Based Social Media Interventions: A Systematic Review. Transl. Behav. Med. 2022, 12, 243–272. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  16. Limaye, R.J.; Holroyd, T.A.; Blunt, M.; Jamison, A.F.; Sauer, M.; Weeks, R.; Wahl, B.; Christenson, K.; Smith, C.; Minchin, J.; et al. Social Media Strategies to Affect Vaccine Acceptance: A Systematic Literature Review. Expert Rev. Vaccines 2021, 20, 959–973. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  17. Allen, J.; Watts, D.J.; Rand, D.G. Quantifying the Impact of Misinformation and Vaccine-Skeptical Content on Facebook. Science 2024, 384, 978–984. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  18. Taubert, F.; Meyer-Hoeven, G.; Schmid, P.; Gerdes, P.; Betsch, C. Conspiracy Narratives and Vaccine Hesitancy: A Scoping Review. BMC Public Health 2024, 24, 3325. [Google Scholar] [CrossRef] [Scilit]
  19. Neff, T.; Kaiser, J.; Wasserman, H.; Ghezzi, A.; Singh, A.; Scacco, A. Vaccine Hesitancy in Online Spaces: A Scoping Review of the Research Literature, 2000–2020. HKS Misinf. Rev. 2021, 2, 1–18. [Google Scholar] [CrossRef] [Scilit]
  20. Glandt, K.; Khanal, S.; Li, Y.; Caragea, D.; Caragea, C. Stance Detection in COVID-19 Tweets. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing; (Volume 1: Long Papers); Association for Computational Linguistics: Stroudsburg, PA, USA, 2021. [Google Scholar] [CrossRef] [Scilit]
  21. Balaji, S.; Joshi, S.; Hari, A.; Singam, A.; Varma, V. COVID-19 Vaccine Stance Classification from Tweets. CEUR Workshop Proc. 2022, 3180, 1–5. [Google Scholar]
  22. Wang, Z.; Lin, Y.; Shen, J.; Zhu, X. A Survey of Large Language Models for Text Classification: What, Why, When, Where, and How. TechRxiv 2025, preprint. [Google Scholar] [CrossRef] [Scilit]
  23. Barberia, L.G.; Lombard, B.; Roman, N.T.; Sousa, T.C.M. Clarifying Misconceptions in COVID-19 Vaccine Sentiment and Stance Analysis. arXiv 2025, arXiv:2503.18095. [Google Scholar] [CrossRef] [Scilit]
  24. Brown, T.B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. Language Models Are Few-Shot Learners. Adv. Neural Inf. Process. Syst. 2020, 33, 1877–1901. [Google Scholar]
  25. Gera, P.; Neal, T. MLSD: A Novel Few-Shot Learning Approach to Enhance Cross-Target and Cross-Domain Stance Detection. arXiv 2025, arXiv:2509.03725. [Google Scholar] [CrossRef] [Scilit]
  26. Khiabani, P.J.; Zubiaga, A. Few-Shot Learning for Cross-Target Stance Detection by Aggregating Multimodal Embeddings. IEEE Trans. Comput. Soc. Syst. 2024, 11, 2081–2090. [Google Scholar] [CrossRef] [Scilit]
  27. Liu, R.; Lin, Z.; Ji, H.; Li, J.; Fu, P.; Wang, W. Target-Aware Contrastive Learning and Consistency Regularization for Few-Shot Stance Detection. In Proceedings of COLING 2022; International Committee on Computational Linguistics: Gyeongju, Republic of Korea, 2022; pp. 1234–1248. [Google Scholar]
  28. Kostina, A.; Dikaiakos, M.D.; Stefanidis, D.; Pallis, G. Large Language Models for Text Classification: Case Study and Comprehensive Review. arXiv 2025, arXiv:2501.08457. [Google Scholar] [CrossRef] [Scilit]
  29. Santos, D.P.D.; Mendonca, F.L.; Lustosa, J.P.; Serrano, A.; Torres, J.A.S.; da Silva, D.A. Large Language Models for Text Classification: A New Era of Accuracy and Efficiency. In CSCI Proceedings; Springer: Berlin/Heidelberg, Germany, 2025; pp. 1–8. [Google Scholar]
  30. Trust, P.; Minghim, R. A Study on Text Classification in the Age of Large Language Models. Mach. Learn. Knowl. Extr. 2024, 6, 2688–2721. [Google Scholar] [CrossRef] [Scilit]
  31. Krishnan, G.S.; Kamath, S.S.; Sugumaran, V. Predicting Vaccine Hesitancy and Vaccine Sentiment Using Topic Modeling and Evolutionary Optimization. In NLDB 2021 Proceedings; Springer: Berlin/Heidelberg, Germany, 2021. [Google Scholar]
  32. Melton, C.A.; Olusanya, O.A.; Ammar, N.; Shaban-Nejad, A. Public Sentiment Analysis and Topic Modeling Regarding COVID-19 Vaccines on Reddit. J. Infect. Public Health 2021, 14, 1505–1512. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  33. Aivazpour, Z.; Zhan, Y. Exploring the Semantic and Linguistic Features of Vaccine Skepticism in Online Discourse. In AMCIS 2025 Proceedings; Association for Information Systems: Atlanta, GA, USA, 2025; Paper 20; Available online: https://aisel.aisnet.org/amcis2025/sig_odis/sig_odis/20 (accessed on 29 November 2025).
  34. Kabir, M.E. Topic and Sentiment Analysis of Responses to Muslim Clerics’ Misinformation Correction About COVID-19 Vaccine: Comparison of Three Machine Learning Models. Online Media Glob. Commun. 2022, 1, 497–523. [Google Scholar] [CrossRef] [Scilit]
  35. Sievert, C.; Shirley, K. LDAvis: A Method for Visualizing and Interpreting Topics. In Proceedings of the Workshop on Interactive Language Learning, Visualization, and Interfaces; Association for Computational Linguistics: Stroudsburg, PA, USA, 2014. [Google Scholar]
  36. Doshi-Velez, F.; Kim, B. Towards a Rigorous Science of Interpretable Machine Learning. arXiv 2017, arXiv:1702.08608. [Google Scholar] [CrossRef] [Scilit]
  37. Herrewijnen, E.; Nguyen, D.; Bex, F.; van Deemter, K. Human-Annotated Rationales and Explainable Text Classification: A Survey. Front. Artif. Intell. 2024, 7, 1260952. [Google Scholar] [CrossRef] [Scilit]
  38. Blei, D.M.; Ng, A.Y.; Jordan, M.I. Latent Dirichlet Allocation. J. Mach. Learn. Res. 2003, 3, 993–1022. [Google Scholar]
  39. Bail, C.A. Breaking the Social Media Prism: How to Make Our Platforms Less Polarizing; Princeton University Press: Princeton, NJ, USA, 2021. [Google Scholar]
  40. Gao, T.; Fisch, A.; Chen, D. Making Pre-Trained Language Models Better Few-Shot Learners. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics; Association for Computational Linguistics: Stroudsburg, PA, USA, 2021; pp. 381–398. [Google Scholar]
  41. Kabir, M.E. #StopAsianHate Counterspeech on Twitter: Effectiveness of Counterspeech Strategies and Geospatial Analysis; Bowling Green State University: Bowling Green, OH, USA, 2023. [Google Scholar]
  42. Kabir, M.E.; Ha, L. Influencers Against Hate: A Comparison of Counter Speech Strategies Among South Asian, East Asian, and Non-Asian American Social Media Influencers. Howard J. Commun. 2025, 37, 543–561. [Google Scholar] [CrossRef] [Scilit]
  43. Yin, W.; Hay, J.; Roth, D. Benchmarking Zero-Shot Text Classification: Datasets, Evaluation and Entailment Approach. In Proceedings of EMNLP-IJCNLP 2019; Association for Computational Linguistics: Stroudsburg, PA, USA, 2019; pp. 3914–3923. [Google Scholar] [CrossRef] [Scilit]
  44. Laurer, M.; van Atteveldt, W.; Casas, A.; Welbers, K. Less Annotating, More Classifying: Addressing the Data Scarcity Issue of Supervised Machine Learning with Deep Transfer Learning and BERT-NLI. Polit. Anal. 2024, 32, 84–100. [Google Scholar] [CrossRef] [Scilit]
  45. Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; et al. Transformers: State-of-the-Art Natural Language Processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations; Association for Computational Linguistics: Stroudsburg, PA, USA, 2020; pp. 38–45. [Google Scholar] [CrossRef] [Scilit]
  46. Jiang, A.Q.; Sablayrolles, A.; Deleu, T.; Lachaux, M.A.; Mensch, A.; Bressand, M.; Bamford, C.; Chaplot, D.S.; de las Casas, D.; Lambert, N.; et al. Mistral 7B. arXiv 2023, arXiv:2310.06825. [Google Scholar] [CrossRef] [Scilit]
  47. Grattafiori, A.; Dubey, A.; Jauhri, A.; Pandey, A.; Kadian, A.; Al-Dahle, A.; Letman, A.; Mathur, A.; Schelten, A.; Vaughan, A.; et al. The Llama 3 Herd of Models. arXiv 2024, arXiv:2407.21783. [Google Scholar]
  48. Bi, X.; Chen, D.; Chen, G.; Chen, S.; Dai, D.; Deng, C.; Ding, H.; Dong, K.; Du, Q.; Fu, Z.; et al. DeepSeek LLM: Scaling Open-Source Language Models with Longtermism. arXiv 2024, arXiv:2401.02954. [Google Scholar] [CrossRef] [Scilit]
  49. Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Ichter, B.; Xia, F.; Chi, E.; Le, Q.V.; Zhou, D. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. Adv. Neural Inf. Process. Syst. 2022, 35, 24824–24837. [Google Scholar]
  50. Min, S.; Lewis, M.; Zettlemoyer, L.; Hajishirzi, H. MetaICL: Learning to Learn in Context. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics; Association for Computational Linguistics: Stroudsburg, PA, USA, 2022; pp. 279–285. [Google Scholar]
  51. Kumar, S.; Goud, J.S.; Choudhury, M. Detecting Stance in Social Media: State-of-the-Art and Trends. Comput. Linguist. 2021, 47, 835–886. [Google Scholar]
  52. Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. Llama 2: Open Foundation and Fine-Tuned Chat Models. arXiv 2023, arXiv:2307.09288. [Google Scholar] [CrossRef] [Scilit]
  53. Sokolova, M.; Lapalme, G. A Systematic Analysis of Performance Measures for Classification Tasks. Inf. Process. Manag. 2009, 45, 427–437. [Google Scholar] [CrossRef] [Scilit]
  54. Gao, X.; Lin, Y.; Li, R.; Wang, Y.; Chu, X.; Ma, X.; Yu, H. Enhancing Topic Interpretability for Neural Topic Modeling Through Topic-Wise Contrastive Learning. arXiv 2024, arXiv:2412.17338. [Google Scholar] [CrossRef] [Scilit]
  55. Röder, M.; Both, A.; Hinneburg, A. Exploring the Space of Topic Coherence Measures. In Proceedings of the Eighth ACM International Conference on Web Search and Data Mining; Association for Computing Machinery: New York, NY, USA, 2015; pp. 399–408. [Google Scholar]
  56. Bird, S.; Klein, E.; Loper, E. Natural Language Processing with Python; O’Reilly Media: Sebastopol, CA, USA, 2009. [Google Scholar]
  57. Řehůřek, R.; Sojka, P. Software Framework for Topic Modelling with Large Corpora. In Proceedings of the LREC 2010 Workshop on New Challenges for NLP Frameworks; University of Malta: Msida, Malta, 2010; pp. 45–50. [Google Scholar]
  58. DiRusso, C.; Kabir, M.E.; Hagan, A.H.; Boatwright, B. Why Is the Southeastern United States Late to Adopt Electric Vehicles? Analyzing Trends in Social Media Conversation to Form Marketing Recommendations. Transp. Res. Interdiscip. Perspect. 2025, 32, 101509. [Google Scholar] [CrossRef] [Scilit]
  59. Diaz, P.; Zizzo, J.; Balaji, N.C.; Reddy, R.; Khodamoradi, K.; Ory, J.; Ramasamy, R. Fear About Adverse Effect on Fertility Is a Major Cause of COVID-19 Vaccine Hesitancy in the United States. Andrologia 2022, 54, e14361. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  60. Gauna, F.; Raude, J.; Khouri, C.; Cracowski, J.-L.; Ward, J.K. Exploring the Relationship Between Experience of Vaccine Adverse Events and Vaccine Hesitancy: A Scoping Review. Hum. Vaccines Immunother. 2025, 21, 2471225. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  61. Kuru, O.; Stecula, D.; Lu, H.; Ophir, Y.; Chan, M.-P.S.; Winneg, K.; Jamieson, K.H.; Albarracín, D. The Effects of Scientific Messages and Narratives About Vaccination. PLoS ONE 2021, 16, e0248328. [Google Scholar] [CrossRef] [Scilit]
  62. Corker, A. Vaccine Hesitancy Has Become a Nationwide Issue: What Can Science Do About It? MUSC Research. 10 April 2023. Available online: https://research.musc.edu/stories/news/2023/04/10/vaccine-hesitancy (accessed on 29 November 2025).
  63. Rosenberg, J.; Syed, S.; Reinecke, K.; DiSalvo, C. Viral Visualizations: How Coronavirus Skeptics Use Orthodox Data Practices to Promote Unorthodox Science Online. arXiv 2021, arXiv:2101.07993. [Google Scholar]
  64. Zimmerman, T.; Shiroma, K.; Fleischmann, K.R.; Xie, B.; Jia, C.; Verma, N.; Lee, M.K. Misinformation and COVID-19 Vaccine Hesitancy. Vaccine 2022, 41, 136–144. [Google Scholar] [CrossRef] [Scilit]
  65. Reid, J.C.; Brown, S.J.; Dmello, J. COVID-19, Diffuse Anxiety, and Public (Mis)Trust in Government: Empirical Insights and Implications for Crime and Justice. Crim. Justice Rev. 2023, 49, 117–134. [Google Scholar] [CrossRef] [Scilit]
  66. Okuhara, T.; Terada, M.; Okada, H.; Yokota, R.; Kiuchi, T. Experiences of Public Health Professionals Regarding Crisis Communication During the COVID-19 Pandemic: Systematic Review of Qualitative Studies. JMIR Infodemiology 2025, 5, e66524. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Timeline of posts showing peak around mid-2021 to mid-2022.
Figure 1. Timeline of posts showing peak around mid-2021 to mid-2022.
Bdcc 10 00159 g001
Figure 2. Coherence scores.
Figure 2. Coherence scores.
Bdcc 10 00159 g002
Figure 3. Volume of Vaccine Hesitancy Topics.
Figure 3. Volume of Vaccine Hesitancy Topics.
Bdcc 10 00159 g003
Figure 4. Word Clouds for Vaccine Hesitancy Topics.
Figure 4. Word Clouds for Vaccine Hesitancy Topics.
Bdcc 10 00159 g004
Figure 5. Evolution of the Topics Over Time.
Figure 5. Evolution of the Topics Over Time.
Bdcc 10 00159 g005
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Kabir, M.E.; Tanim, S.H.; Sellnow, D.D.; Luteria, G.L.P.; Rennert, L. Teaching AI to Decode Vaccine Hesitancy Narratives: A Few-Shot Learning and Topic Modeling Approach. Big Data Cogn. Comput. 2026, 10, 159. https://doi.org/10.3390/bdcc10050159

AMA Style

Kabir ME, Tanim SH, Sellnow DD, Luteria GLP, Rennert L. Teaching AI to Decode Vaccine Hesitancy Narratives: A Few-Shot Learning and Topic Modeling Approach. Big Data and Cognitive Computing. 2026; 10(5):159. https://doi.org/10.3390/bdcc10050159

Chicago/Turabian Style

Kabir, Md Enamul, Shakhawat H. Tanim, Deanna D. Sellnow, Geneva Lei P. Luteria, and Lior Rennert. 2026. "Teaching AI to Decode Vaccine Hesitancy Narratives: A Few-Shot Learning and Topic Modeling Approach" Big Data and Cognitive Computing 10, no. 5: 159. https://doi.org/10.3390/bdcc10050159

APA Style

Kabir, M. E., Tanim, S. H., Sellnow, D. D., Luteria, G. L. P., & Rennert, L. (2026). Teaching AI to Decode Vaccine Hesitancy Narratives: A Few-Shot Learning and Topic Modeling Approach. Big Data and Cognitive Computing, 10(5), 159. https://doi.org/10.3390/bdcc10050159

Article Metrics

Back to TopTop