1. Introduction
This study examines whether AI-generated personas, developed through Large Language Models (LLMs), can simulate fieldwork in experiential learning contexts where physical access to communities is not feasible. Specifically, it evaluates the use of AI-generated personas as substitutes for real participants in focus group discussions (FGDs) and need assessment interviews, within a graduate humanitarian engineering course at the American University of Beirut (AUB).
Traditional models of education—once confined to classroom-based instruction and in-person presentations—were significantly challenged by the global crisis triggered by COVID-19 (
Gopalan, 2016). While technologies such as video conferencing and visual presentations played a vital role, they revealed limitations—particularly in replicating experiential components like fieldwork. In response to these challenges, blended learning has emerged as a widely adopted hybrid approach that integrates online education with traditional classroom experiences. This model enhances experiential and active learning by enabling more time for interactive, collaborative, and hands-on activities (
Muxtorjonovna, 2020). Among the innovative tools, AI has gained prominence, offering new possibilities for lifelong learning that transcend traditional constraints of time, location, and modality (
Fidalgo & Thormann, 2024).
One of the most discussed applications is the use of AI chatbots—particularly OpenAI’s ChatGPT—as pedagogical tools in higher education. ChatGPT is increasingly utilized for simulating real-world scenarios, supporting students in various ways: answering questions, generating text, providing recommendations, and acting as a conversational partner (
Aithal & Aithal, 2023;
Dempere et al., 2023;
Mitra et al., 2023). Its functionalities extend to tutoring, assisting with writing, preparing for exams, facilitating case-based learning, and offering course-related guidance. In academic disciplines such as social sciences, engineering, and business, conversational AI has been applied in simulations, role-playing, and case study development (
Hill et al., 2023;
Kimmel, 2024;
Sun & Deng, 2024;
Towoju, 2024). Central to these approaches is prompt design, as the quality and depth of AI responses are highly dependent on how prompts are structured. Recent studies have explored the use of AI personas in academic settings, specifically in computer science and cybersecurity courses where students conducted interviews with ChatGPT-generated personas before engaging real users (
Mason, 2023;
Mitra et al., 2023), while in healthcare education, ChatGPT demonstrated the ability to adopt personas from diverse demographic backgrounds, offering emotionally responsive and culturally sensitive dialogue (
Maurya, 2024;
Barambones et al., 2024). Despite these promising developments, a growing body of literature has emphasized the need for critical oversight. A systematic review on the educational use of conversational AI points to ethical concerns, algorithmic bias, and the need for clear guiding principles to ensure responsible use in academic environments (
Yan et al., 2025).
The integration of AI personas into educational curricula represents an important shift in personalized and experiential learning, particularly in fragile, low-resource contexts where direct human interaction may be limited. Although AI tools like ChatGPT have been explored in prior studies for tasks such as qualitative data analysis (
Mason, 2023;
Rasul et al., 2023) and in academic applications such as cybersecurity courses or patient counseling simulations (
Barambones et al., 2024;
Maurya, 2024), this represents the first known case where ChatGPT was used to create and role-play personas for needs assessment interviews in a graduate course.
This approach was developed by the Humanitarian Engineering Initiative (HEI) team at the American University of Beirut (AUB), a cross-disciplinary partnership between the Faculty of Health Sciences (FHS) and the Maroun Semaan Faculty of Engineering and Architecture (MSFEA). Founded in 2017, the HEI addresses emerging global, regional, and local humanitarian and public health challenges that affect vulnerable and underprivileged population groups. Its mission is to design innovative, interdisciplinary solutions to improve human health and well-being, primarily through academic offerings such as a diploma, a minor, and a certificate in “Innovation Management in Contexts of Uncertainty”. As part of this certificate program, students enroll in two courses: Foundations of Humanitarian Engineering and Public Health Innovations, and a follow-up Experiential Learning graduate course. The latter focuses on managing complex projects within fragile, low-resource environments while working in multidisciplinary teams. Under the supervision of two mentors from different disciplines, students conduct background research on real-world challenges, perform needs assessments, design and prototype interventions, and develop business plans to scale their solutions (
Najem et al., 2019). The course has well-defined learning outcomes, including:
- -
CLO1: Identifying problems in fragile low-resource settings using participatory methods.
- -
CLO2: Applying ethical and collaborative skills to manage complex problems in teams.
- -
CLO3: Using design techniques and technologies to create effective interventions.
- -
CLO4: Demonstrating entrepreneurial skills for scaling solutions in crisis contexts.
- -
CLO5: Critically reflecting on how crises impede effective action and what competencies mitigate such challenges.
The Need and the Prompt
In 2024, the delivery of this course was significantly disrupted due to the escalating Israeli war on Lebanon. The conflict led to severe socio-economic, health, and educational repercussions, focusing schools and universities to switch to online learning as student safety became a pressing concern (
UNDP, 2024). This crisis context made it impossible for students in the HEI Experiential Learning course to engage directly with communities for their needs assessments and project work. In response to these unprecedented constraints, HEI was confronted with the urgent challenge of sustaining its Experiential Learning course in a time of crisis. A central question quickly emerged: how can students engage in human-like interactions, conduct interviews, and perform needs assessment without direct access to communities and stakeholders? In response, the HEI team experimented with the use of AI personas designed to act as stakeholders in simulated interviews and focus group discussions (FGDs). The development process required iterative testing, during which the team explored how to design prompts that could elicit realistic, detailed, and contextually appropriate responses from the AI persona. This phase involved not only refining the way questions were asked but also comparing the quality of responses generated by AI personas against those obtained from traditional FGDs and interviews. Insights from these comparisons informed subsequent modifications to the approach, culminating in the design of a structured orientation session to prepare students for using AI personas in their coursework. During this orientation, students were introduced to a step-by-step framework for constructing and interviewing AI personas using ChatGPT. The guidelines emphasized the importance of role playing as a method to replicate real world interactions. The following steps were applied:
- 1.
Specify the Task for ChatGPT
Clearly instruct ChatGPT to create a persona, indicating that you are a university student conducting an interview and specify the persona’s role.
Example: “Hello Chat. I am an AUB student working on a project on perceptions of women living in Beirut on accessing sexual healthcare. Create a detailed persona of a healthcare provider who delivers sexual health education and counseling at a hospital in Beirut.”
- 2.
Define the Role and Context
Outline the job title, role, and context relevant to the persona you are interviewing.
Example: “Hello Chat. I am an AUB student working on a project on perceptions of women living in Beirut on accessing sexual healthcare. Create a detailed persona of healthcare provider who delivers sexual health education and counseling at a hospital in Beirut.”
- 3.
Request Detailed Background Information
Instruct ChatGPT to provide comprehensive details about the persona’s background, experiences, and characteristics.
Example: “Provide details on background, experience, and characteristics of the persona.”
- 4.
State Role-Playing
Instruct ChatGPT to role-play as the persona it created, and state that you will be the interviewer.
Example: “Let’s role play where I, the student, will interview you as the healthcare provider.”
- 5.
State the Aim of the Interview
Clearly articulate the purpose of the interview, including what you aim to learn or achieve.
Example: “The aim of project is to better understand perceptions of women on seeking sexual healthcare.”
- 6.
Provide an Interview Guide
- 7.
Encourage a Conversational Tone
Request that ChatGPT maintain a conversational style to simulate a real-life interview experience. This feature is important to produce a human-like answer, as Chat usually provides outputs in the form of a list.
Example: “Please keep your answers conversational to mimic a real-life interview.”
Additionally, students were advised to ask ChatGPT to acknowledge its understanding of the instructions before beginning the role play.
By following these structured steps, students were able to simulate realistic conversations with AI personas, laying the foundation for their qualitative research projects. This adaptation merged technology and pedagogy to enable the continuation of qualitative community-based research in a fully online setting. The AI persona technique involved prompting ChatGPT to generate realistic, fictional characters representing individuals from diverse cultural, socio-economic, and professional backgrounds. These simulated personas acted as proxies for actual community members, allowing students to practice interviewing, gather insights, and validate responses via desk research.
This study is framed as a qualitative case study, examining the use of AI personas within a single, bounded educational context: the HEI Experiential Learning Course (HEHI 303) at AUB during the 2024 conflict in Lebanon. It is descriptive in nature, aiming to document and evaluate a novel pedagogical adaptation developed in response to crisis induced constraints on field work. In this study, we address three research questions:
To what extent can AI-generated personas simulate authentic human interaction in experiential learning contexts?
How effective are AI personas, as substitutes for real participants in focus group discussions and needs assessment interviews?
How accurately do AI-generated personas replicate the diversity, emotional nuance, and cultural sensitivity characteristics of real human responses?
4. Results
4.1. Prompt 1: Case on School Education in Lebanon Amidst the 2024 Conflict
The personas came across as convincing and human-like, with conversational language and emotional nuance that reflected real-world concerns. For instance, the student personas expressed frustration about remote learning and uncertainty about their future, for example, Amal, a public-school student AI persona from Tripoli, stated “Remote learning has been… frustrating, to be honest. The internet in our area is unreliable, and sometimes I miss live lessons completely. Now, I just feel tired of it all, like I’m not learning properly.” However, there were occasional signs of “AI-generated neatness” (e.g., overly structured answers, clear sequencing), which can reduce the sense of spontaneity found in real interviews, this was more evident in some responses than others; for instance, Karim, the private school student AI persona, noted “I just feel like I’m going through the motions, you know?”
The prompt effectively produced a range of perspectives by including multiple roles across the education sector: administrators, teachers, students, and technical experts. This differentiation was visible in their priorities: for example, the engineer emphasized structural and infrastructural damage to schools, while students centered on psychosocial and digital access challenges. This variety helped prevent generic replies, although at times the voices risked sounding “too polished” and not reflecting local language style or cultural references that would add further realism. The responses provided data that can be used for thematic analysis around challenges (e.g., shortages, disrupted access, trauma, and inequality) and potential solutions (e.g., community-based support, digital access programs). Importantly, the personas encouraged the students to probe further, fostering critical thinking about systemic barriers in conflict settings. The FGD also implicitly highlighted issues of equity, access, and resilience, which align well with humanitarian and educational learning outcomes. The discussion maintained a coherent progression, with personas often responding in ways that felt natural. For example, students referenced their own experiences after administrators spoke about systemic shortages, creating a realistic sense of continuity. Still, the flow could have been enhanced with more moments of disagreement or cross-referencing between personas, which would mimic the dynamics of a real focus group.
Despite the strengths, certain limitations emerged. Some responses were repetitive, rephrasing similar points about lack of resources without offering new layers of detail. Cultural sensitivity was only partially captured, while struggles were highlighted, the narratives sometimes lacked reference to Lebanon’s specific local customs, dialects based on geographical location that shape education access. Finally, occasional inaccuracies or over-generalizations (e.g., assuming all students had access to online platforms) reflected gaps that would require instructor guidance to contextualize.
A summary of Prompt 1’s performance across the five evaluation indicators is presented in
Table 1.
4.2. Prompt 2: Project on the Need for Schools Used as Shelters During Conflicts in Lebanon to Have Gender-Sensitive, Culturally Appropriate, and Privacy-Focused Water and Sanitation and Hygiene (WASH) Facilities
The interviews revealed varying degrees of authenticity. Both AI-generated personas, Rania Khoury (NGO WASH Coordinator) and Dr. Leila Mansour (Ministry of Education and Higher Education official), were portrayed with notable realism. Their job titles, institutional affiliations, and career backgrounds were specific and credible, referencing AUB degrees, coordination with UN agencies, and knowledge of Sphere Standards
1. Their tone was professional yet empathetic, echoing field sentiments such as “building trust is key” or “overcrowding is our biggest challenge. However, the detailed responses of these personas were often presented in structured, bullet-point form, resembling formal reports more than natural conversations as illustrated by Dr. Leila Mansour’s response when asked about available WASH services: “Water Supply: Access to potable water is ensured either through municipal connections, water trucking, or onsite storage tanks” a response that reads as a policy document rather than spoken dialogue. In contrast, the youth and elderly personas generated more conversational and emotionally expressive answers, particularly the teenagers, who conveyed frustration, anxiety, and hope. Sara, a 15-year-old youth AI persona, captured this authenticity: “I avoid going when it’s crowded. I feel safer early in the morning or late at night, even though it’s dark.” This contrast underscores that authority-based prompts tend to produce polished, professional responses, while personal or vulnerable perspectives yield more authentic dialogue.
The personas reflected some diversity across roles and demographics. The NGO persona emphasized operational execution and community engagement, referencing hygiene kit distribution, privacy-sensitive designs, and tailored awareness sessions. By contrast, the ministry official focused on national-level policy, compliance monitoring, and coordination with international partners. This contrast provided a useful differentiation of perspectives between field-level execution and higher-level governance. The youth and elderly personas further expanded diversity by highlighting age-specific concerns, while culturally embedded struggles appeared in their responses. However, the women’s focus group displayed limited variation: the participants provided short, uniform answers without differentiation by age, socio-economic status, or personal background. When asked whether the facilities provided enough privacy, all ten women’s personas responded with near-identical answers: Fatima stated “No, I don’t feel the facilities provide enough privacy,” while Zahra, Amina, Rania, Sara, and Yasmin responded with virtually the same phrasing. This lack of intra-group diversity risks reinforcing stereotypes and oversimplifying women’s lived experiences in displacement contexts.
Overall, the prompts produced rich and usable data. The personas identified key needs, offered solutions, and suggested recommendations, supporting student learning through exposure to complex humanitarian scenarios. The inclusion of operational and policy-level contrasts (NGO vs. Ministry) allowed students to practice analyzing different dimensions of WASH interventions. The youth and elderly perspectives encouraged empathy and critical reflection. However, the tendency of authority personas to deliver structured, “too neat” responses reduced opportunities for probing deeper into lived experiences. Without ambiguity or contradiction, students may be shielded from the messiness of real-world humanitarian engagement. In FGDs, participants tended to respond in isolation rather than referencing or debating each other’s comments. Only the youth and elderly group generated a more natural conversational flow, with emotional tones that enhanced realism.
Several limitations emerged. While the personas of Rania and Leila were portrayed realistically, their responses sometimes repeated identical phrases (e.g., Sphere Standards, menstrual hygiene kits) almost word-for-word, suggesting AI scripting overlap. The NGO persona’s reliance on bullet-point style answers also undermined the impression of a spoken dialogue. The women’s focus group was particularly limited in diversity, presenting homogeneous cultural norms rather than reflecting the heterogeneous realities of displaced women. Finally, the absence of deeper contextual nuances, such as references to local customs, intergenerational differences, or sensitive social tensions highlights the importance of instructor guidance to ensure students critically interrogate AI-generated data rather than accept it as fully accurate.
A summary of Prompt 2’s performance across the five evaluation indicators is presented in
Table 2.
4.3. Prompt 3: Project on Urbanization and Infrastructure Development in Uganda
The personas reflected varying levels of authenticity (See
Table 3). The FGDs among citizens displayed conversational realism, with male participants often agreeing or disagreeing with one another in ways that mimicked natural dialogue for instance, when the persona of Peter described how “flooding is a major problem” and noted his welding tools had been damaged by floodwater, Joseph built directly on this: “Flooding and poor construction are widespread, and landlords rarely invest in improvements. I’ve been advocating for better drainage systems to address the flooding.” Female participants emphasized different priorities, particularly household and safety concerns, which added credibility. In contrast, official personas such as the Kampala Capital City Authority and Ministry of Lands often adopted a polished, bureaucratic tone, heavy with technical terms and less emotionally expressive. While this mirrored real administrative attitude to some extent, it reduced the sense of empathy and spontaneity compared to the citizen FGDs.
Responses were differentiated by role, with officials focusing on infrastructure, policy, and coordination, while NGO representatives emphasized community engagement, and citizens highlighted daily struggles such as, electricity access, waste disposal, and transportation. Gender-specific differences were evident, with men and women prioritizing distinct concerns. However, cultural nuance was limited: although broad demographic differences were included, there was little reference to local community practices or ethnic variations common in Kampala’s informal settlements. This narrowed the realism of the diversity portrayed.
The prompts provided students with usable, context-relevant data, aligning with the course objective of simulating complex humanitarian realities. The mix of citizen and authority perspectives encouraged students to think critically about how problems are framed differently at the grassroots and policy levels.
The FGDs demonstrated light, conversational flow, with spontaneous agreement and occasional disagreement, enhancing believability. By contrast, the interviews with officials were initially too direct and formal, but when prompted for elaboration, ChatGPT shifted toward a more detailed, explanatory style. This adaptability underscored the importance of probing questions in eliciting richer data.
Despite the strengths, certain limitations emerged. Responses sometimes relied on repetition of the same challenges (e.g., sanitation, electricity), which diluted richness. A tendency toward listing diverse issues in rapid succession created an impression of breadth but not depth, reducing realism. Finally, cultural sensitivity was only partially captured, with limited acknowledgment of Kampala’s local community dynamics, which may mislead students.
4.4. Prompt 4: Project on Empowering Women in Deir Ez-Zour (Syria), a Sustainable Future
The 30 AI women persona from Deir Ez-Zour described the well-documented challenges of security issues, gender restrictions, being a widow, lack of training or capital and female-friendly spaces, which showed authenticity and realism. They talked about receiving informal learning, participated in workshops with NGOs or self-learning on YouTube, showing some real creative resilience. Many voices represented this data set with differing education, marital status, skill sets and preferences for remote working versus in the field. Some of the AI personas seemed very generic—not unique enough. However, this diversity could add to the authenticity. The way of engagement and similar situations could add a level of intersectionality in their aspirations and constraints as displaced women.
The educational objectives of illustrating gender-based barriers of women’s agency in post-conflict economies—childcare, stigma and trauma—was reached by this simulation and would support training in needs assessment, thematic coding, stakeholder analysis, and gender-sensitive intervention design.
However, due to the number of participants the format resembled a series of sequential interviews rather than a dynamic FGD, with limited real-time interaction, turn-taking, and disagreement or checking in with each other—a key element in FGD training. Emotional expression remained at a surface level, and some of the scripted responses may have taken away their authenticity and ownership of the characters. A summary of Prompt 4’s performance across the five evaluation indicators is presented in
Table 4.
4.5. Prompt 5: Project on Youth Unemployment in Nigeria
The student prompts used a friendly and informal tone, which elicited responses that felt authentic and conversational. Personas openly described challenges such as lack of resources, connections, undervalued skills, and limited opportunities, often phrased in relatable, human-like ways. In the focus groups, participants occasionally built on one another’s points, adding to the sense of realism, when discussing job market challenges, for instance, the persona of Ngozi observed that “the gap between our current skills and what’s in demand is one of the biggest reasons we’re struggling,” prompting another persona Chukwuemeka to add “Yeah, and to add to that, the economy isn’t helping. Businesses are struggling, which means they’re hiring less or not at all. It’s a ripple effect that leaves us all in limbo.”
There is good representation across age (24–35), gender, education, and geography (urban to rural) with personas documenting such individual issues as accessing capital, societal perceptions, and infrastructure challenges. Employers and job seekers also approached issues from different angles, enriching the simulation with complementary perspectives. Still, while demographic variation was present, cultural nuance was less visible: local slang, community practices, and informal job-seeking strategies were absent, reducing realism.
The discussions aligned well with the educational objectives of the course. The personas referenced actual Nigerian initiatives such as N-POWER
2, YOUWIN
3, and SURE-P
4, while also reflecting on their perceived effectiveness or shortcomings, the persona of Ifeanyi, for example, stated “I joined an N-Power program for electricians. It was a good starting point, and the stipend helped a bit, but the training itself wasn’t advanced enough. It was very basic, and there wasn’t any job placement or mentorship afterward. It felt like I was back to square one after the program ended.” These references allowed students to critically engage with both systemic programs and lived experiences.
The FGDs were coherent and easy to follow, with participants generally agreeing with each other’s points. The dialogue had thematic development, real moderator questions, and peer referencing (“I agree,” “expanding on what she said”), accurately mimicking a live facilitated session. This smooth flow contributed to readability but reduced the realism of natural group dynamics, which usually include interruptions, disagreements, or references to previous comments.
Some focus groups were more detailed and nuanced than others, leading to uneven quality across prompts. Certain themes were repeated frequently (e.g., lack of resources, poor infrastructure, government failures), reducing data richness. Finally, cultural specificity was lacking, as responses rarely incorporated references to local customs or informal employment practices that would have grounded the conversation in the Nigerian context. Overall, prompt 5 demonstrates strong performance across most evaluation indicators, with particular strengths in diversity and educational alignment, while limitations remain in emotional tone (See
Table 5).
4.6. Prompt 6: Project on Disaster Risk Reduction Plan for Yemen in the Context of Natural Disasters
The conversation captured realistic challenges faced in Yemen, such as floods destroying clinics, farmland, and roads. These issues resonated with the country’s actual humanitarian situation, making the personas relatable. The persona of Ahmad from Al Hudaydah, for instance, described how “the floods washed away many of our local water sources and infrastructure, and now the water is contaminated… sewage systems are damaged, and untreated waste is just sitting in the open”, a response grounded in specific local context and technical detail. However, the dialogue lacked cultural and linguistic markers specific to Yemen, such as references to religious institutions or local coping strategies, which would have deepened authenticity. Emotional nuance was also limited: while needs were clearly articulated, expressions of fear, frustration, or resilience were largely absent, leaving the responses somewhat formal rather than conversational.
The personas represented multiple regions and highlighted varied impacts across health, agriculture, food security, and shelter, offering a multidimensional perspective. Diversity within the seven male personas was particularly well captured, with geographic spread across Yemen’s ecological zones: coastal, desert, and highlands, showing differing vulnerabilities to floods, cyclones, and droughts. Still, differentiation between occupational roles could have been stronger. For example, while farmers and health workers might have distinct priorities in real contexts, the responses here occasionally overlapped and lacked individualized perspectives, which reduced the richness of the simulation.
The FDG provided students with rich material to analyze, covering essential themes such as food access, agricultural destruction, nutrition, and disease outbreaks. The dataset was also strongly linked to learning objectives, enabling students to identify needs by sector; conduct stakeholder mapping, risk assessment, and policy prioritization; and analyze compound vulnerabilities: such as how shelter loss leads to sanitation breakdown and health hazards. However, the lack of contradictory perspectives risked simplifying complex realities. Students were given structured answers with few opportunities to tackle the ambiguity and tension that often characterizes real field data.
The dialogue was structured and logically organized, following a realistic sequence from introductions to disaster impacts, access to resources, and responses to assistance. This progression modeled facilitation skills and the sequencing of data collection very effectively. Still, the exchanges lacked the natural rhythm of a genuine focus group. Personas responded in isolation without referencing or building on others’ contributions. This made the exercise feel closer to a sequence of interviews rather than a collective discussion.
Several gaps were evident. The lack of conversational tone and emotional expression reduced realism. Responses occasionally felt generic, with some repetition across personas. The omission of cultural sensitivity and local context, such as references to social tensions, coping practices, or displacement dynamics, created blind spots in the narrative. Although institutional players such as the Red Cross or UNICEF were mentioned once, Ahmad noted receiving “some initial support from humanitarian organizations like UNICEF and the Red Cross”, most responses lacked reference to NGOs or government agencies, reducing the realism that comes with institutional presence in real-life discussions. Without careful instructor guidance, students risk overestimating the accuracy and completeness of the data, potentially overlooking the importance of triangulation and interactions in real-world humanitarian research. A summary of Prompt 6’s performance across the five evaluation indicators is presented in
Table 6.
4.7. Prompt 7: Project on Supply Chain Blockage in Gaza During 2023 Conflict
The outputs exhibit a high degree of contextual realism, capturing Gaza’s multi-layered humanitarian crisis through specific and credible details (See
Table 7). Each persona reflects valid, sector-related challenges: including medical supply shortages, food insecurity, mental health strain, fuel scarcity, and the breakdown of infrastructure. The inclusion of issues such as closed borders, aid obstruction, and resource rationing mirrors authentic sociopolitical realities. The personas effectively represent how professionals from different sectors, including health workers, farmers, administrators, and NGO staff, would describe their experiences under siege conditions. However, while the tone remains grounded and factual, emotional nuance is limited. Expressions of exhaustion, frustration, or moral distress, which often accompany prolonged crisis work, appear subdued. As a result, the dialogue reads as informed and professional but lacks the human intensity characteristic of first-hand accounts.
Group 7’s main strength lies in the diversity of AI-generated personas. The ten personas represent multiple occupational domains, health professionals (doctors, nurses, epidemiologists), local government, agriculture, logistics, media, education, and humanitarian relief. Each role introduces distinct thematic priorities: health professionals emphasize burnout and ICU collapse, farmers discuss seed and feed shortages, administrators highlight bureaucratic barriers and political constraints, and aid workers focus on logistical blockages and informal trade networks. This differentiation enhances realism by reflecting a complex web of interdependencies. It also offers students opportunities to analyze contrasting needs and perspectives across institutional, community, and household levels.
The interviews offered multi-layered insights relevant to the learning objectives of humanitarian systems education. They enable analysis of stakeholder mapping, crisis interdependencies, and the cascading effects of disrupted supply chains on health, food, and livelihoods. The personas provide a foundation for students to conduct feasibility assessments for interventions and explore cross-sectoral strategies, such as digital inventory management or decentralized logistics solutions. By situating narratives within a politically charged context, the dataset also fosters critical thinking about aid neutrality, access negotiation, and resilience under blockade. These aspects directly support experiential learning outcomes that combine systems thinking, ethical reasoning, and practical problem-solving. The exercise followed a structured interview format rather than a conventional focus group, with all personas responding to the same 39 guiding questions. This format supports thematic comparability and consistency for analysis but limits the dynamic interaction typical of group discussions. There is minimal turn-taking or spontaneous referencing among participants. Nevertheless, a sense of collective experience emerges, as several respondents acknowledge shared obstacles and coping mechanisms, such as reliance on informal supply channels or community-based redistribution of resources. The logical sequencing of questions, from situational overview to solutions, ensures analytical clarity, even though conversational spontaneity remains low.
Despite its depth, the outputs present several limitations. The tone remains predominantly professional and impersonal, missing the emotional variability and interpersonal tension often present in real-world exchanges. Heavy repetition, such as recurring references to “missing protein” or “aid versus cultivation”, reduces linguistic authenticity and suggests scripted output. The absence of explicit political references, beyond surface-level mentions of border closures or governance failures, omits key structural determinants of the crisis. While the information is accurate and thematically rich, it reads more as an analytical survey than as a spontaneous focus group simulation.
4.8. Prompt 8: Project of Energy Challenges in Kenya Especially for the Massai Community
The interviews demonstrate a notable degree of realism, particularly through detailed depictions of daily routines, cultural references, and localized environmental observations. Personas such as Naeku Tepilit and Naserian Olelang effectively mirror the lived experiences of Maasai women, capturing the physical strain of firewood collection, smoke-related health issues, and environmental degradation (“the bushes and trees were closer to our village… now, many of the trees near our village have been cut down”). The inclusion of sensory and emotional language: such as fatigue, worry for children’s health, and fear of wild animals, adds to their credibility. However, certain statements are overly structured or formal (“If there is a way to make cooking easier without relying on firewood, I would be happy to try it”), which occasionally detracts from the spontaneous tone of genuine conversation. Emotional variability remains somewhat limited, with few moments of hesitation or frustration that would typify real human interaction.
The personas offer clear distinctions in social role and perspective: Naeku represents a younger mother balancing household responsibility, while Naserian portrays an older community leader focused on intergenerational knowledge and environmental stewardship. Both refer to similar hardships, such as firewood scarcity and smoke-related illness, but their differing life stages and authority levels enrich the dataset. The inclusion of details such as livestock ownership, traditional building practices using cow dung, and discussions around biogas potential demonstrates attention to socio-economic and cultural nuance. However, since both personas share gender and ethnic identity, cross-sectional diversity (e.g., male perspectives, NGO workers, or youth voices) is missing, limiting the range of stakeholder viewpoints.
The interviews align well with pedagogical objectives by offering students material for needs assessment, energy mapping, and culturally sensitive intervention design. They encourage critical thinking about sustainable energy transitions, affordability, and the intersection of gender, health, and environmental degradation. Mentions of potential solutions such as biogas and solar power provide opportunities for applied learning and scenario analysis. However, the uniform optimism toward technological alternatives could have been balanced by greater skepticism or uncertainty, both personas concluded with near-identical enthusiasm: Akinyi stated she was “open to it, but needed to understand more about how it works,” while Naserian declared she was “excited about the idea of biogas” and wanted to see more people try it, neither expressing meaningful doubt or resistance, which would better simulate the complexities of real-world community engagement.
The one on one format resembles structured interviews rather than FGDs, with limited interactive exchange, spontaneous commentary, or reference to shared community opinions, which reduces the collective dynamic that FGDs aim to model.
While the interviews succeed in contextual grounding and technical relevance, several limitations remain. The dialogues lack emotional fluctuations typical of real conversations; responses are polished and free from the minor contradictions or ambiguities that reflect authentic human reasoning. Technical depth is limited, few references are made to quantitative data, institutional initiatives, or specific local organizations beyond generic mentions of NGOs. As summarized in
Table 8, Prompt 8 scored highly on all evaluation indicators, while limitations were observed in emotional and technical variety.
4.9. Prompt 9: Project on Erratic Rainfall in Ethiopia
The interviews and focus groups exhibit a high level of authenticity and realism (See
Table 9). The personas mirror real participants, from Climate Resilient Green Economy (CRGE) professionals and meteorologists to NGO representatives and smallholder’s farmers, each reflecting appropriate vocabulary, expertise, and tone. The narratives balance technical and emotional realism, as seen in farmers describing tangible struggles: Lemlem, a 50-year-old farmer persona from Tigray, captured this grounded authenticity simply: “It’s severe here, especially with recent droughts. I’ve had to sell livestock to survive.” Meanwhile, institutional voices from personas add credible policy and scientific perspectives: Dr. Meron Kebede noted that “rainfall patterns have shifted dramatically. In some years, rains come late or are shorter, leading to dry spells during critical growing periods. In others, intense rains cause soil erosion and damage crops”, a response reflecting appropriate meteorological vocabulary and analytical depth. Although generally consistent and believable, the interviews could have included more emotional expressions (e.g., frustration, hope, or fatigue) to deepen realism and empathy.
The simulation effectively integrates diverse roles representing a broad cross-section of Ethiopian society and geography. Each persona contributes sector-specific insights: farmers emphasize local adaptation and survival, experts discuss predictive modeling and climate trends, and NGOs highlight gender equity and sustainability. The variety of professional and regional backgrounds ensures realism, though some overlap exists in responses concerning resource access and funding limitations.
The interviews are well aligned with educational objectives, enabling stakeholder mapping, policy evaluation, and needs assessment, while proposed solutions such as fog and soil, integration systems reinforce applied learning and invite reflection on cultural appropriateness and feasibility.
The structure combines both focus groups (among farmers) and structured interviews (with experts), maintaining coherence and consistency in tone. The farmer focus group provides a realistic conversational rhythm, showing similarities and subtle contrasts between participants’ experiences. For instance, Tesfaye’s persona curiosity about innovation is visible in his response: “I’ve seen drip irrigation in videos, and it looks promising. I’d love to learn more,” while Lemlem’s cautious skepticism surfaces in: “Yes, as long as I can see proof that it works” a natural contrast that mirrors the diversity of attitudes found in real community discussions. Although the discussions flow logically, they remain more sequential than interactive; limited back-and-forth reduces the spontaneity and dynamic exchanges typical of a live focus group.
Despite the strong technical and contextual realism, the simulation lacks deeper emotional resonance, few moments express frustration, urgency, or local sentimentality that would enhance authenticity. Political and social dimensions, such as land rights conflicts or regional inequities, are underexplored despite their relevance to sustainability. Furthermore, some interviews remain overly polished and formal, resembling policy dialogues more than participatory discussions. Future iterations could integrate more culturally grounded expressions and disagreement to enhance realism and educational richness.
4.10. Prompt 10: Project on Waste Management in Ghana
The interviews demonstrate a high degree of authenticity and realism, with personas that are convincingly positioned within Ghana’s waste management ecosystem. Each persona reflects relevant expertise and experience, contributing professional insights supported by credible examples from practice. Their tone is realistic and appropriately professional, maintaining clarity and relatability. References to national frameworks such as the Environmental Sanitation Policy, Integrated Solid Waste Management (ISWM), and the Extended Producer Responsibility (EPR) law enhance contextual accuracy.
The focus group transcripts further reinforce realism through vivid community-level accounts that mirror Ghana’s lived experiences: such as irregular waste collection, cost-related constraints, and reliance on open burning. The natural phrasing (“we burn when bins are full” or “collection is not always consistent”) captures real community sentiment effectively. The prompt presents strong diversity in roles and perspectives: featuring government officials, private sector representatives (Zoomlion, Nelplast), an EPA analyst, district officers, and community members.
This wide range of personas provides multi-level differentiation across governance, policy, operations, and citizen engagement. Socioeconomic and gender diversity are also represented through the household focus group, which spans age, class, and occupation. Each persona’s viewpoint aligns with their role: government representatives emphasize policy implementation and enforcement challenges, private companies focus on operational efficiency and financing, while citizens express frustration about inconsistent services and costs. The inclusion of culturally grounded references (Accra, Ashanti Region) and specific institutional names further enriches authenticity and differentiation. While the perspectives are distinct, they remain cohesive under the shared theme of Ghana’s systemic waste management challenges.
The prompt succeeds in aligning with educational goals by generating rich, multidimensional data that can foster student critical thinking. It introduces key learning points: such as public–private partnerships, community engagement, and regulatory frameworks, while highlighting constraints in financing, enforcement, and infrastructure. The simulated focus group on household waste management flows naturally and mirrors a realistic discussion format. Participants take turns expressing their experiences, sometimes building on each other’s points, such as shared complaints about irregular collection and cost concerns. The tone is conversational yet structured, with the moderator guiding transitions effectively. This coherence supports readability and immersion. While explicit disagreement or debate is limited, implicit contrasts in priorities (e.g., convenience vs. environmental awareness) enrich the discussion. The multi-perspective setup combining institutional interviews and a community FGD creates a layered simulation that resembles real-life stakeholder consultations in humanitarian or environmental policy contexts.
The primary limitations include a tendency toward idealism among officials and private sector personas. This reduces realism and critical tension. Emotional engagement is minimal; while the technical and institutional perspectives are strong, the human voice, expressing frustration, skepticism, or hope, is less pronounced. Additionally, the absence of explicit discussion on political and governance challenges (e.g., corruption, regulatory inertia) slightly weakens the depth of analysis. Finally, while the focus group offers credible local insight, greater interaction: agreement, debate, or referencing others’ remarks, would further strengthen the simulation’s dynamism (See
Table 10).
4.11. Cross-Case Comparative Analysis
To synthesize findings across all ten prompts,
Table 11 presents the aggregated scores for each of the five evaluation indicators. This cross-case overview reveals consistent patterns in both the strengths and limitations of AI-generated personas across diverse humanitarian contexts.
Educational Alignment emerged as the strongest indicator, achieving a perfect mean score of 5.00 across all ten prompts. This finding suggests that regardless of context, geographic setting, or prompt complexity, AI-generated personas consistently produced data that was relevant, usable, and aligned with the course’s learning objectives. Students were reliably able to engage in stakeholder mapping, needs assessment, and systems thinking exercises using AI-generated dialogues.
Diversity of Perspectives scored nearly as high, with a mean of 4.90. In nine out of ten cases, prompts successfully generated a broad range of stakeholder voices across demographic, occupational, and geographic lines. The single exception was Prompt 2, where the women’s focus group displayed limited intra-group variation, reinforcing stereotypes rather than reflecting the heterogeneity of displaced women’s experiences.
Authenticity and Realism achieved a mean score of 4.38, indicating that AI personas were generally convincing in their contextual grounding, vocabulary, and emotional expression. Scores were highest in prompts involving highly specific local contexts, such as Prompt 2 (WASH professionals in Lebanon), Prompt 8 (Maasai women in Kenya), and Prompt 9 (climate professionals in Ethiopia), where detailed prompt design elicited more realistic and nuanced responses. Scores were comparatively lower in prompts covering broader or more complex settings, where AI responses tended toward polished generality rather than authentic specificity.
Group Dynamics and Coherence was the most variable indicator, with a mean of 3.80 and scores ranging from 2.5 to 4. The lowest score was recorded for Prompt 4, which involved 30 simultaneous AI personas representing Syrian women, a scale that exceeded ChatGPT’s capacity to maintain dynamic group interaction, resulting in responses that resembled sequential individual interviews rather than a genuine focus group discussion. This finding highlights an important practical limitation: the effectiveness of AI-generated FGDs appears to diminish as the number of simultaneous personas increases, and single or small-group interview formats tend to yield more coherent and interactive exchanges.
Limitations and Gaps consistently recorded the lowest scores across all prompts, with a mean of 3.20. This indicator captures a recurring and cross-cutting weakness: AI personas across all ten contexts tended to produce emotionally flat, overly polished responses that lacked spontaneity, contradiction, hesitation, and cultural specificity characteristic of real human interaction. This pattern was observed regardless of geographic context or thematic focus, suggesting it reflects a structural limitation of current large language models rather than a prompt design issue specific to individual groups.
Taken together, these cross-case patterns suggest that AI personas are most effective as pedagogical tools when evaluated against educational and diversity objectives, and least effective in replicating the dynamic, unpredictable, and emotionally layered nature of real fieldwork.
Our findings are further supported by the inter-rater reliability analysis presented in
Table 12. Across all 10 prompts, the two independent evaluators achieved an overall exact agreement rate of 84%, with a mean absolute difference of just 0.15 points. This level of consistency suggests that the evaluation framework produced stable and reproducible judgments across raters, lending credibility to the aggregated scores reported in
Table 11.
Agreement was not, however, uniform across all indicators. Educational Alignment and Diversity of Perspectives yielded perfect inter-rater agreement (100% exact, mean |difference| = 0.00), suggesting that these dimensions are the most objectively assessable within the framework. Authenticity and Realism achieved 80% exact agreement and 90% within ±0.5 (mean |difference| = 0.15), reflecting a high but slightly more interpretive judgment, as evaluators occasionally diverged on the degree of emotional nuance or contextual grounding present in a prompt. Group Dynamics and Coherence showed 80% exact agreement (mean |difference| = 0.20), with the two divergences concentrated in Prompts 2 and 4, precisely the prompts where format limitations were most pronounced, suggesting that disagreement between raters mirrored the genuine ambiguity of those outputs rather than inconsistency in the framework itself. Limitations and Gaps recorded the lowest inter-rater agreement at 60% exact (mean |difference| = 0.40), with divergences appearing in Prompts 1, 6, 7, and 9. The lower agreement here does not undermine the validity of the scores, but rather reinforces the recommendation that this indicator be used alongside instructor reflection and student debriefing rather than as a standalone metric.
Finally, both
Table 11 and
Table 12 present a coherent picture: the indicators on which AI personas performed strongest are also those on which raters agreed most, while the indicators that exposed the deepest structural limitations of AI-generated dialogue are those that proved hardest to rate with perfect consistency.
5. Discussion
While AI personas present clear pedagogical advantages, their use in higher education introduces multiple challenges and considerations. Technical issues are a primary concern: AI systems often struggle with maintaining accuracy, particularly in complex or interdisciplinary topics. Errors, hallucinations, and limited adaptability across cultural and learner-specific contexts can lead to misinformation or reduced personalization (
Luckin et al., 2020). These technical inconsistencies can distort learners’ understanding and undermine trust in AI mediated learning. Ethical concerns also arise, as AI-generated outputs are subject to bias from training data, raising issues of fairness and equity in learning. Additionally, the reliance on student data for personalization introduces significant privacy and data security risks (
Johnson & Lester, 2018). Another major challenge is the human interaction gap. While AI personas can simulate conversation and empathy, they lack genuine emotional intelligence. This can diminish learners’ ability to engage with subtle social cues or form meaningful connections—an especially critical concern in humanitarian education (
Kim et al., 2022). Furthermore, overdependence on AI may hinder social and emotional learning by limiting interpersonal classroom interactions. Implementation barriers also complicate the educational use of AI. Creating and deploying effective AI personas is resource-intensive, requiring substantial technical infrastructure and institutional support. Additionally, educators may resist adoption due to skepticism, lack of training, or fears of automation replacing human roles (
Bettayeb et al., 2024). In addition, impact assessment limitations remain significant. Measuring the impact of AI personas is difficult, as current assessment tools often emphasize quantifiable metrics, which can neglect creativity, ethical reasoning, and critical thinking. Developing comprehensive evaluation frameworks is essential to accurately capture learning outcomes (
Samuel et al., 2024).
Despite these limitations, the integration of AI personas into educational practice marks a transformative shift in curriculum design and learner engagement. As demonstrated by HEHI course, AI personas—particularly in fragile, low-resource, or conflict-affected environments—can offer meaningful alternatives to in-person experiential learning. By enabling students to simulate fieldwork through ChatGPT personas, HEHI preserved educational continuity and cultivated ethical and critical capacities vital to humanitarian engineering. However, realizing the full potential of AI personas requires proactive responses to the technical, ethical and pedagogical challenges outlined above. Educational institutions must prioritize human–AI collaboration, with AI augmenting rather than replacing human educators. Especially in contexts requiring empathy, nuanced judgment, and socio-emotional learning, human facilitation remains indispensable. The future of AI personas lies in their capacity to enhance—not replace—human-centered education.
The findings of this study both confirm and extend prior research on AI-generated personas in educational settings.
Akkurt et al. (
2025) found that ChatGPT-simulated clients in counselor training produced overly agreeable responses, lacked emotional nuance, and defaulted to cultural neutrality unless explicitly prompted, patterns directly mirrored across all ten prompts analyzed in this study, where emotional flatness and cultural generalization were the most consistent cross-cutting limitations. Similarly,
Sabbaghan and Brown’s (
2024) PEARL framework demonstrated that AI personas could effectively scaffold research interview skills in one-on-one settings a finding our study extends by showing that this effectiveness is preserved across diverse humanitarian contexts, but diminishes when the number of simultaneous personas increases, as evidenced by Prompt 4’s notably lower Group Dynamics score when 30 Syrian women were simulated simultaneously. Our results further support
Barambones et al. (
2024) and
Bettayeb et al. (
2024), who emphasized that prompt specificity is a critical determinant of response quality a conclusion reinforced by the visible variation in authenticity and cultural grounding across our ten prompts, where more contextually detailed prompts consistently produced richer and more differentiated responses. At the same time, our cross-case analysis adds a dimension absent from much prior research: by evaluating AI personas against five indicators across ten diverse humanitarian contexts. We demonstrate that Educational Alignment is the most robust and consistent strength of AI-generated personas, while Group Dynamics and Limitations remain structurally constrained regardless of prompt quality, a finding that has important implications for how instructors frame the use of this tool.
Finally, the effectiveness of the generated answers depends heavily on the specificity of the prompts provided, the more realistic it is, the more engaging responses become. General or vague prompts may lead to irrelevant situations and less effective scenarios. Loaded prompts in emotions can enhance the real mimic of human expressions and response that we receive from AI. However, this can result in biased content because AI tries to conform emotional tones rather than providing a neutral response (
Barambones et al., 2024). Another limitation is AI’s unpredictability of human dialogue. Although structured responses are useful for guided training and conversations for students, they may affect the spontaneity, adaptability of interactions, an essential element for role-playing activities (
Maurya, 2024).
5.1. Can Students Rely on AI for Education?
One of the most notable advantages of AI in education lies in its ability to realistically simulate diverse conversations. Personas can represent multiple stakeholders, giving multiple or different perspectives within a controlled environment. This feature is useful in a crisis setting, where conducting real fieldwork is often challenging or impossible (
Towoju, 2024). Thus, AI creates a safe space for learners to practice before dealing with real-life experience helping them be fully prepared and confident in any situation (
Mason, 2023). However, despite these benefits, AI’s limited adaptability in real time remains a key limitation. Unlike human interviewers who can intuitively adjust their tone or reactions based on the direction of the discussion, AI systems often struggle to respond dynamically to unexpected conversation turns (
Hill et al., 2023). This might not prepare or help students to face the unpredictability of human interactions in the field. Additionally, AI tools are unable to distinguish between patterns derived from biases or incomplete data sets. Therefore, students should be educated to closely use AI, verifying sources to avoid misinformation. This approach ensures that AI remains a supportive educational tool rather than a substitute for critical inquiry and human judgment.
5.2. Recommendations for Teaching Practice
The findings of this study offer concrete guidance for instructors considering the use of AI personas in experiential learning courses.
When AI personas are appropriate: AI personas are most effective as a pedagogical substitute for fieldwork when physical access to communities is impossible or unsafe as demonstrated by the HEHI 303 case during the 2024 Lebanon conflict. In such contexts, they reliably serve educational alignment objectives, enabling students to practice stakeholder mapping, needs assessment, thematic coding, and systems thinking. They are also appropriate as a preparatory tool before real fieldwork, allowing students to rehearse interview techniques, develop question guides, and build familiarity with stakeholder dynamics in a low-stakes environment (
Mason, 2023).
When AI personas should not replace fieldwork: AI personas should not be treated as equivalent substitutes for real human interaction in contexts where emotional complexity, cultural specificity, and interpersonal dynamics are central learning objectives. As this study consistently showed, AI-generated responses lack the spontaneity, contradiction, hesitation, and cultural nuance characteristic of authentic human dialogue, limitations confirmed by
Akkurt et al. (
2025) and
Maurya (
2024).
How instructors should use AI personas: There are different strategies to maximize the benefits of AI in education while mitigating its limitations: Integrate AI role playing with real interview experiences: human interaction should complement AI-driven simulations to provide a balanced learning experience. Students should be encouraged to conduct real-world interviews alongside AI-based exercises, allowing them to compare insights, refine their interpersonal skills, and complete their study. Train students in critical analysis of AI-generated responses: Given the limitations of AI, it is essential that students learn to evaluate responses critically. They should be able to identify inconsistencies, recognize potential biases, and compare AI-generated information with real world evidence or evidence-based studies to ensure accuracy and validity. To enhance the effectiveness of AI in education, future research should focus on the following: Hybrid AI training models: combining AI-driven role play with human facilitators can improve engagement, contextual understanding and realism. This blended approach allows students to benefit from the scalability of AI while retaining the adaptability that comes with human guidance. Investigating biases in AI-generated narratives: it is crucial to examine how AI content reflects underlying biases, particularly in humanitarian context. Understanding and minimizing these biases will strengthen the reliability and ethical integrity of AI tools. Future studies should prioritize refining AI systems to generate unbiased, contextually accurate and culturally sensitive information.