Next Article in Journal
Poly(Neutral Red)–Silver Nanorods–Carbon Nanotubes Composite-Based Ratiometric Electrochemical Sensor for Rapid Detection of Histamine in Crayfish
Previous Article in Journal
Characterization of Species- and Sex-Dependent Volatile and Non-Volatile Flavor Profiles in Three Commercially Important Mussel Species
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Can ChatGPT Reflect Professional HACCP Judgments? A Comparative Study in Hospitality Food Safety

by
Despoina Maria Konstantinidi
1,
Elisavet Stavropoulou
1,
Agathangelos Stavropoulos
1,
Chrysoula (Chrysa) Voidarou
2,
Christina Tsigalou
1,
Vassiliki Pitiriga
3 and
Christos Stefanis
1,*
1
Laboratory of Hygiene and Environmental Protection, Medical School, Democritus University of Thrace, 68100 Alexandroupolis, Greece
2
Department of Agriculture, School of Agriculture, University of Ioannina, 47100 Arta, Greece
3
Department of Microbiology, Medical School, University of Athens, 11527 Athens, Greece
*
Author to whom correspondence should be addressed.
Foods 2026, 15(16), 2916; https://doi.org/10.3390/foods15162916
Submission received: 23 July 2026 / Revised: 13 August 2026 / Accepted: 17 August 2026 / Published: 20 August 2026
(This article belongs to the Section Food Quality and Safety)

Abstract

The implementation of Hazard Analysis and Critical Control Points (HACCP) systems remains a cornerstone of food safety management. However, their effectiveness is influenced by organizational, human, and technological factors. This study investigates professional perceptions of HACCP implementation and explores the extent to which a large language model (LLM)—ChatGPT 4.1—can approximate human judgments in this domain. A structured questionnaire was administered to 90 professionals operating in food service, hospitality, industry, and consultancy. Responses were compared with outputs generated by ChatGPT 4.1 via a standardized multi-persona prompting protocol simulating five professional roles. To enhance response stability and minimize stochastic variation, each question was submitted in independent zero-shot sessions over multiple iterations. Responses were analysed across three thematic dimensions: barriers to HACCP implementation, perceived benefits, and digital readiness. Spearman correlation analysis of human responses revealed a systemic “training–turnover association,” where high staff turnover (r = 0.62, p < 0.01) was significantly associated with difficulties in maintaining continuous training. Statistical benchmarking using one-sample t-tests showed that the tested model generated significantly higher ratings for implementation barriers (p < 0.001) and digital readiness (p < 0.001) than the corresponding human assessments. These findings suggest that while ChatGPT can approximate aggregated professional perceptions in certain areas, notable divergences persist in operational and readiness-related domains. The study contributes methodological insights into human–AI comparative research and highlights opportunities and limitations of AI-supported decision-making in food safety management systems.

1. Introduction

Over the past decade, the regulatory landscape governing food safety in the European Union has evolved considerably, strengthening expectations for systematic hazard control and more rigorous verification of food safety practices across the hospitality and mass catering sectors. Between 2016 and 2025, successive updates to EU hygiene legislation and the modernisation of the farm-to-fork strategy shifted Hazard Analysis and Critical Control Points (HACCP) from a procedural compliance tool toward a performance-oriented framework that emphasises preventive control, documentation accuracy and continuous improvement. These changes have particularly affected operationally complex environments such as hotels and large catering units, where multiple production sites complicate the consistent application of HACCP procedures [1,2,3,4,5,6,7,8].
The HACCP (Figure 1) system constitutes the primary international framework for ensuring food safety in food service and hospitality environments. Numerous studies conducted in Europe and globally have demonstrated that the degree to which HACCP is successfully implemented depends heavily on staff training and competence, the overall food safety culture within an establishment, and the commitment of management to continuous enforcement of the system [9,10,11]. Despite its widespread regulatory adoption, substantial variability in compliance persists across food service and hotel operations, suggesting that legislative requirements do not necessarily translate into consistent day-to-day practice [12,13,14].
In parallel with regulatory developments, increased attention has been directed toward the financial and reputational consequences of foodborne disease outbreaks in hospitality settings. Evidence from recent years indicates that food safety failures may lead not only to direct economic losses but also to long-term reputational damage, intensified inspections, and operational disruptions, particularly in complex food service environments such as hotels and large catering units [16,17]. These risks have reinforced the importance of robust HACCP systems as core components of organizational risk management strategies within the hospitality sector.
Alongside these structural challenges, advances in digital technologies and computational risk assessment have introduced new opportunities for supporting food safety management. In particular, large language models (LLMs) and related artificial intelligence (AI) tools have demonstrated the capacity to process unstructured textual data, synthesize regulatory information, and identify potential food safety signals in diverse domains [18,19,20]. While such technologies have shown promise in decision support and analytical applications, the literature consistently emphasises that their outputs require careful interpretation, reproducibility controls, and alignment with domain-specific expertise [21,22].
Recent developments in large language models (LLMs) have shown their capacity to synthesize domain-specific information and produce structured technical reasoning. In food safety management, where professional decisions frequently rely on practical experience and contextual interpretation of guidelines, examining whether LLM-generated responses align with the perceptions of professionals involved in HACCP implementation is of increasing research interest. This perspective allows the evaluation of whether AI-generated reasoning reflects real operational conditions or primarily reproduces generalized regulatory logic.
Within this context, the present study provides a structured comparative assessment between the perceptions of professionals involved in food service and hospitality operations and the responses generated by a large language model. By examining areas of convergence and divergence between human judgments and LLM outputs across key HACCP dimensions, the study aims to contribute to the understanding of the perceptual alignment, limitations, and potential role of AI-based simulations as complementary tools in food safety management and professional training.
Accordingly, the study tested the following hypotheses: (H1) Professional perceptions of HACCP implementation barriers, benefits, and AI/digitalization differ across professional roles; (H2) LLM-generated assessments differ from human professional assessments across the examined HACCP dimensions; and (H3) Significant associations exist among selected organizational, implementation, and AI-related perceptions within the human dataset. These hypotheses were examined within the exploratory framework of the study, without implying causal relationships.

2. Materials and Methods

2.1. HACCP Questionnaire Design and Data Collection Procedures

The questionnaire was developed to systematically investigate perceptions, experiences, and practical aspects related to the implementation of the Hazard Analysis and Critical Control Points (HACCP) system in food service establishments and hotel operations (Supplementary Material S1). Its design was informed by contemporary methodological frameworks for KAP (Knowledge, Attitudes, Practices) studies in food safety and public health research, drawing on recent evidence from the food service and food technology literature [23,24,25,26].
The development process followed a multi-stage structure. An initial pool of items was generated based on international literature concerning HACCP implementation, regulatory compliance, operational barriers, and organizational food safety culture [27,28]. These items were organized into three sections:
  • Barriers to HACCP implementation,
  • Benefits of HACCP implementation,
  • AI and Digitalization.
This thematic structure is consistent with contemporary multi-factor survey instruments used in food industry research [23,24].
The first validation stage involved a qualitative professional review by four specialists (n = 4) with expertise in food safety, food technology, hospitality management, and quality assurance. Reviewers assessed the clarity, relevance, conceptual structure, and coherence of the proposed items [23,24,29]. A formal quantitative Content Validity Index (CVI) was not calculated, as this stage was designed as a qualitative content review rather than a psychometric content-validity assessment. Based on the reviewers’ feedback, items requiring clarification were revised with respect to wording, terminology, and conceptual consistency. Prior to dissemination, a pilot test (n = 10) was subsequently conducted with food service and hospitality professionals to assess question comprehension, thematic flow, terminology precision, questionnaire functionality, and completion time [23,30].
The pilot sample was intentionally limited to 10 participants because its purpose was not statistical or psychometric validation but rather qualitative pre-testing of questionnaire comprehensibility, terminology, thematic flow, functionality, and completion time before full-scale dissemination. Methodological guidance indicates that approximately 10 participants, or even fewer, may be sufficient when pilot testing focuses on item wording, clarity, formatting, and ease of administration. Cognitive questionnaire pre-testing commonly employs small samples of approximately 5–15 participants per testing round. Feedback from the pilot resulted in further refinement of item wording, simplifications of technical terminology for non-specialists, and the inclusion of response options suitable for academic or research-oriented participants who evaluate HACCP conceptually rather than operationally [31,32].
The finalised questionnaire was distributed via Google Forms between 1 April and 30 April, allowing participants with professional or research experience in HACCP to complete the survey anonymously and voluntarily [23,24].
A purposive sampling strategy was employed, targeting individuals with professional or research experience related to HACCP—including food handlers, chefs and kitchen supervisors, quality managers, hotel F and B managers, HACCP consultants, and academic or research personnel. Distribution occurred through professional associations, hospitality networks, food service organizations, and academic channels.
The standardized format facilitated the reliable transformation of each item into a prompt and supported valid comparisons between human responses and model-generated outputs, in line with emerging methodological frameworks for human–LLM comparative research [21,22]. The operationalization of this process is described in detail in Section 2.2. The original Greek introductory text and the full questionnaire are provided in the Supplementary Material of this article.

2.2. Standardised ChatGPT Simulation Protocol and Comparative Statistical Analysis of Human and AI-Generated Responses

This study compares participants’ responses to the HACCP questionnaire in food service and hotel settings with those generated by the large language model ChatGPT 4.1. To capture stakeholder-specific perspectives, five distinct professional persona-based prompts were employed, simulating a multi-stakeholder expert panel (e.g., HACCP Consultant, Production Staff, Manager/Owner, Auditor, Researcher/Academic). Persona definitions were limited strictly to professional role framing and did not involve narrative role-playing, emotional context, or adaptive interaction. Each questionnaire item was administered in an independent, zero-shot session using standardized prompts to ensure consistency, neutrality, and comparability across personas. This standardized approach minimizes bias and ensures comparability between human and computational responses [22]. Zero-shot prompting was applied in the sense that no examples, prior answers, or iterative refinement were provided, and each response was generated independently. The prompts specified only the response format (e.g., a Likert 1–5 rating accompanied by a brief justification), following best practices outlined in recent methodological studies on human–AI comparative research [21]. Reproducibility was ensured by using identical prompts across all trials and submitting each question in an independent, memory-free session. For each professional persona, the prompt was executed in multiple independent runs (n = 30) to reduce stochastic variability and obtain a stable representative score. A fully standardized prompt structure and identical generation conditions were maintained across runs to minimize variability attributable to prompt formulation. The mean across repeated runs was therefore used as a stabilized computational reference for each questionnaire item. Importantly, this LLM-derived value was not interpreted as a population mean or population parameter, but as an experimental benchmark generated under the predefined simulation protocol. Accordingly, the one-sample t-test was used as an exploratory benchmark-based inferential procedure rather than as an estimation of an underlying LLM population parameter. The resulting p-values were therefore interpreted only as evidence of divergence between the distribution of human ratings and the predefined computational benchmark under the specified prompting conditions, while the direction and magnitude of divergence were assessed descriptively through Δ values.
Correlation analysis using a Spearman coefficient was applied exclusively to the human dataset to assess internal coherence and structural associations among questionnaire items, whereas alignment between human and ChatGPT responses was examined through descriptive comparisons and one-sample t-tests. Cronbach’s α was calculated exclusively for the human dataset to assess the internal consistency of the questionnaire scales [25,33]. The coefficient was not computed for the computational dataset, as ChatGPT 4.1 provides a single deterministic answer per question without within-item variance. Therefore, no multi-observation sample exists to permit a perceptual alignment estimate [21]. LLM mean scores represent the average of responses generated across the five professional personas for each questionnaire item. Because ChatGPT 4.1 outputs are probabilistic, a standardized prompt and identical generation conditions were used to reduce variability and the resulting persona responses were aggregated to obtain a stable benchmark value.
Differences between ChatGPT 4.1 and human mean scores (Δ) were calculated to facilitate comparative interpretation. A one-sample t-test was applied to examine whether human participant mean scores differed significantly from a fixed AI-derived reference value. The LLM mean score for each item was treated as a benchmark rather than a sampled population. This approach allowed assessments of whether human responses statistically confirmed or deviated from AI-based expectations [22,26].
All LLM experiments were conducted in May 2026 after completion of the online questionnaire distribution, using GPT-4.1 through the ChatGPT web interface (OpenAI, San Francisco, CA, USA). No API was used, therefore API version, API-specific model snapshot/ID, or API settings are not applicable. The model designation displayed to the researchers through the interface at the time of inference was GPT-4.1. The complete prompts used for all five professional personas, together with the response format instructions, are provided in the Supplementary Material. Because the experiments were conducted through the ChatGPT 4.1 web interface, backend system instructions and API-level configuration parameters were neither researcher-controlled nor directly accessible and therefore cannot be reported as experimental settings.
Figure 2 presents the sequential stages of the study: literature review; development, professional validation, and pilot testing of the HACCP questionnaire; online data collection through purposive sampling; data cleaning and coding; statistical analysis of human responses; zero-shot prompt–based LLM simulation using ChatGPT 4.1; and the final comparative evaluation of human and AI-generated responses.

3. Results

3.1. Professional Survey Outcomes and ChatGPT Simulation Outputs

The results of the ChatGPT 4.1 Simulation Protocol indicate that the model consistently produces high scores across all sections of the questionnaire, reflecting consistently high model-generated ratings across the examined items related to HACCP compliance, operational barriers, and the potential benefits of both HACCP and digitalization. As shown in Table 1, all five sets of simulated responses—corresponding to the selected professional profiles (“Consultant,” “Production Staff,” “Manager/Owner,” “Auditor,” and “Researcher/Academic”)—converge toward similarly high ratings regarding perceived barriers, perceived benefits, and the potential of artificial intelligence to support HACCP systems.
In Section B (Barriers), all profiles assign particularly high scores to human-resource-related challenges, such as lack of staff motivation (Q6) and high employee turnover (Q7), with means approaching or exceeding 4.5. The “Consultant” and “Auditor” profiles produced the highest scores for staff turnover, while “Auditor” and “Academic/Researcher” identified deficiencies in prerequisite programs (Q11), reflecting a more systemic and standards-oriented interpretation of HACCP obstacles. By contrast, the “Production Staff” profile assigns the highest value to the documentation burden (Q9), while the “Manager/Owner” profile consistently rates this specific barrier lower than the other groups (3.50).
Section C (Benefits) shows near-uniform and very strong agreement that HACCP improves legal compliance (Q14), reduces the likelihood of customer complaints (Q15), enhances business reputation and credibility (Q16), and fosters a food safety culture among employees (Q18). Although a small number of profiles exhibit scores slightly below 4.5—primarily the Production Staff profile—the overall distribution in Section C remains very high (4.30–4.97), indicating a strong, cross-role consensus on the perceived benefits of HACCP. Τhe “Researcher/Academic” profile showsthe highest agreement regarding the role of HACCP in promoting a food safety culture (4.93).
In Section D (AI/Digitalization), the five simulated professional profiles present consistently positive evaluations of the potential contribution of artificial intelligence to HACCP processes. Mean scores across questions Q19–Q22 range from 3.70 to 4.83, indicating substantial agreement on the usefulness of AI-driven tools in food safety management. The strongest endorsement appears for automated monitoring of CCPs and PRPs (Q20), which achieves the highest overall mean (4.65), with Auditors (4.83) and Researchers/Academics (4.80) expressing the greatest confidence. For AI-supported decision-making through Large Language Models (Q21), the highest values are again observed for the Researcher/Academic (4.73) and Auditor (4.67) profiles, suggesting a technologically receptive stance among roles traditionally engaged in analytical evaluation and regulatory oversight.
A similar pattern emerges for AI-assisted HACCP training (Q19), where Consultants (4.40) and Researchers/Academics (4.77) provide the most favourable assessments while Production Staff report the lowest value (3.97), indicating a more cautious orientation among personnel working primarily in operational settings. The same gradient is observed in willingness to participate in AI-based HACCP training programs (Q22), with Researchers/Academics showing the highest level of openness (4.60), followed by Auditors (4.33) and Managers/Owners (4.07), and again the lowest score assigned by Production Staff (3.70). Taken together, these findings demonstrate that enthusiasm for digital transformation increases with greater exposure to structured decision-making and decreases among roles whose daily responsibilities are dominated by manual tasks. The ChatGPT 4.1 simulation produced consistently high scores across all five professional profiles (overall mean X¯ = 4.50), indicating broad alignment with the principles and perceived value of HACCP.
In Section B, human-resource-related barriers—particularly high staff turnover (Q7, X¯ = 4.77) and limitations in continuous training—were rated as the most critical across profiles. The second largest discrepancy concerned the documentation burden (Q9), to which Production Staff assigned the highest importance (4.83) and Managers/Owners the lowest (3.50), reflecting contrasting operational versus strategic viewpoints. Section C showed near-unanimous agreement, with all benefit-related items exceeding a mean of 4.50. The strongest overall benefit concerned enhanced business reputation (Q16, X¯ = 4.80), while the Researcher/Academic profile placed the highest emphasis on promoting a food safety culture (Q18, 4.93). In Section D, attitudes toward digitalization diverged: Auditors and Researchers/Academics demonstrated the highest confidence in AI-supported monitoring and decision-making, whereas Production Staff expressed the lowest willingness to engage in AI-based training (Q22, 3.70).
The study sample consisted of 90 professionals (n = 90) active in the food safety and hospitality sectors in Greece. The demographic and professional characteristics of the participants are summarized in Table 2.
A total of 90 respondents participated in the study, representing a diverse range of professional roles and levels of experience related to HACCP implementation. With respect to experience in HACCP, the majority of participants reported substantial professional exposure: 37% had more than 10 years of experience, while 22% reported 4–7 years and 18% reported 8–10 years of experience. Participants with limited experience were fewer, with 15% reporting 1–3 years and 8% reporting less than 1 year of involvement in HACCP-related activities.
Regarding organization type, respondents were primarily employed in mass catering or catering units (42%) and hotel units (35%), reflecting the study’s focus on hospitality and food service settings. Smaller proportions were affiliated with consulting or auditing firms (12%) and public authorities or regulatory bodies (11%), providing additional regulatory and advisory perspectives. In terms of training background, a large majority of participants (82%) reported having received certified HACCP training, while 18% indicated no official certification, suggesting that most responses were informed by formal training and professional qualification.
The largest groups consisted of HACCP consultants or auditors (n = 33) and quality or food safety managers (n = 32). Additional roles included managers or business owners (n = 10), researchers or academics (n = 7), and production or kitchen staff (n = 5), while a small number of respondents (n = 3) reported other food safety–related positions, such as regulatory authority staff or inspectors.
Finally, participants reported professional activity across multiple geographical regions of Greece, with the highest representation from Attica (n = 38) and Central Macedonia (n = 24). Smaller numbers were recorded from Eastern Macedonia and Thrace (n = 6), the Peloponnese (n = 5), Crete (n = 4), and the Ionian Islands (n = 4), while nine respondents reported activity in other regions, including Central Greece Regions, the North and South Aegean, Thessaly, Western Greece, Epirus, and Western Macedonia.
The results derived from the human questionnaire responses indicate a high level of overall agreement regarding both the challenges and benefits associated with HACCP implementation, while also revealing meaningful variations across professional roles, particularly in relation to operational constraints and digital transformation (Table 3).
In Section B (Barriers), the highest mean scores consistently relate to human resource–driven challenges. High staff turnover (Q7) emerges as the most critical barrier overall (Mean = 4.59), with particularly elevated scores reported by Auditors (4.70) and Researchers/Academics (4.65). Similarly, the inability to ensure continuous training (Q12) is rated as a major obstacle (Mean = 4.42), underscoring the dependence of HACCP systems on sustained knowledge transfer and skill development.
Issues related to staff motivation and commitment (Q6) and management support (Q10) also receive high evaluations, with mean scores exceeding 4.3. Auditors and Researchers assign greater importance to management commitment (4.55 and 4.45, respectively), reflecting a more systemic understanding of HACCP performance as an outcome of organizational governance rather than isolated operational practices. In contrast, the lack of specialized technical consultants (Q13) records the lowest overall mean (3.63), indicating that while external expertise is valued, it is perceived as less critical than internal organizational and human capital factors.
In Section C (Benefits), near-unanimous agreement is observed across all professional profiles. All items yield mean scores above 4.5, confirming a strong shared perception of HACCP as an effective and value-adding system. The highest overall rating is observed for improved compliance with legislation (Q14) (Mean = 4.82), with Auditors and Researchers/Academics assigning almost maximal scores (4.90 and 4.95, respectively). Benefits related to business reputation and credibility (Q16) and customer trust (Q17) are also rated extremely highly across all roles, highlighting that HACCP is widely perceived not merely as a regulatory obligation but as a strategic asset. The promotion of a food safety culture (Q18) shows slightly greater variation, with Researchers/Academics assigning the highest value (4.80), consistent with a longer-term and more conceptual perspective on organizational change.
Section D (AI and Digitalization) reveals a more nuanced pattern. While overall perceptions remain positive, mean scores are comparatively lower than those observed for HACCP benefits. The potential of AI to support HACCP training (Q19) and automated monitoring of CCPs and PRPs (Q20) receives favourable evaluations (Means = 4.10 and 4.31, respectively), particularly from Auditors and Researchers/Academics. In contrast, Production Staff consistently report lower scores, most notably regarding LLM-based decision support (Q21) (3.50), indicating greater scepticism toward AI systems that directly influence daily operational decision-making. Nevertheless, willingness to participate in AI-enabled HACCP training programs (Q22) remains relatively high across all groups (Mean = 4.19), suggesting openness to gradual digital integration when framed as supportive rather than substitutive.

3.2. Benchmarking Model-Human Consistency

The below bar chart (Figure 3) displays Cronbach’s alpha (α) coefficients for the three main sections; Barriers to HACCP (α = 0.84), Perceived Benefits (α = 0.89), and AI/Digitalization Perspectives (α = 0.86) and the overall (Total) questionnaire. All values exceed the 0.80 threshold, indicating high perceptual alignment and professional-grade internal consistency for the study sample (n = 90).
Figure 4 presents statistically significant correlation coefficients (r) derived exclusively from the human participant HACCP questionnaire (n = 90). Colour intensity represents the strength of positive associations, ranging from weaker (red) to stronger (green) correlations. Only correlations reaching statistical significance at the p < 0.05 level are displayed.
The correlation analysis of the human participant HACCP questionnaire (n = 90) identified several statistically significant positive associations between selected survey items (p < 0.05). The strongest correlation was observed between AI training (Q19) and willingness to participate in AI-based training programs (Q22) (r = 0.68). A high correlation was also recorded between high staff turnover (Q7) and inability to ensure continuous training (Q12) (r = 0.62).
Moderate positive correlations were identified between lack of staff motivation (Q6) and management support and commitment (Q10) (r = 0.54), as well as between AI-based monitoring of CCPs and PRPs (Q20) and the use of Large Language Models for HACCP decision-making (Q21) (r = 0.59). Additionally, a moderate association was observed between improved legislative compliance (Q14) and enhanced business reputation and credibility (Q16) (r = 0.51). Finally, a statistically significant but weaker correlation was found between documentation burden (Q9) and financial constraints/equipment renewal costs (Q5) (r = 0.41).
Table 4 summarizes the mean scores of the human participants (n = 90) and ChatGPT-generated responses across all HACCP questionnaire items (Q5–Q22), along with the calculated difference Δ (LLM − Human). Overall, the tested model produced consistently high ratings across all dimensions, with an overall mean score of 4.51 compared to 4.31 for human respondents, indicating a general upward bias in model-generated evaluations.
In the Barriers to HACCP section, the LLM consistently produced higher mean scores compared to human respondents across almost all items, indicating a systematic tendency to assign greater weight to structural and organizational constraints. The largest positive deviations were observed for financial constraints and equipment renewal costs (Q5), lack of specialized technical consultants (Q13), and insufficient prerequisite programs (Q11), where the ChatGPT exceeded human ratings by more than 0.45 points. Moderate positive differences were also recorded for limited technical expertise of staff (Q8) and inability to ensure continuous training (Q12). In contrast, time required for documentation and record keeping (Q9) was the only barrier item for which the LLM assigned a lower score than human participants, suggesting a divergence in the perceived relative importance of administrative burden. Overall, the results indicate that while both humans and the ChatGPT 4.1 identify similar barriers, the model tends to amplify the perceived severity of resource- and infrastructure-related challenges.
In the Benefits of HACCP section, human and LLM assessments demonstrated a high degree of convergence, with all absolute differences remaining small. Human respondents marginally rated legislative compliance (Q14), reduction of customer complaints (Q15), and customer satisfaction and trust (Q17) higher than the ChatGPT 4.1, whereas the model slightly exceeded human scores only for the promotion of food safety culture among employees (Q18). The remaining item, enhancement of business reputation and credibility (Q16), showed near-identical ratings between the two sources. These results indicate strong alignment between human expertise and LLM-generated evaluations regarding the perceived benefits of HACCP implementation, with no item exhibiting a substantial divergence.
In the AI and Digitalization section, the tested model consistently assigned higher scores than human respondents across all AI-related functional items, reflecting a clear directional difference in assessment. The most pronounced discrepancy was observed for the role of ChatGPT 4.1 in supporting HACCP decision-making (Q21), followed by AI-based automatic monitoring of CCPs and PRPs (Q20) and the contribution of AI to staff HACCP training (Q19). These differences suggest that the LLM attributes greater perceived effectiveness to advanced digital and AI-driven applications than human professionals. Notably, willingness to participate in AI-based HACCP training programs (Q22) showed complete agreement between human and LLM scores, indicating convergence at the attitudinal level despite divergence in functional assessments.
Figure 5 presents a graphical comparison between human and LLM (ChatGPT 4.1) mean scores, including variability across professional roles, allowing visual assessment of agreement patterns and response dispersion.
Figure 6 represents mean Likert scores (1–5) reported by human participants for each HACCP dimension, while points indicate the corresponding AI-derived benchmark values used as test references in the one-sample t-tests. The ChatGPT 4.1 output was treated as a fixed reference value rather than a sampled distribution; thus, variance was derived solely from human responses. p-values denote statistically significant deviations between human responses and the AI benchmark.
Figure 6 summarizes the results of the one-sample t-tests comparing human questionnaire responses with AI-derived benchmark values across the three HACCP dimensions. Statistically significant deviations were observed for the Barriers to HACCP (Section B) and AI Readiness (Section D) dimensions, where human mean scores were significantly lower than the LLM reference values (p < 0.001 for both sections). In contrast, for the Benefits of HACCP dimension (Section C), human responses were slightly but significantly higher than the AI benchmark (p = 0.011), although the absolute difference in mean scores was minimal.

4. Discussion

Several studies conducted in Southern Europe and Greece have consistently reported persistent gaps between formal HACCP adoption and its effective operational implementation [34,35,36,37]. The findings of the present study reinforce this body of evidence, as the elevated scores observed in the Barriers section (Average: 4.18) are consistent with previously documented infrastructural deficiencies, limited technical expertise, and organizational constraints in small and medium-sized hospitality establishments [38,39,40]. In hotel environments in particular, the structural complexity of food production systems, combined with seasonal workload fluctuations, further challenges the stability of HACCP monitoring and verification procedures [41,42,43,44,45]. Within the present Greek sample, ChatGPT 4.1-derived ratings were higher than the corresponding human ratings for several HACCP implementation barriers.
A relatively strong positive association was observed between high staff turnover and difficulty in ensuring continuous training (r = 0.62), indicating that these barriers tend to co-occur. Given the cross-sectional design, however, the direction of this relationship cannot be established. Therefore, the finding should not be interpreted as evidence that turnover disrupts training or that insufficient training increases turnover. Correlation analysis of human responses revealed a strong association between high staff turnover and the inability to ensure continuous training (r = 0.62, p < 0.01), empirically substantiating earlier qualitative observations that workforce instability undermines long-term food safety capacity building [41,46]. In parallel, the significant relationship observed between lack of staff motivation and insufficient management support (r = 0.54, p < 0.01) highlights the central role of leadership commitment in shaping food safety culture, in line with prior research emphasising management-driven compliance dynamics [11].
The importance of leadership commitment is further illustrated by differences in how organizational roles perceive HACCP implementation challenges and emerging digital tools. Previous research has shown that personnel occupying different positions within HACCP-based food safety management systems may prioritize barriers differently, with frontline or food safety team members placing greater emphasis on production burdens, employee-related constraints, management attributes, and infrastructure, whereas leadership roles may prioritize different system-level concerns [47,48]. Indeed, food safety team leaders and team members have been shown to develop distinct perceptions of barriers even within the same food safety management system [47]. In the present study, this role-dependent pattern may also help explain the comparatively lower rating assigned by Production Staff to LLM-supported HACCP decision-making (Q21 = 3.50). Rather than indicating generalized scepticism toward AI, this finding may reflect differences in role-specific exposure, perceived usefulness, workflow compatibility, and the immediate operational relevance of AI-supported decision-making. This interpretation is consistent with evidence showing that food safety compliance is shaped by interacting organizational and human factors, including staff resistance, insufficient knowledge and training, managerial support, supervision, and resource limitations [48,49]. Similar role-dependent differences have also been observed in other standardized quality-management settings, where personnel closer to operational activities may perceive documentation and procedural requirements differently from senior management [50]. Importantly, factors such as workload, psychological safety, digital familiarity, organizational support, and trust in AI were not directly measured in the present study. They should therefore be regarded as plausible explanatory mechanisms rather than demonstrated mediators of the observed differences and warrant direct investigation in future studies [38,47,48,49,50].
Multiple studies indicate that documentation burden is consistently experienced more negatively by operational-level staff than by management, primarily due to its direct interference with core professional activities and time spent away from frontline tasks [51,52]. Frontline personnel frequently describe documentation either as an essential but time-consuming obligation or as a perceived administrative burden disconnected from practical risk prevention, whereas senior managers tend to emphasise its role in accountability, compliance, and quality assurance. Although HACCP-specific studies addressing documentation perceptions remain limited, record keeping is explicitly identified as a central component of HACCP systems, serving to demonstrate adherence and control verification [7,53].
Studies examining HACCP compliance consistently report that financial limitations, lack of training, and insufficient managerial support operate as interdependent barriers rather than competing ones [54,55,56]. Cost-related pressures may restrict investments in training, digital tools, and system optimization, while human resource limitations amplify the perceived workload associated with documentation and monitoring activities. This interdependence supports the interpretation that documentation intensity and financial constraints jointly shape practitioners’ experiences of HACCP implementation, particularly in resource-constrained hospitality settings.
Multiple studies highlight that cost-related pressures frequently result in partial or reactive implementation of food safety systems, rather than proactive risk prevention strategies [57,58]. These constraints tend to intensify during periods of economic stress, amplifying the trade-offs between regulatory compliance and operational viability.
The moderate but significant correlation between documentation burden and financial costs (r = 0.41, p < 0.05) suggests that administrative workload and economic constraints are experienced as mutually reinforcing pressures. These interdependencies appear to have been intensified during the COVID-19 pandemic, which increased operational costs and documentation demands through stricter hygiene and monitoring requirements [17,19]. At the same time, the strong alignment between legislative compliance and business reputation (r = 0.51, p < 0.01) indicates that regulatory adherence is framed by practitioners not merely as a legal obligation, but as a strategic investment in organizational credibility and trust [59,60,61,62].
The comparative analysis between human responses and the LLM-derived benchmark further extends these findings by revealing systematic patterns of convergence and divergence. While the one-sample t-test demonstrated close alignment between the ChatGPT 4.1 and human professionals regarding perceived HACCP benefits (p = 0.011), statistically significant deviations were observed in both implementation barriers (p < 0.001) and digital readiness (p < 0.001). In particular, the LLM baseline for AI readiness (4.39) exceeded the corresponding human mean score (4.12), suggesting a tendency toward higher assessments of the practical feasibility of digital HACCP adoption under the present experimental conditions. This divergence likely reflects limited sensitivity to contextual constraints such as technological lag, resource limitations, and continued reliance on manual record keeping practices prevalent in Southern European hospitality settings [18].
This technological lag limits real-time monitoring, increases the likelihood of documentation errors, and constrains the scalability of food safety management systems in complex operational environments. Importantly, statistical significance should not be interpreted as equivalent to practical significance. Although statistically significant Human–ChatGPT differences were identified, the observed absolute score differences were generally modest and should therefore be interpreted cautiously, as they indicate systematic perceptual divergence rather than necessarily substantial disagreement in everyday HACCP practice. In the context of developing and operating an HACCP system, modest overestimations of digital readiness or AI-driven mitigation of barriers could lead to operational overestimation if not filtered through human expertise and insight [63,64,65,66,67]. The practical value of this comparison lies primarily in identifying domains in which ChatGPT-generated assessments converge with, or systematically diverge from, professional perceptions, thereby informing the design of human-supervised AI tools for HACCP training, documentation support, and preliminary decision support. These findings should not be interpreted as evidence that LLMs can independently perform operational HACCP tasks. Validation through scenario-based hazard analysis, CCP identification, critical-limit setting, and corrective-action assessment is required before such applications can be considered.
Analysis of the ChatGPT 4.1-generated justifications indicates that model responses predominantly reflect standardized regulatory frameworks and generalized best-practice assumptions, whereas human participants consistently emphasise operational feasibility, contextual adaptation, and the necessity of human-in-the-loop oversight. This contrast highlights the complementary nature of human experiential knowledge and AI-generated reasoning, particularly in safety-critical decision-making contexts. Importantly, the results also point toward a pragmatic pathway for digital transition. The strong association between perceived AI usefulness for training and willingness to participate in AI-based programs (r = 0.68, p < 0.01) positions training-oriented AI applications as a credible and socially acceptable entry point for digital transformation within HACCP systems. The observed differences across roles suggest that perceptions of HACCP barriers and digital readiness vary between operational and regulatory contexts. The LLM appears to align more closely with regulatory/theoretical frameworks, which may partly explain divergences from operational, frontline professional responses [68,69,70,71,72].
In summary, several structural divergences emerge when viewed from a broader perspective. The discrepancies between human judgement and AI-generated reasoning underscore the inherent limitations of relying solely on algorithmic outputs. While training Large Language Models on global datasets, technical manuals, and regulatory frameworks provides extensive generalized knowledge, these models often fail to detect the local nuances and operational realities that shape workforce behaviour in specific sectors, such as hospitality. In the present study, this is clearly evidenced by the divergent perceptions regarding barriers to HACCP implementation.
A risk of algorithmic bias appears to be present, favouring a ‘checklist and compliance’ orientation rather than a proactive food safety culture. Consequently, professional judgement must remain the primary authority in governance and ethical decision-making, with AI serving in a subsidiary, supportive role. As AI tools become increasingly integrated into industrial safety systems, their role should be clearly defined: primarily assisting in training and documentation validation while the human factor remains central to operational control, ethical responsibility and the contextual interpretation of risks [73,74,75]. Overall, the effective integration of AI into HACCP frameworks requires continuous validation against empirically grounded human perspectives, strong governance mechanisms, and leadership commitment to ensure that digital tools enhance—rather than replace—professional judgement in food safety management systems [76,77,78,79].
From a practical perspective, the present findings support a clearly bounded, human-supervised role for ChatGPT-4.1 in HACCP-related applications. Appropriate use cases may include staff training, synthesis of regulatory or technical information, documentation support, and preliminary decision support. Conversely, the tested model should not replace qualified professionals in safety-critical functions such as site-specific hazard identification, CCP determination, establishment of critical limits, interpretation of monitoring deviations, or selection of corrective actions. Any LLM-generated recommendation should be independently verified against applicable regulatory requirements, validated HACCP plans, and facility-specific operational conditions before implementation. This is particularly important because inaccurate, incomplete, or contextually inappropriate recommendations could result in ineffective control measures and potentially compromise food safety. Accordingly, human oversight and clearly defined accountability should remain integral to any LLM-assisted HACCP workflow. Future research should move beyond perception-based benchmarking toward controlled, scenario-based validation using real or realistically simulated HACCP decision cases with predefined expert or regulatory reference standards.

5. Conclusions and Limitations

This exploratory study acknowledges certain limitations, primarily the relatively small sample size (n = 90), purposive sampling strategy, and geographic specificity of the human dataset. The geographical distribution of participants was concentrated primarily in Attica and Central Macedonia, limiting the external validity of the findings. Accordingly, the results should not be directly extrapolated to other countries, cultural or regulatory contexts, or to other food-sector environments such as food manufacturing, retail, and healthcare food service, where HACCP implementation conditions may differ substantially. The observed Human–LLM differences should therefore be interpreted as context-specific findings rather than as universally generalisable patterns. Future studies should validate these findings using larger and more geographically diverse samples across different food industry sectors and regulatory environments. Accordingly, the tendency of the tested model to assign higher ratings to several HACCP barriers should be interpreted as a context-specific comparative finding rather than evidence of systematic LLM overestimation across populations or food service environments. Importantly, the observed Human–ChatGPT alignment refers exclusively to similarity in perceived ratings across the examined HACCP dimensions and should not be interpreted as evidence of equivalent professional judgement, HACCP decision quality, hazard identification accuracy, or effectiveness in real-world food safety practice, none of which were directly evaluated in the present study.
Furthermore, as LLM technology evolves rapidly, these results represent a specific technological snapshot. Thus, issues of output stability, inherent bias, and reproducibility remain critical open questions. Given the tendency of the ChatGPT 4.1 to overestimate implementation barriers and digital readiness relative to human responses and the identified “training–turnover association” in the Greek hospitality sector, ChatGPT 4.1 should be viewed as supportive tool rather than an autonomous risk assessment system. While it may offer potential for training, regulatory information synthesis, and documentation support, it cannot replace professional judgement and requires rigorous human-in-the-loop oversight to ensure safety-critical reliability. The benchmark-based statistical comparison does not explicitly model uncertainty across individual LLM generations. Future studies could incorporate run-level LLM variability using bootstrap, permutation, or hierarchical modelling approaches.
An additional limitation is that the computational comparison was conducted exclusively with ChatGPT 4.1. Consequently, the observed Human–AI convergence and divergence patterns are specific to this model and its configuration at the time of inference and should not be generalized to other LLMs. Future research should conduct cross-model validation using multiple LLM architectures to determine whether the observed patterns are model-specific or reproducible across different systems.
A final limitation is the inherent ‘contextual gap’ between the datasets. The tested model is trained on vast global repositories that represent idealized standards, whereas the human sample reflects the specific socio-economic conditions of the Greek hospitality sector. This may explain discrepancies where the AI aligns with theoretical best practices while human professionals prioritize practical, context-dependent challenges. Future research should explore multi-regional samples to further address this geographic and demographic variability.
Based on the findings of this study, several actionable recommendations can be proposed for stakeholders. For industry practitioners, LLMs should be utilized as documentation aids rather than autonomous decision-makers, employing a ‘human-in-the-loop’ approach where AI-generated drafts are rigorously validated by experienced professionals. Regulatory bodies should consider developing frameworks that emphasise the necessity for AI-assisted safety plans to explicitly account for local operational constraints. For researchers, the focus should shift toward developing context-aware prompting techniques tailored to specific operational environments.
Future studies should systematically compare zero-shot, context-enriched, few-shot, and retrieval-augmented generation (RAG) approaches to determine whether grounding LLMs in validated local regulatory, organizational, and sector-specific knowledge improves alignment with professional HACCP assessments.
Such efforts will facilitate the development of transparent, scalable digital tools capable of complementing professional expertise. Ultimately, the transition toward digital HACCP systems must prioritize educational value and leadership commitment, leveraging AI as a strategic entry point to foster a resilient food safety culture in resource-constrained environments.

Supplementary Materials

The following supporting information can be downloaded at https://www.mdpi.com/article/10.3390/foods15162916/s1, File S1: HACCP Implementation Questionnaire Greek-English version.

Author Contributions

Conceptualisation, D.M.K. and C.S.; methodology, C.S. and C.V.; software, E.S. and A.S.; investigation, C.T. and V.P.; data curation, C.S.; writing—original draft preparation, D.M.K. and C.S.; writing—review and editing, D.M.K.; visualization, C.S. and C.V.; supervision, C.S.; project administration, C.T. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

The Institutional Review Board Statement to explicitly state that the Research Ethics and Deontology Committee of Democritus University of Thrace confirmed that, under the applicable institutional guidelines and national regulations, this type of anonymous, non-interventional professional survey did not require formal ethical review or prior ethics committee approval.

Informed Consent Statement

Informed consent was obtained from all participants involved in the study. Participation was voluntary and anonymous, and completion of the questionnaire constituted consent. Written informed consent has been obtained from the participants to publish this paper.

Data Availability Statement

The original contributions presented in this study are included in the article/supplementary material. Further inquiries can be directed to the corresponding author.

Acknowledgments

During the preparation of this work, the authors used GPT-v.5pro (OpenAI) in order to improve the readability and language of the manuscript. The authors have reviewed and edited the output and take full responsibility for the content of this publication. This work was supported by the master’s programme in “Food, Nutrition and Microbiome” of the Medical School, Democritus University of Thrace, Greece.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Aslam, M.U.; Aslam, E.; Shahbaz, M.; Aslam, M.U.; Shahbaz, M. Emerging trends in food safety and quality management. Insights J. Health Rehabil. 2025, 3, 851–859. [Google Scholar] [CrossRef] [Scilit]
  2. Błaszczyk, I. The management of food safety in beverage industry. In Safety Issues in Beverage Production; Grumezescu, A.M., Holban, A.M., Eds.; Academic Press: London, UK, 2020; pp. 1–38. [Google Scholar] [CrossRef] [Scilit]
  3. Hyde, R.; Hoflund, A.B.; Pautz, M. One HACCP, two approaches: Experiences with HACCP food safety management systems in the United States and the EU. Adm. Soc. 2016, 48, 962–987. [Google Scholar] [CrossRef] [Scilit]
  4. Zarid, M. The green HACCP approach: Advancing food safety and sustainability. Sustainability 2025, 17, 7834. [Google Scholar] [CrossRef] [Scilit]
  5. Sariq, M. The effectiveness of HACCP and FSMS in enhancing food safety in meat industry. Int. J. Res. Appl. Sci. Eng. Technol. 2025, 13, 2945–2952. [Google Scholar] [CrossRef] [Scilit]
  6. Uzoigwe, D.; Kongolo, D. Integration of hazard analysis and critical control points with maintenance practices: Enhancing food safety in the food and beverage industry. Int. J. Latest Technol. Eng. Manag. Appl. Sci. 2024, 13, 88–101. [Google Scholar] [CrossRef] [Scilit]
  7. Radu, E.; Dima, A.; Dobrota, E.M.; Badea, A.M.; Madsen, D.Ø.; Dobrin, C.; Stanciu, S. Global trends and research hotspots on HACCP and modern quality management systems in the food industry. Heliyon 2023, 9, e18232. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  8. Dervilly, G.; Besselink, H.; Bover, S.; Hou, J.; Rantsiou, K.; Yue, M.; Zwietering, M.H.; Engel, E. The SAFFI project: Fostering alignment and collaboration in EU-China food safety management. Food Res. Int. 2025, 213, 116600. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. Osimani, A.; Milanović, V.; Aquilanti, L.; Polverigiani, S.; Garofalo, C.; Clementi, F. Hygiene auditing in mass catering. Public Health 2018, 159, 17–20. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  10. Cenci-Goga, B.T.; Ortenzi, R.; Bartocci, E.; Codega de Oliveira, A.; Clementi, F.; Vizzani, A. Implementation of HACCP and microbiological quality of meals. Foodborne Pathog. Dis. 2005, 2, 138–145. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  11. Worsfold, D.; Worsfold, P. Increasing HACCP awareness. J. R. Soc. Promot. Health 2005, 125, 129–135. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  12. Jevšnik, M.; Hlebec, V.; Raspor, P. Barriers identification during HACCP implementation. Acta Aliment. 2006, 35, 319–353. [Google Scholar] [CrossRef] [Scilit]
  13. Casolani, N.; Signore, A. Factors influencing HACCP applications in HoReCa sector. Br. Food J. 2016, 118, 1195–1207. [Google Scholar] [CrossRef] [Scilit]
  14. Fletcher, S.; Maharaj, S.; James, K. Food safety systems in hotels. J. Travel Med. 2009, 16, 35–41. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  15. FAO. Hazard Analysis and Critical Control Point System. Available online: https://www.fao.org/3/w8088e/w8088e05.htm (accessed on 20 November 2025).
  16. Moreira, A.; Léon, M.; Coda Moscarola, F.; Roumpakis, A. In the eye of the storm… again! Social policy responses to COVID-19 in Southern Europe. Soc. Policy Adm. 2021, 55, 339–357. [Google Scholar] [CrossRef] [Scilit]
  17. Vagionaki, T. Linking compliance and policy learning. The case of EU soft law in Greece and Spain. Int. Rev. Public Policy 2022, 4, 219–240. [Google Scholar] [CrossRef] [Scilit]
  18. McClements, D.J.; Barrangou, R.; Hill, C.; Kokini, J.L.; Lila, M.A.; Meyer, A.S.; Yu, L.L. Building a resilient, sustainable, and healthier food supply through innovation and technology. Annu. Rev. Food Sci. Technol. 2020, 12, 1–28. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  19. Galanakis, C.M.; Rizou, M.; Aldawoud, T.M.; Ucak, I.; Rowan, N.J. Innovations and technology disruptions in the food sector within the COVID-19 pandemic and post-lockdown era. Trends Food Sci. Technol. 2021, 110, 193–200. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  20. Laurence, T.; Harris, J.; Loman, L.; Douglas, A.; Chan, Y.; Hounsome, L.; Larkin, L.; Borowitz, M. Review GIDE—Restaurant review gastrointestinal illness detection and extraction with large language models. arXiv 2025, arXiv:2503.09743. [Google Scholar]
  21. Salekpay, F.; van den Bergh, J.; Savin, I. Comparing advice on climate policy between academic experts and ChatGPT. Ecol. Econ. 2024, 226, 108352. [Google Scholar] [CrossRef] [Scilit]
  22. Charalampidou, S.; Zeleskidis, A.; Dokas, I.M. Hazard analysis in the era of AI: Assessing the usefulness of ChatGPT4 in STPA hazard analysis. Saf. Sci. 2024, 178, 106608. [Google Scholar] [CrossRef] [Scilit]
  23. Chen, Y.; Yin, R.; Guo, L.; Zhao, D.; Sun, B. Consumer sensory evaluation scale for pale lager beer. Foods 2025, 14, 2834. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  24. Shen, C.; Meng, W.; Chen, X.; Liu, K.; Wu, X.; Yu, Q. Consumers’ perception of food safety risks. Foods 2025, 14, 3463. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  25. Yuan, X.; Chen, Y.; Yin, R.; Guo, L.; Song, Y.; Zhong, B.; Zhao, D. Relationship between personal characteristics and alcohol consumption. Foods 2025, 14, 3536. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  26. Pînzariu, S.; Pînzariu, A. HACCP standards in military feeding services. Knowl. Based Organ. 2025, 31, 166–170. [Google Scholar] [CrossRef] [Scilit]
  27. Xia, T.; Shen, X.; Li, L. AI food and consumer trust. Foods 2024, 13, 2983. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  28. Wang, W.; Chen, Z.; Kuang, J. Artificial Intelligence-Driven Recommendations and Functional Food Purchases: Understanding Consumer Decision-Making. Foods 2025, 14, 976. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  29. Górska, P.; Górna, I.; Miechowicz, I.; Przysławski, J. Eating behaviour during COVID-19 pandemic. Foods 2021, 10, 1624. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  30. Alzhrani, W.F.; Shatwan, I.M. Food safety knowledge of restaurant handlers. Foods 2024, 13, 2176. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  31. Hertzog, M.A. Considerations in determining sample size for pilot studies. Res. Nurs. Health 2008, 31, 180–191. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  32. Willis, G.B. Cognitive Interviewing. In Sage Research Methods; SAGE Publications, Inc.: Thousand Oaks, CA, USA, 2005. [Google Scholar] [CrossRef] [Scilit]
  33. Pourhoseingholi, M.A.; Vahedi, M.; Rahimzadeh, M. Sample size calculation in medical studies. Gastroenterol. Hepatol. Bed Bench 2013, 6, 14–17. [Google Scholar] [PubMed]
  34. Awuchi, C.G. HACCP, quality, and food safety management in food and agricultural systems. Cogent Food Agric. 2023, 9, 2176280. [Google Scholar] [CrossRef] [Scilit]
  35. Lee, J.C.; Darabă, A.; Voidarou, C.; Rozos, G.; El Enshasy, H.A.; Varzakas, T. Implementation of food safety management systems along with other management tools (HAZOP, FMEA, Ishikawa, Pareto): The case study of Listeria monocytogenes and correlation with microbiological criteria. Foods 2021, 10, 2169. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  36. Mutua, A. Role of food management systems on food safety in hotels. J. Food Sci. 2021, 2, 37–50. [Google Scholar] [CrossRef] [Scilit]
  37. Salavelis, A.; Pavlovsky, S.; Lazarenko, N. Obstacles to the implementation of HACCP in small food industry enterprises and in restaurant business establishments. Sci. Messenger LNU Vet. Med. Biotechnol. 2025, 27, 103. [Google Scholar] [CrossRef] [Scilit]
  38. Jevšnik, M.; Raspor, P. Food safety knowledge and behaviour among food handlers in catering establishments: A case study. Br. Food J. 2021. Epub ahead of printing. [Google Scholar] [CrossRef] [Scilit]
  39. Arvanitoyannis, I.; Samourelis, K.; Kotsanopoulos, K.V. A critical analysis of ISO audits results. Br. Food J. 2016, 118, 2126–2139. [Google Scholar] [CrossRef] [Scilit]
  40. Gkrintzali, G.; Pexara, E.; Carayanni, V.; Boskou, G. Consumer protection and food safety in Greece. J. Hell. Vet. Med. Soc. 2018, 69, 965–972. [Google Scholar] [CrossRef] [Scilit]
  41. Rizzo, C.E.; Venuto, R.; Genovese, G.; Squeri, R.; Genovese, C. Food hygiene non-compliance assessment. Foods 2025, 14, 3364. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  42. Agkoli, P. SDGs strategies in Greek hospitality SMEs: Overcoming barriers to sustainability implementations. HAPSc Policy Briefs Ser. 2024, 5, 132–138. [Google Scholar] [CrossRef] [Scilit]
  43. Chatzimpyrou, O.; Chaidoutis, E.; Keramydas, D.; Papalexis, P.; Thomaidis, N.S.; Pitiriga, V.C.; Langi, P.; Koutsiari, F.; Drikos, L.; Giannari, M.; et al. Health inspections of restaurants in Greece. J. Food Prot. 2025, 88, 100452. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  44. Mellou, K.; Mplougoura, A.; Mandilara, G.; Papadakis, A.; Chochlakis, D.; Psaroulaki, A.; Mavridou, A. Swimming pool regulations in the COVID-19 era. Assessing acceptability and compliance in Greek hotels in two consecutive summer touristic periods. Water 2022, 14, 796. [Google Scholar] [CrossRef] [Scilit]
  45. Vaithinathan, A.G.; Anwar, N.; Sulieman, A.; AlAlawi, F.; Mohammed, M.Y.; Ahmed, S. Coronavirus disease and food safety in hospitality sector. In Handbook of Research on the Impacts of COVID-19 on Tourism; IGI Global: Hershey, PA, USA, 2021; pp. 603–626. [Google Scholar] [CrossRef] [Scilit]
  46. Varotsis, N. Quality standards in hospitality industry. J. Hosp. Tour. Manag. 2019, 8, 417. [Google Scholar]
  47. Osman, N.E.; Abdallah, M.A. Difficulties and barriers for the implementing of HACCP and food safety systems in food businesses in Khartoum-Sudan. Total Qual. Manag. 2018, 19, 73–79. [Google Scholar]
  48. Marule, L.; Du Rand, G.; Marx-Pienaar, N. Gauteng’s managers’ implementation of food safety protocols and practices in their QSR environments. J. Food Consum. Sci. 2024, 1, 172–188. [Google Scholar] [CrossRef] [Scilit]
  49. Psomatakis, M.; Papadimitriou, K.; Souliotis, A.; Drosinos, E.H.; Papadopoulos, G. Food Safety and Management System Audits in Food Retail Chain Stores in Greece. Foods 2024, 13, 457. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  50. Dzwolak, W.; Anim, B. Barriers hindering maintenance of standardised HACCP-based food safety management systems in small Polish food businesses. Food Control 2025, 168, 110849. [Google Scholar] [CrossRef] [Scilit]
  51. Nicolaisen, A.; Bogh, S.B.; Churruca, K.; Ellis, L.A.; Braithwaite, J.; von Plessen, C. Managers’ perceptions of the effects of a national mandatory accreditation program in Danish hospitals: A cross-sectional survey. Int. J. Qual. Health Care 2019, 31, 331–337. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  52. de Groot, K.; de Veer, A.J.E.; Munster, A.M.; Francke, A.L.; Paans, W. Nursing documentation and its relationship with perceived nursing workload: A mixed-methods study among community nurses. BMC Nurs. 2022, 21, 34. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  53. Weinroth, M.D.; Belk, A.D.; Belk, K.E. History, development, and current status of food safety systems worldwide. Anim. Front. 2018, 8, 9–15. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  54. Noor Hasnan, N.Z.; Kadir, R.; Mohd Amin, N.A.; Aziz, N.; Mohd Ramli, S.H. Analysis of the most frequent nonconformance aspects related to GMP among SMEs in the food industry and their main factors. Food Control 2022, 141, 109205. [Google Scholar] [CrossRef] [Scilit]
  55. Galstyan, S.; Harutyunyan, T. Barriers and facilitators of HACCP adoption in the Armenian dairy industry. Br. Food J. 2016, 118, 2676–2691. [Google Scholar] [CrossRef] [Scilit]
  56. Borovčanin, D.; Kilibarda, N. Assuring good food handling practices in hospitality: Financial costs and employees’ attitudes—A case study from Serbia. Meat Technol. 2020, 61, 82–94. [Google Scholar] [CrossRef] [Scilit]
  57. Fotopoulos, C.; Kafetzopoulos, D.; Psomas, E.L. Assessing the critical factors and their impact on the effective implementation of a food safety management system. Int. J. Qual. Reliab. Manag. 2009, 26, 894–910. [Google Scholar] [CrossRef] [Scilit]
  58. Semos, A.; Kontogeorgos, A. HACCP implementation in Northern Greece. Br. Food J. 2007, 109, 5–19. [Google Scholar] [CrossRef] [Scilit]
  59. Bertella, G. Rethinking sustainability and food in tourism. Ann. Tour. Res. 2020, 84, 103005. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  60. Molina-Collado, A.; Santos-Vijande, M.L.; Gómez-Rico, M.; Madera, J.M. Sustainability in hospitality and tourism. Int. J. Contemp. Hosp. Manag. 2022, 34, 3029–3064. [Google Scholar] [CrossRef] [Scilit]
  61. Ruiz-Molina, M.E.; Belda-Miquel, S.; Hytti, A.; Gil-Saura, I. Addressing sustainable food management in hotels. Br. Food J. 2022, 124, 462–492. [Google Scholar] [CrossRef] [Scilit]
  62. Zhylenko, K.; Yarovenko, T.; Stavytska, A.; Samoilenko, A. Sustainable development of hotel and restaurant business. Econ. Financ. Law 2024, 5, 64–67. [Google Scholar] [CrossRef] [Scilit]
  63. Ibrahim, A.Z.; Megahed, M.; Farida, A.; Tamer, A. Investigating the effect of food safety practices on hotel performance. Int. J. Tour. Hosp. Manag. 2021, 4, 243–264. [Google Scholar] [CrossRef] [Scilit]
  64. Peistikou, M. Restaurants industry in the COVID-19 era: Challenge or opportunity? In Strategic Innovative Marketing and Tourism in the COVID-19 Era; Springer Proceedings in Business and Economics; Kavoura, A., Havlovic, S.J., Totskaya, N., Eds.; Springer: Berlin/Heidelberg, Germany, 2021; pp. 153–162. [Google Scholar]
  65. Hassani, S. Enhancing legal compliance and regulation analysis with large language models. In Proceedings of the 2024 IEEE 32nd International Requirements Engineering Conference (RE), Reykjavik, Iceland, 24–28 June 2024; pp. 507–511. [Google Scholar]
  66. Hassani, S.; Sabetzadeh, M.; Amyot, D. An empirical study on LLM-based classification of requirements-related provisions in food-safety regulations. Empir. Softw. Eng. 2025, 30, 3. [Google Scholar] [CrossRef] [Scilit]
  67. Görgen, L.; Müller, E.; Triller, M.; Nast, B.; Sandkuhl, K. Large language models in enterprise modeling: Case study and experiences. In Proceedings of the 12th International Conference on Model-Based Software and Systems Engineering (MODELSWARD 2024), Rome, Italy, 21–23 February 2024. [Google Scholar]
  68. Collier, Z.A.; Gruss, R.; Abrahams, A.S. How good are large language models at product risk assessment? Risk Anal. 2025, 45, 766–789. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  69. Esposito, M.; Palagiano, F.; Lenarduzzi, V. Beyond words: On large language models actionability in mission-critical risk analysis. In Proceedings of the 18th ACM/IEEE International Symposium on Empirical Software Engineering and Measurement, Catalunya, Spain, 20–25 October 2024. [Google Scholar]
  70. Huang, Y.; Song, J.; Wang, Z.; Chen, H.; Ma, L. Look before you leap: An exploratory study of uncertainty measurement for large language models. arXiv 2023, arXiv:2307.10236. [Google Scholar]
  71. Kattamreddy, A.R.; Chinnam, H. The future of large language models in toxicological risk assessment: Opportunities and challenges. Public Health Toxicol. 2025, 5, 3. [Google Scholar] [CrossRef] [Scilit]
  72. Ma, P.; Tsai, S.; He, Y.; Jia, X.; Zhen, D.; Yu, N.; Wang, Q.; Ahuja, J.K.; Wei, C.-I. Large language models in food science: Innovations, applications, and future. Trends Food Sci. Technol. 2024, 148, 104488. [Google Scholar] [CrossRef] [Scilit]
  73. Ahmed, N.; Kour, R.; Jan, T.; Sharma, S.; Singh, T.P.; Chauhan, P.; Ghanghas, S.; Sheikh, I.; Rafatullah, M.; Setyawan, H.Y.; et al. The intersection of artificial intelligence and food systems: Exploring technological breakthroughs and data-driven agriculture. Cogent Food Agric. 2026, 12, 2615165. [Google Scholar] [CrossRef] [Scilit]
  74. Abdi, Y.H.; Bashir, S.G.; Abdullahi, Y.B.; Abdi, M.S.; Ahmed, N.I. Artificial intelligence applications for strengthening global food safety systems. Discov. Food 2025, 5, 391. [Google Scholar] [CrossRef] [Scilit]
  75. Harikrishnan, S.; Kaushik, D.; Rasane, P.; Kumar, A.; Kaur, N.; Reddy, C.K.; Proestos, C.; Oz, F.; Kumar, M. Artificial intelligence in sustainable food design: Technological, ethical consideration, and future. Trends Food Sci. Technol. 2025, 163, 105152. [Google Scholar] [CrossRef] [Scilit]
  76. Dokas, I. From hallucinations to hazards: Benchmarking LLMs for hazard analysis in safety-critical systems. Saf. Sci. 2025, 194, 107056. [Google Scholar] [CrossRef] [Scilit]
  77. Fan, L.; Li, L.; Ma, Z.; Lee, S.; Yu, H.; Hemphill, L. A bibliometric review of large language models research from 2017 to 2023. ACM Trans. Intell. Syst. Technol. 2023, 15, 1–25. [Google Scholar] [CrossRef] [Scilit]
  78. Yu, H.; Fan, L.; Li, L.; Zhou, J.; Ma, Z.; Xian, L.; Hua, W.; He, S.; Jin, M.; Zhang, Y.; et al. Large language models in biomedical and health informatics: A review with bibliometric analysis. J. Healthc. Inform. Res. 2024, 8, 658–711. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  79. Sindhu, B.; Prathamesh, R.P.; Sameera, M.B.; Kumaraswamy, S. The evolution of large language model: Models, applications and challenges. In Proceedings of the 2024 International Conference on Current Trends in Advanced Computing (ICCTAC), Bengaluru, India, 8–9 May 2024; pp. 1–8. [Google Scholar]
Figure 1. Flow diagram illustrating the sequential implementation of the seven HACCP principles in food safety management [15].
Figure 1. Flow diagram illustrating the sequential implementation of the seven HACCP principles in food safety management [15].
Foods 15 02916 g001
Figure 2. Flowchart of the study design, questionnaire development, LLM simulation protocol, and comparative analysis workflow.
Figure 2. Flowchart of the study design, questionnaire development, LLM simulation protocol, and comparative analysis workflow.
Foods 15 02916 g002
Figure 3. Internal consistency reliability of the questionnaire.
Figure 3. Internal consistency reliability of the questionnaire.
Foods 15 02916 g003
Figure 4. Statistically significant correlations between selected items of the human participant HACCP questionnaire (n = 90, p < 0.05).
Figure 4. Statistically significant correlations between selected items of the human participant HACCP questionnaire (n = 90, p < 0.05).
Foods 15 02916 g004
Figure 5. Comparison of human and LLM (ChatGPT 4.1) responses across HACCP questionnaire items (Q5–Q22). Bars represent mean Likert scores for human participants and LLM-generated responses. Error bars illustrate variability across professional roles.
Figure 5. Comparison of human and LLM (ChatGPT 4.1) responses across HACCP questionnaire items (Q5–Q22). Bars represent mean Likert scores for human participants and LLM-generated responses. Error bars illustrate variability across professional roles.
Foods 15 02916 g005
Figure 6. One-sample t-test results comparing human questionnaire responses (n = 90) against AI-derived benchmark values across HACCP dimensions (Mean Likert score, 1–5; p < 0.05).
Figure 6. One-sample t-test results comparing human questionnaire responses (n = 90) against AI-derived benchmark values across HACCP dimensions (Mean Likert score, 1–5; p < 0.05).
Foods 15 02916 g006
Table 1. Mean Scores of ChatGPT 4.1-Generated Responses Across Five Professional Roles (N = 30 runs per role).
Table 1. Mean Scores of ChatGPT 4.1-Generated Responses Across Five Professional Roles (N = 30 runs per role).
Question NoItem ConsultantProduction StaffManager/
Owner
AuditorResearcher/
Academic
Average
Score
SECTION B—
BARRIERS
Q5Financial constraints/Equipment renewal costs4.634.374.874.474.704.61
Q6Lack of staff motivation and commitment4.704.874.434.604.504.62
Q7High staff turnover4.804.774.704.804.774.77
Q8Limited technical expertise of staff4.534.674.334.704.434.53
Q9Time required for documentation (record keeping)3.904.833.504.073.233.91
Q10Management support and commitment4.234.434.534.504.734.48
Q11Insufficient prerequisite programs (PRPs)4.304.104.174.874.674.42
Q12Inability to ensure continuous training4.674.734.404.774.834.68
Q13Lack of specialized technical consultants4.133.873.974.234.334.11
SECTION C—
BENEFITS
Q14Improves compliance with legislation4.704.404.604.974.904.71
Q15Reduces the likelihood of customer complaints4.534.304.774.674.504.55
Q16Enhances business reputation and credibility4.874.534.904.874.834.80
Q17Increases customer satisfaction and trust4.734.474.774.634.604.64
Q18Promotes food safety culture among employees4.634.404.734.874.934.71
SECTION D—
AI/DIGITALIZATION
Q19AI can contribute to staff HACCP training4.403.974.104.574.774.36
Q20AI systems can assist with automatic monitoring of CCPs and PRPs4.704.274.634.834.804.65
Q21Large Language Models (LLMs) can support HACCP decision-making4.373.904.134.674.734.36
Q22Willingness to participate in AI-utilizing HACCP training program4.233.74.074.334.64.19
Average Score 4.504.374.424.634.604,5
Table 2. Demographic and professional characteristics of the participants (n = 90).
Table 2. Demographic and professional characteristics of the participants (n = 90).
CategoryCharacteristicPercentage (%)/
Frequency
Experience in HACCP<1 year8%
1–3 years15%
4–7 years22%
8–10 years18%
>10 years37%
Organization TypeMass Catering/Catering Units42%
Hotel Units35%
Consulting/Auditing Firms12%
Public Authorities/Regulatory Bodies11%
TrainingCertified HACCP Training82%
No Official Certification18%
Professional roleHACCP Consultant/Auditor33
Quality or Food Safety Manager32
Manager/Business Owner10
Researcher/Academic7
Production/Kitchen Staff5
Other food safety–related roles (e.g., regulatory authority staff, assistants, inspectors):3
Geographical region of
professional activity
Attica38
Central Macedonia24
Eastern Macedonia and Thrace6
Peloponnese5
Crete4
Ionian islands4
Other regions (Sterea Ellada, North Aegean, South Aegean, Thessaly, Western Greece, Epirus, Western Macedonia)9
Table 3. Human participant survey results: Mean scores per HACCP dimension and professionals (n = 90).
Table 3. Human participant survey results: Mean scores per HACCP dimension and professionals (n = 90).
Question NoItem ConsultantProduction StaffManager/
Owner
AuditorResearcher/
Academic
Average
Score
SECTION B—
BARRIERS
Q5Financial constraints/Equipment renewal costs4.153.854.34.13.954.07
Q6Lack of staff motivation and commitment4.44.24.154.54.354.32
Q7High staff turnover4.554.64.454.74.654.59
Q8Limited technical expertise of staff4.24.13.954.354.254.17
Q9Time required for documentation (record keeping)4.14.454.254.153.84.15
Q10Management support and commitment4.354.054.44.554.454.36
Q11Insufficient prerequisite programs (PRPs)3.93.753.854.24.13.96
Q12Inability to ensure continuous training4.454.34.24.64.554.42
Q13Lack of specialized technical consultants3.653.43.553.83.753.63
SECTION C—
BENEFITS
Q14Improves compliance with legislation4.854.654.754.94.954.82
Q15Reduces the likelihood of customer complaints4.64.54.74.754.654.64
Q16Enhances business reputation and credibility4.754.64.854.84.84.76
Q17Increases customer satisfaction and trust4.74.554.84.754.74.7
Q18Promotes food safety culture among employees4.54.354.454.74.84.56
SECTION D—
AI/DIGITALIZATION
Q19AI can contribute to staff HACCP training4.13.83.94.254.454.1
Q20AI systems can assist with automatic monitoring of CCPs and PRPs4.353.954.154.54.64.31
Q21Large Language Models (LLMs) can support HACCP decision-making3.853.53.74.14.33.89
Q22Willingness to participate in AI-utilizing HACCP training program4.23.94.054.34.54.19
Average Score 4.504.374.424.634.604.5
Table 4. Comparative Statistical Analysis between Human Responses and LLM (ChatGPT 4.1) Outputs.
Table 4. Comparative Statistical Analysis between Human Responses and LLM (ChatGPT 4.1) Outputs.
Question NoItem
Description
LLM
Average
Human
Average
Δ (LLM − Human)
SECTION B—
BARRIERS
Q5Financial constraints/Equipment renewal costs4.614.070.54
Q6Lack of staff motivation and commitment4.624.320.3
Q7High staff turnover4.774.590.18
Q8Limited technical expertise of staff4.534.170.36
Q9Time required for documentation (record keeping)3.914.15−0.24
Q10Management support and commitment4.484.360.12
Q11Insufficient prerequisite programs (PRPs)4.423.960.46
Q12Inability to ensure continuous training4.684.420.26
Q13Lack of specialized technical consultants4.113.630.48
Average 4.454.18
SECTION C—
BENEFITS
Q14Improves compliance with legislation4.714.82−0.11
Q15Reduces the likelihood of customer complaints4.554.64−0.09
Q16Enhances business reputation and credibility4.84.760.04
Q17Increases customer satisfaction and trust4.644.7−0.06
Q18Promotes food safety culture among employees4.714.560.15
Average 4.624.69
SECTION D—
AI/DIGITALIZATION
Q19AI can contribute to staff HACCP training4.364.10.26
Q20AI systems can assist with automatic monitoring of CCPs and PRPs4.654.310.34
Q21Large Language Models (LLMs) can support HACCP decision-making4.363.890.47
Q22Willingness to participate in AI-utilizing HACCP training program4.194.190
Average 4.394.120.19
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Konstantinidi, D.M.; Stavropoulou, E.; Stavropoulos, A.; Voidarou, C.; Tsigalou, C.; Pitiriga, V.; Stefanis, C. Can ChatGPT Reflect Professional HACCP Judgments? A Comparative Study in Hospitality Food Safety. Foods 2026, 15, 2916. https://doi.org/10.3390/foods15162916

AMA Style

Konstantinidi DM, Stavropoulou E, Stavropoulos A, Voidarou C, Tsigalou C, Pitiriga V, Stefanis C. Can ChatGPT Reflect Professional HACCP Judgments? A Comparative Study in Hospitality Food Safety. Foods. 2026; 15(16):2916. https://doi.org/10.3390/foods15162916

Chicago/Turabian Style

Konstantinidi, Despoina Maria, Elisavet Stavropoulou, Agathangelos Stavropoulos, Chrysoula (Chrysa) Voidarou, Christina Tsigalou, Vassiliki Pitiriga, and Christos Stefanis. 2026. "Can ChatGPT Reflect Professional HACCP Judgments? A Comparative Study in Hospitality Food Safety" Foods 15, no. 16: 2916. https://doi.org/10.3390/foods15162916

APA Style

Konstantinidi, D. M., Stavropoulou, E., Stavropoulos, A., Voidarou, C., Tsigalou, C., Pitiriga, V., & Stefanis, C. (2026). Can ChatGPT Reflect Professional HACCP Judgments? A Comparative Study in Hospitality Food Safety. Foods, 15(16), 2916. https://doi.org/10.3390/foods15162916

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop