Next Article in Journal
Influencing Factors of Green Smart City Development Under Government–Enterprise Cooperation: An Integrated DEMATEL-ISM-MICMAC Approach
Previous Article in Journal
Embedding Innovation in Engineering Education: An Iterative Approach to Curriculum Integration at Bachelor’s and Master’s Levels
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

An Intelligent Museum Agent Framework (IMAF): A Design Science Research Approach to Agentic AI in Museums

College of Business Administration, Hongik University, Seoul 04066, Republic of Korea
Systems 2026, 14(8), 954; https://doi.org/10.3390/systems14080954
Submission received: 2 April 2026 / Revised: 9 May 2026 / Accepted: 18 May 2026 / Published: 7 August 2026

Abstract

Although museums are increasingly integrating large language models (LLMs) into their operations, current applications remain largely confined to narrow tasks such as chatbots and document management. While the emergence of agentic AI offers more advanced, autonomous capabilities, existing frameworks are primarily designed for general-purpose environments. Consequently, they provide limited guidance for museums, which function as complex socio-technical systems governed by strict institutional mandates, ethical responsibilities, and operational constraints. Due to these complexities, the practical deployment of agentic AI in real-world museum settings remains highly challenging. At present, a system-level architecture for agentic AI tailored specifically to the unique requirements of museums is notably absent. To address this gap, this study employs the Design Science Research (DSR) paradigm to propose the Intelligent Museum Agent Framework (IMAF). The proposed framework extends an existing general agent architecture by incorporating a bifurcated module for contexts to systematically enforce institutional and operational constraints. Additionally, it features an integrated memory structure grounded in cultural heritage data to ensure factual reliability. Ultimately, the IMAF provides a conceptual systems integration guideline for museums exploring domain-constrained autonomy and offers a basis for future implementation and empirical validation.

1. Introduction

Recent advances in artificial intelligence, particularly with large language models (LLMs), have enabled the development of conversational systems and automated knowledge services across a wide range of application domains. In parallel, research on agentic AI has progressed rapidly, proposing architectures in which LLMs are combined with memory, planning, and tool use to support multi-step reasoning and task execution [1,2]. These developments have shifted attention from isolated model capabilities toward integrated systems capable of interacting with external resources and operating in complex environments.
Existing studies on agentic AI have primarily focused on general-purpose frameworks designed for broad problem-solving across domains. These frameworks typically emphasize task execution, adaptive reasoning, and tool integration in neutral or commercial environments. As a result, they tend to provide abstract models of autonomy that are largely independent of institutional settings. While such approaches are effective for generic applications, they offer limited guidance for domains in which knowledge authority, public accountability, and ethical responsibility are central concerns.
In the museum domain, research on LLM-based systems has concentrated on specific functions such as visitor guidance, collection documentation, and recommendation services. These studies demonstrate the feasibility of applying LLMs to individual tasks but do not provide an integrated agentic AI framework that reflects the institutional characteristics of museums. More broadly, although general conceptual models of agentic AI are well established, frameworks that explicitly address museum-specific operational conditions and governance requirements remain difficult to identify in the literature. This situation indicates a gap between general agentic AI architectures and the practical needs of museums as knowledge institutions.
To respond to this gap, this study adopts the Design Science Research (DSR) paradigm, which focuses on the development of artifacts through systematic design and evaluation [3]. Building on existing work in agentic AI, this research proposes the Intelligent Museum Agent Framework (IMAF) as a domain-specific architectural model for museums. Rather than introducing a new general theory of autonomous agents, the framework is designed to support museums in applying agentic AI technologies in ways that are compatible with their institutional roles and practices as socio-technical knowledge institutions.
The remainder of this paper is organized as follows. Section 2 reviews related work on LLMs, agentic AI, and digital transformation in museums. Section 3 presents the research methodology based on the DSR paradigm. Section 4 describes IMAF. Section 5 demonstrates the framework through illustrative case study scenarios. Section 6 discusses the evaluation and technical feasibility of the framework, and Section 7 discusses implications, limitations, and directions for future research.

2. Related Work

2.1. LLM and Autonomous Agents

Recent artificial intelligence research has reached a major turning point with the development of LLMs. Since the introduction of the Transformer architecture [4], large-scale language models have demonstrated the ability to learn linguistic patterns from massive corpora and to perform a wide range of text-related tasks [5]. Subsequent work incorporating human feedback further improved their performance in responding to user queries and task instructions [6]. Beyond basic language understanding, text generation, or simple question answering, LLMs have very quickly advanced toward more complex reasoning, planning, and difficult problem-solving capabilities. As a result, LLM-based systems are now adopted across diverse domains, including education, culture, scientific research, and business applications [7,8].
On the basis of these technological developments, research has increasingly attempted to extend LLMs beyond conversational chat-based systems. Although LLMs’ fundamental function is to handle and transform textual inputs, recent studies have explored mechanisms by which they can invoke external systems, execute actions, and exchange data with their environment. Such systems are commonly referred to as agents or agentic AI. ReAct, for example, integrates reasoning processes with actions, enabling LLMs to solve problems in a step-by-step manner while making use of external tools [2]. Generative Agents further demonstrated that combining observation, planning, and reflection mechanisms allows simulated agents to exhibit behavior patterns that resemble those of humans in interactive environments [9]. A recent survey synthesizing this line of work also emphasizes that LLM-based autonomous agents can go beyond single-turn question-answering paradigms and behave like humans for autonomous decision-making and independent task completion [10].
Weng [1] proposed a widely cited conceptual framework for structuring agentic AI. Weng characterizes LLM-based autonomous agents in terms of four core components: a Brain (LLM-based reasoning engine), Memory (short- and long-term storage), Planning (goal decomposition and strategy formation), and Tool Use (integration with external tools and APIs). This framework identifies the functional components for moving from passive language models to systems capable of performing complex tasks in dynamic environments, and it has served as a reference model for subsequent agent-oriented studies. In addition, AutoGen proposes a multi-agent framework in which multiple LLM-based agents collaborate through structured conversations to solve problems collectively [11]. Taken together, these studies indicate that AI is evolving from text processing and generation toward autonomous agentic architectures that integrate planning, memory, and tool utilization.

2.2. Evolution of Museums with Digital Technology and AI

Since the 2000s, research in museum studies has examined how digital technologies extend traditional exhibition and operational practices. Parry [12] described this transformation as “recoding,” fundamentally reshaping core museum functions such as collection, authenticity, visitor experience, and narrative construction beyond mere technology utilization. In terms of visitor experience, technologies including handheld guides, multimedia systems, and interactive interfaces enabled personalized and participatory experiences of museum visitors [13,14]. These approaches continued to evolve with advances in data mining, artificial intelligence, and mobile computing, supporting adaptive and context-aware museum experiences.
With the emergence of LLM-based AI technologies, research in the museum domain has increasingly explored the use of LLMs to extend existing practices of information provision and collection management. These efforts can be broadly categorized into three areas: (1) accessibility and documentation of collections, (2) visitor guidance and interaction, and (3) personalized and context-aware recommendation services.
First, in the domain of collection accessibility and documentation, LLMs have been applied to automate the processing of domain knowledge and to improve searchability. For instance, Reusens et al. [15] and Weaver et al. [16] demonstrated that LLMs can be used to generate collection descriptions and to transcribe specimen labels, thereby enhancing the efficiency of record management and the accessibility of archival materials. Second, in the area of visitor experience and guidance services, LLM-based chatbots and conversational agents have been actively investigated. Wang and Matviienko [17], Ho et al. [18], Ariya et al. [19], and Vasic et al. [20] reported systems that apply generative AI or LLMs in physical and virtual museum environments to provide artwork interpretation, conversational guidance, and personalized virtual tours. Third, LLMs have also been explored in personalized and context-aware recommendation services. Trichopoulos et al. [21] and Ferrato [22] proposed recommendation and guidance systems that take into account visitor preferences, spatial context, and situational factors. These approaches move beyond static information delivery toward adaptive services that respond to user profiles and visiting conditions.
In parallel, broader research in cultural and artistic domains has begun to explore the potential of Agentic AI, where LLMs are investigated not only as tools for supporting specific functions but also as autonomous agents capable of coordinating multiple tasks. For example, in virtual and digital heritage environments, multi-agent interaction models have been proposed to support diverse interpretive perspectives. Su et al. [23] developed SimViews, a system in which multiple visitor agents engage in simulated conversations to present alternative viewpoints on artifacts, thereby reducing reliance on a single narrative. Aragon and D’Haro [24] introduced ASSIST, a multi-agentic framework integrating visual perception, speech interfaces, and retrieval-augmented generation (RAG) to enable context-aware guidance in cultural heritage settings. These studies suggest that multi-agent structures support interpretive diversification beyond single-channel information delivery.
Although still in the early stages, conceptual discussions about Agentic AI [25] have started to appear in the tourism, service, and arts sectors. For example, in the hospitality and tourism industry, Agentic AI systems are described as having autonomy, context awareness, and adaptive decision-making skills [26]. Rieder et al. [27] explored how AI in the art world is now seen as an independent actor rather than just a tool. More specifically for museums, Abdelfattah and Atef [28] proposed MuseePal, a multimodal agent architecture that uses structured knowledge representations to guide the generation of outputs.

2.3. Comparative Positioning and Design Implications

The literature reviewed above suggests that research relevant to museum-oriented agentic AI can be organized into several streams: general LLM-based autonomous agent architectures, museum-specific LLM applications, visitor-facing AI systems, context-aware recommendation systems, multi-agent and multimodal cultural heritage systems, cultural heritage knowledge representation, and socio-technical studies of museum digital transformation. These streams provide important foundations for the present study, but they remain only partially connected. In particular, prior work does not yet provide a system-level architectural model that explains how agentic autonomy can be aligned with the institutional authority, operational constraints, and knowledge-governance requirements of museums. Table 1 summarizes this positioning by comparing the main contributions and remaining limitations of prior research streams.
This comparison clarifies that the gap addressed by this study lies in the lack of a museum-specific architectural model that connects agentic reasoning, institutional governance, operational context, authoritative heritage knowledge, and role-differentiated interaction within a single system structure. The limitations identified in Table 1 are therefore reflected in the design of the IMAF in several ways. First, because general agent architectures provide reasoning, planning, memory, and tool-use mechanisms but remain domain-general, the IMAF adapts them to the museum domain by adding a dedicated Context component. Second, because museum LLM applications and visitor-facing systems tend to focus on narrow tasks, the IMAF is designed as an integrated architecture that can support multiple roles across visitor-facing and staff workflows. Third, while existing museum systems often treat context mainly as a personalization variable, the IMAF also stresses the institutional aspect of context so that ethical, interpretive, rights-related, conservation, spatial, and facility-related constraints can guide planning and action. Fourth, in order to utilize cultural heritage knowledge infrastructures that provide structured and reliable knowledge, the IMAF incorporates them within its design to support grounded reasoning and provenance-aware retrieval. In this sense, the IMAF is positioned not as a new foundation model or a general theory of autonomous agents, but as a museum-specific application architecture that reorganizes existing agentic AI capabilities into a governable socio-technical system.
These design implications directly inform the framework presented in Section 4. The proposed architecture extends prior agentic AI models by specifying how autonomy should be constrained and grounded in museum environments.

3. Research Methodology

3.1. Design Science Research Paradigm in Information Systems and Cultural Heritage

This study adopts the DSR paradigm as its methodological foundation, following the view that information systems research may contribute not only through explanation but also through the development of artifacts intended to address organizational problems. According to Hevner et al. [3], DSR focuses on the design and evaluation of artifacts, including constructs, models, methods, and instantiations.
The application of DSR is relevant to the field of digital cultural heritage, where technological innovation must be aligned with institutional practices and values. Previous work has underscored the importance of human-centered and design-oriented approaches for enabling successful digital transformation in museums, where design thinking practices help address challenges such as technology adoption, organizational change, and visitor engagement [33]. Related work has also shown that design-based frameworks can be used to structure museum experience in conjunction with digital technologies [34]. Within this context, DSR provides a suitable methodological basis for developing and evaluating an architectural framework intended for museum environments.
By adopting DSR, this study treats IMAF as a designed artifact rather than as a purely conceptual proposal. The framework is developed to address a specific problem identified in the museum domain, namely the lack of an architectural model that supports the application of agentic AI in ways that are compatible with institutional roles and operational conditions.

3.2. Problem Identification: Evolution of Agentic AI and Institutional Alignment Challenges

The motivation for developing the IMAF is associated with recent shifts in artificial intelligence research toward agentic AI, often discussed in terms of autonomy and proactivity—i.e., systems that can pursue objectives through adaptive decision-making with reduced need for continuous human intervention [35]. While LLMs have demonstrated substantial capabilities in language processing and reasoning, their deployment in museum environments raises domain-specific challenges.
A central issue concerns institutional alignment. Museums operate under formal mandates that govern interpretation, public communication, and the management of cultural materials [36]. In this context, autonomous systems that are primarily optimized for task completion may produce outputs that are inconsistent with institutional policies or ethical guidelines. Prior research has indicated that digital transformation in cultural organizations requires attention to organizational digital readiness and strategic alignment [32]. These findings suggest that the application of agentic AI in museums cannot be treated as a purely technical problem.
Another challenge concerns contextual awareness. General-purpose agents are typically developed without explicit consideration of spatial and temporal constraints. In museum environments, however, system behavior is shaped by operational conditions such as opening hours, gallery layouts, and conservation requirements. Failure to account for such conditions may result in recommendations or actions that are impractical or inappropriate within physical exhibition spaces.
These considerations indicate that the primary design problem addressed in this study is not the improvement of text generation performance using LLMs, but the development of an architectural model that supports the alignment of agentic autonomy with institutional and operational conditions in museums.

3.3. The DSR Process for IMAF Development and Evaluation

To address the identified problem, this study follows the DSR Methodology proposed by Peffers et al. [37], which defines a sequence of activities from problem identification to communication. The process begins with the identification of a domain-specific gap between general agentic AI architectures and museum requirements. This is followed by the definition of objectives focused on the development of an architectural framework that supports the application of agentic AI in museum contexts.
The design and development phase centers on the construction of the IMAF as the primary artifact. The framework builds on established agentic AI models, particularly the architectural components described by Weng [1], while adapting them to the museum domain. In this phase, the IMAF is specified as a multi-component architecture intended to support reasoning, memory, planning, and action within institutional constraints.
The demonstration phase employs scenario-based case studies to illustrate how the framework may be applied to representative museum tasks. These scenarios are designed to reflect both visitor-facing services and staff-oriented decision support functions. Because a full technical deployment is outside the scope of this study, the evaluation phase relies on qualitative analysis based on logical walkthroughs. This evaluation examines whether the proposed framework addresses the alignment and contextual challenges identified in Section 3.2. The methodological role of DSR in this study is therefore not to introduce a new general research method, but to make explicit the design logic through which a general agentic AI architecture is translated into a museum-specific artifact. The novelty of the artifact lies in this structured translation from domain requirements to architectural components, particularly the incorporation of institutional and operational constraints into the agentic loop.

4. Description of the Proposed Artifact: The Intelligent Museum Agent Framework

4.1. Overview of the Intelligent Museum Agent Framework

This study introduces IMAF as a domain-oriented architectural artifact developed through the DSR process. The framework is intended to support the application of agentic AI in museum environments by specifying how autonomous reasoning, institutional constraints, and structured knowledge resources may be integrated within a unified architecture. While the general autonomous agent model proposed by Weng [1]—comprising Brain, Planning, Memory, and Action—provides a functional blueprint for task-oriented problem solving, it is primarily formulated for general-purpose settings. Museums, by contrast, operate as socio-technical knowledge institutions in which factual accuracy, ethical responsibility, and institutional authority are central requirements. Accordingly, the IMAF extends Weng’s architecture by introducing a dedicated Context component and by refining existing modules to reflect the institutional and operational conditions characteristic of museums. This extension is not intended as a new algorithmic model of autonomous agents. Rather, its novelty lies in the architectural specification of how agentic AI components can be reorganized for museum environments.
The development of the IMAF was guided by four design objectives. First, the framework aims to support academic integrity and information reliability by emphasizing the use of curated and structured knowledge sources rather than unrestricted text generation. Second, it seeks to incorporate institutional policies and regulatory considerations by introducing governance mechanisms that constrain autonomous behavior in accordance with museum missions and professional standards. Third, the framework is designed to account for situational conditions such as spatial configuration and facility status, enabling the agent to respond to operational contexts in a manner consistent with physical and organizational environments. Fourth, the framework is intended to support both visitor-facing and staff-oriented use cases by enabling differentiated interaction modes aligned with user roles.
The principal components of the IMAF and their functional roles are summarized in Table 2. The framework is organized around five interrelated components: Brain, Context, Memory, Planning, and Action. These components are conceptualized as interacting modules within an agentic loop rather than as independent subsystems. Incoming user requests are interpreted by the Brain component, which coordinates task decomposition and response generation. During this process, the Context component provides institutional and operational constraints that shape permissible actions. The Memory component supplies access to short-term interaction history and long-term structured knowledge resources. The Planning component supports multi-step reasoning and workflow construction, while the Action component interfaces with internal and external tools to execute tasks. Detailed descriptions of the individual components are provided in the following subsections.
Figure 1 illustrates the hierarchical organization of the IMAF, delineating the relationship and interaction between the components in a hierarchical way. The architecture consists of three vertical layers: the Context layer (Layer 1) at the top establishes the regulatory boundaries through institutional and operational constraints; the Brain layer (Layer 2) serves as the central orchestrator; and the third layer at the bottom provides Memory and Planning resources (Layer 3). This vertical structure is intersected by a horizontal interaction axis that connects Users on the left with Action-mediated system integration on the right. By anchoring the reasoning process within this bifurcated governance framework, the IMAF ensures that agentic autonomy is systematically aligned with organizational policies and operational constraints.
In systems terms, the IMAF defines its boundary around the governed agentic loop and its interfaces to institutional repositories and operational sensing systems. Users and stakeholders, including visitors and staff shape system behavior through role-specific requests and oversight channels, while governance constraints regulate permissible actions.

4.2. Component 1: Context (Governance and Constraints)

The Context module operationalizes the institutional and operational conditions that constrain and guide agent reasoning and action. Unlike general-purpose agentic AI systems, which are typically designed for abstract task environments, museum-oriented agents must account for formal mandates, professional responsibilities, and physical operating conditions. In the IMAF, context is therefore not treated as general background information. Rather, it functions as a regulatory layer that determines whether a candidate response, recommendation, or action is permissible, requires revision, needs human confirmation, or must be blocked.
The Context module is bifurcated into Institutional Context and Operational Context. Institutional Context refers to relatively stable normative and policy-related conditions that shape interpretation, communication, rights management, ethical responsibilities, and professional conduct. Operational Context refers to more dynamic situational and environmental conditions, including gallery availability, opening hours, facility status, technical infrastructure, and visitor-flow constraints. This bifurcation allows the IMAF to distinguish between what the agent is institutionally permitted to do and what is practically feasible in a given operational situation.
During operation, the Context module is consulted by the Brain and Planning components before output generation or tool invocation. When a user request enters the system, the Brain identifies the user role and task type, while the Planning component constructs a candidate response or workflow. The Context module then evaluates the candidate plan through three checks: institutional permissibility, operational feasibility, and data sufficiency. Institutional permissibility examines whether the plan is consistent with interpretive policy, ethical requirements, rights and licensing restrictions, conservation rules, and institutional voice. Operational feasibility examines whether the plan is executable under current spatial, temporal, technical, and logistical conditions. Data sufficiency examines whether the available information is authoritative, complete, and reliable enough to support the proposed response or action. Only plans that satisfy these checks may proceed to the Action component without further intervention.
To support interpretation and practical application of the framework, this study presents an illustrative taxonomy showing how such constraints may be structured in museum environments.
The illustrative taxonomy is grounded in two sources that define the fundamental roles, responsibilities, and entities of common museums. First, it draws on institutional role definitions articulated by the International Council of Museums (ICOM), characterizing museums through functions such as collecting, conserving, interpreting, exhibiting, and educating [36]. Second, it reflects the conceptual distinctions formalized in CIDOC CRM (Conceptual Reference Model) [29], which differentiate among entities such as physical objects, events, actors, places, time-spans, and rights. Together, these sources provide a principled basis for illustrating how context may be organized for museum-oriented agentic AI.

4.2.1. Institutional Context

Institutional context encompasses the policy-related constraints and normative requirements derived from the formal roles and social responsibilities of cultural organizations. These constraints ensure that an agent’s outputs—such as explanations, recommendations, or autonomous decisions—remain aligned with established professional standards and institutional authority. Table 3 presents representative categories of institutional context, illustrating how professional mandates and organizational norms can be translated into structured constraints for agentic behavior. This classification serves as an exemplary framework for system design rather than an exhaustive or prescriptive set of requirements.
Operationally, institutional context functions as a high-priority constraint layer because it reflects the museum’s normative authority and public responsibilities. When institutional constraints conflict with visitor preferences, personalization goals, or operational convenience, the institutional constraint should generally take precedence. For example, if a visitor requests a speculative provenance story, the Interpretation and Voice constraint restricts the agent from producing unsupported claims. If a staff user requests the public release of a high-resolution image with unresolved rights status, the Rights and Licensing constraint prevents direct execution and requires staff confirmation. Similarly, ethics- or conservation-sensitive cases should not be resolved solely by the agent. Instead, the agent should present the relevant constraint, identify the source of uncertainty, and escalate the issue to an authorized museum professional.

4.2.2. Operational Context

Operational context encompasses constraints associated with the immediate physical, technical, and procedural conditions of museum management. Unlike the institutional context, which is rooted in relatively long-term and usually persistent mandates, these constraints are derived from the real-time availability of resources and the logistical requirements of everyday operations. They represent the situational variables that determine whether the actions that can be derived from IMAF’s reasoning are practically feasible. Table 4 summarizes these categories, providing a schema for synchronizing agent behavior with the museum’s lived environment. This classification also serves as an exemplary framework that can be adapted to the specific logistical circumstances of individual institutions.
Operationally, operational context functions as a feasibility and adaptation layer. In many cases, operational constraints do not block the task itself but require the agent to revise the plan. For example, if a preferred gallery is temporarily closed or overcrowded, the agent may reroute a visitor to an alternative space while preserving the original interpretive objective. If an interactive display, elevator, or climate-control system is unavailable, the agent may modify the recommendation or suggest a different gallery. However, when operational constraints involve safety, conservation risk, or access control, they may become blocking constraints rather than simple adaptation variables. In such cases, the agent should not execute the action autonomously but should either provide a safe alternative or escalate the issue to staff.

4.2.3. Contextual Constraint Handling and Human Oversight

The bifurcated Context module is intended not only to classify institutional and operational constraints, but also to regulate agent behavior when relevant information is incomplete, conflicting, or uncertain. In such situations, the IMAF treats uncertainty as a condition for constraint-aware limitation rather than as permission for autonomous generation or execution. When contextual data are incomplete, the agent should avoid unsupported claims, indicate the limits of available information, and rely on additional authoritative sources or human review before producing consequential recommendations. When institutional and operational constraints conflict, institutional constraints generally take priority over personalization goals or operational convenience because they reflect the museum’s public responsibilities, professional mandates, and legal or ethical obligations. However, operational constraints may override an otherwise valid plan when physical feasibility, visitor safety, accessibility, conservation conditions, or facility status are at stake. If a conflict cannot be resolved through predefined priorities, or if the decision involves sensitive cultural interpretation, rights and licensing, conservation risk, contested provenance, or changes to institutional records, the agent should not resolve the issue independently. Accordingly, human review is required for tasks involving professional authority, sensitive ethical or legal judgment, conservation or provenance concerns, rights-related decisions, conflicting data, or irreversible institutional consequences. Instead, it should present the relevant constraint, identify the source of uncertainty or conflict, and escalate the case to authorized museum staff. Through this mechanism, the Context module functions as an enforcement and escalation layer that keeps agentic autonomy within institutional and operational boundaries while preserving human oversight for high-risk decisions.

4.3. Component 2: The Brain

The Brain functions as the central coordinating component of the IMAF, using the reasoning capabilities of LLMs to interpret user intent and to manage interactions among other modules. While this component follows the general concept of an LLM-based reasoning core proposed by Weng [1], the IMAF refines it by incorporating a Persona and Tone Controller. In museum contexts, the manner in which information is communicated is closely related to institutional authority and interpretive responsibility. Accordingly, the Brain operates through a Dual Interface that distinguishes between visitor-facing and staff-oriented interaction modes. In visitor-facing interactions, it adopts a docent-like persona oriented toward explanation and learning support, whereas in staff-oriented interactions it adopts an assistant-like persona oriented toward task execution and professional communication. This design allows the system to adapt its outputs to different user roles while maintaining consistency with the museum’s institutional identity.

4.4. Component 3: Memory

Memory in the IMAF adapts and refines Weng’s [1] memory categories to support both short-term interaction continuity in museum environments and long-term access to persistent knowledge resources. Consistent with the framework overview in Section 4.1, it comprises a Perceptual Buffer for immediate sensory inputs, an Active Interaction Context for short-term interaction history, and a long-term Heritage and Archive layer that provides access to structured knowledge resources, including knowledge graph representations and exhibition records.
The Perceptual Buffer processes transient inputs such as images of artifacts and spoken user queries, enabling the system to register information that may not be available in textual form. The Active Interaction Context maintains the immediate conversational state (e.g., prior turns and referenced objects) to support coherence across multi-turn interactions.
The long-term Heritage and Archive layer can be implemented using a knowledge graph approach to represent explicit relations among entities [30]. This layer aligns with the ontological and relational structuring of knowledge, such as those formalized in CIDOC CRM [29], which organize cultural heritage data by linking physical objects, actors, and places through mediating events and time-spans. While vector-based retrieval supports similarity-oriented access to content, the knowledge graph encodes these formal relational structures essential for interpretive, historical, and provenance-oriented queries. The two mechanisms are treated as complementary: vector retrieval facilitates approximate semantic matching, whereas the knowledge graph enables relation-based access and multi-step traversal across the ontological network. For example, rather than relying on surface-level textual similarity, the knowledge graph can represent how a specific historical event influenced an artist’s choice of themes, and how that influence is manifested in extant artworks (Figure 2). This configuration ensures grounded access to cultural heritage knowledge by situating objects within their broader historical and social contexts [31].
The Heritage and Archive layer should be understood as a governed access layer rather than as a mechanism for independently updating museum knowledge. In the IMAF, updates to authoritative cultural heritage records—such as object metadata, provenance information, rights status, interpretive descriptions, or exhibition records—remain the responsibility of existing museum information systems and institutional procedures, including collection management systems, archival databases, curated knowledge graphs, staff review, version control, and source attribution. The agent retrieves and checks these sources by considering authority, provenance, recency, and consistency, but it does not autonomously overwrite institutional records. Dynamic operational knowledge, such as gallery occupancy, facility status, environmental conditions, or equipment availability, may be updated automatically through sensors, IoT devices, and facility-management APIs and then used as contextual input for planning and action. When retrieved sources are incomplete, outdated, or inconsistent, the agent should avoid unsupported generation, indicate the uncertainty, prioritize institutionally verified records where available, and request staff review for consequential interpretive, rights-related, or provenance-related decisions.

4.5. Component 4: Planning

The Planning component supports the decomposition of complex requests into a sequence of executable sub-tasks within the IMAF. It employs Chain-of-Thought (CoT) prompting [1,38] and self-reflection mechanisms to structure multi-step reasoning processes. For instance, when a staff user requests assistance in designing an exhibition layout, the Planning component may divide the task into stages such as artifact selection, narrative sequencing, and spatial arrangement.
During this process, candidate plans are assessed with reference to both the Institutional Context and the Operational Context defined in Section 4.2. This reflective step is intended to function as a constraint mechanism rather than as an optimization procedure: proposed workflows are examined for consistency with interpretive principles, ethical considerations, and situational conditions prior to execution. Through this configuration, the Planning component integrates task decomposition with context-aware validation, supporting the generation of workflows that remain aligned with institutional requirements and operational conditions. In this process, the constraint-handling logic defined in Section 4.2.3 determines whether a candidate plan can proceed, should be revised, requires clarification, needs human confirmation, or must be blocked before execution.

4.6. Component 5: Action

The Action component provides the interface through which the IMAF executes the outcomes of reasoning and planning by interacting with external systems. While the Brain and Planning components generate and validate task structures under institutional and operational constraints, the Action component enables the controlled execution of these structures through tool invocation [2]. However, tool invocation in the IMAF is not treated as an unrestricted final step of planning. Before any tool is invoked, the Action component receives only those plans that have passed the contextual checks conducted by the Planning and Context components. Plans that require confirmation or escalation are withheld from automatic execution, while plans that violate blocking constraints are not executed. This gated execution establishes the delegation boundary for agentic action in museum environments.
Within this boundary, the degree of delegation depends on the institutional risk, reversibility, and authority required for the action. Low-risk and reversible operations, such as retrieving records, querying occupancy data, generating draft explanations, producing analytical summaries, or recommending routes, may be automated after contextual checks. Actions with institutional, legal, ethical, or operational consequences, such as publishing content, updating collection metadata, approving image reuse, modifying schedules, or sending external communications, should require explicit human confirmation before execution. Decisions that require professional judgment or institutional authority—such as provenance determinations, conservation judgments, rights and licensing decisions, repatriation-related decisions, acquisition or deaccession decisions, and final approval of sensitive interpretation—should not be delegated to the agent as final decisions. In these cases, the Action component may support staff by retrieving evidence, preparing drafts, or comparing options, but final authority remains with authorized museum professionals.
To support conceptual understanding and practical use of the IMAF, this study proposes three tool categories for the Action component as a reference framework, consistent with the definition presented in Section 4.1: a Code Interpreter, an Internal API Hub, and an External API Hub (Table 5).
First, the Code Interpreter supports computational and analytical operations that cannot be reliably performed through natural language generation alone. Prior research in data-driven museum management highlights the importance of analytical processing for decision-making based on various museum operation data [32]. Within the IMAF, the Code Interpreter is therefore conceptualized as a mechanism for executing numerical calculations, statistical analyses, and data mining that complement the Planning component’s reasoning processes. This design reflects the practical requirements that language models require auxiliary analytical tools to handle structured data and quantitative tasks.
Second, the Internal API Hub provides controlled access to museum-managed systems, including collection management systems (CMS), facility and sensor infrastructures (e.g., IoT-based occupancy monitoring), and other administrative platforms. The Internal API Hub is structurally aligned with the Heritage and Archive and the Operational Context components described in previous sections, enabling the agent to retrieve authoritative institutional data and to incorporate real-time operational constraints into planning and execution processes.
Third, the External API Hub connects the IMAF to information resources maintained outside the institution, such as cultural heritage databases, general-purpose information services (e.g., weather), and public communication platforms. Accordingly, it plays the role of extending institutional knowledge by referencing external, potentially authoritative sources while preserving the distinction between internal and external provenance.
Together, these three categories structure the Action component as an interface layer to useful and essential tools outside the LLM-based AI agent itself. Through this configuration, the Action component connects abstract reasoning outcomes to concrete technical infrastructures, ensuring that the execution of agentic behavior remains grounded in institutional data, operational conditions, and established channels of information exchange.

5. Artifact Demonstration: Case Study Scenarios

This section presents two scenario-based case studies to illustrate the operationalization of the IMAF under representative socio-technical conditions. The two scenarios were selected to represent two common classes of museum use cases: visitor-facing engagement and staff-oriented institutional support. Together, they demonstrate the applicability of the IMAF to both public interpretation and internal museum operations. The demonstrations examine how the hierarchical architecture mediates agent behavior through vertical governance and horizontal interaction axes, ensuring that autonomous reasoning remains situated within institutional and operational constraints.

5.1. Scenario 1: Personalized Visitor Engagement and Context-Aware Curation

The first scenario considers a visitor accompanied by an eight-year-old child who requests a personalized one-hour tour focused on the theme of “nature in art” with the additional constraint of avoiding overcrowded galleries. Upon receiving this input through the Perceptual Buffer, the Brain (Layer 2) activates its Persona and Tone Controller to adopt a docent-like mode for visitor engagement. This setting is strictly regulated by the Interpretation and Voice category within the Institutional Context (Layer 1), which mandates that generated narratives remain factually accurate while employing a pedagogical style appropriate for a minor. The Planning component (Layer 3) subsequently decomposes the request into discrete sub-tasks, including thematic artifact selection and narrative sequencing. During this process, the system consults the Operational Context (Layer 1), specifically referencing the Capacity and Flow Logistics category, to integrate real-time occupancy data. When a primary gallery is identified as operating at peak capacity, the Planning module dynamically reroutes the tour toward a less congested exhibition space. The system then utilizes relation-based retrieval within the Heritage and Archive layer (Memory) to identify connections between nineteenth-century industrialization and thematic shifts in landscape art, providing a grounded interpretive pathway through the knowledge graph. Finally, the Action component operationalizes the plan through the Internal API Hub, delivering a navigational map and synchronized audio guidance via the museum’s mobile application. The dynamic message flow and inter-component interactions for this visitor-facing service are illustrated in the sequence diagram in Figure 3.

5.2. Scenario 2: Operational Staff Support and Ethical Decision-Making

The second scenario examines the use of the IMAF as a support system for curatorial staff, where a junior curator requests assistance in identifying artifacts for a “digital heritage” exhibition proposal. Recognizing the staff role through the Dual Interface, the Brain (Layer 2) switches to an assistant-like persona oriented toward technical accuracy and professional task execution. The Planning component (Layer 3) decomposes the request into stages involving artifact retrieval, ethical assessment, and spatial feasibility analysis. During the retrieval stage, the Heritage and Archive layer (Memory) identifies an eighteenth-century artifact relevant to the thematic focus. However, the Institutional Context (Layer 1) identifies this artifact within the Ethics and Sensitivity category as being subject to an ongoing repatriation discussion. Based on this normative constraint, the system recommends the use of a high-resolution digital surrogate rather than the physical object to maintain consistency with institutional ethical guidelines. Simultaneously, the system evaluates the proposal against the Operational Context (Layer 1), specifically the Infrastructure and Asset Status category, which indicates that the curator’s initially selected gallery has a climate control malfunction. The system proposes an alternative gallery with appropriate environmental facilities. To further refine the proposal, the Action component invokes the Code Interpreter to analyze historical visitor flow data, supporting the assessment of display configurations through statistical processing. The resulting output is a structured proposal that integrates ethical considerations, spatial feasibility, and data-driven insights. Figure 4 provides a sequence diagram detailing the multi-step reasoning and validation process for this staff-oriented decision support function.

6. Evaluation and Technical Feasibility

6.1. Goal-Oriented Qualitative Evaluation

The evaluation of the IMAF is conducted through a criteria-based logical walkthrough, which is appropriate for this research, where a conceptual ex ante artifact needs to be evaluated in an artificial environment. Although this approach may not capture the full details and complexity of real-world deployment environments, it offers the distinct advantage of enabling rapid assessment of novel technological concepts without fully implementing the system or exposing participants or institutions to operational risks. Within DSR, such evaluation in an artificial environment is appropriate for early-stage conceptual artifacts, particularly when naturalistic deployment is premature or difficult. In the present case, where agentic AI in museums remains at an early stage and real-world experimentation may be constrained, the walkthrough can be used as a legitimate ex ante evaluation step rather than as a substitute for empirical validation [39]. Consistent with this approach, the four evaluation criteria—Reliability, Alignment, Adaptability, and Versatility—are derived directly from the design objectives formulated in Section 4.
Four evaluation criteria are operationalized as follows. Reliability corresponds to the design objective of supporting information accuracy and interpretive authority through the use of curated and structured knowledge sources. Alignment corresponds to the objective of incorporating institutional policies and regulatory considerations into agent behavior through governance mechanisms. Adaptability corresponds to the objective of accounting for situational conditions such as spatial configuration and facility status, enabling context-sensitive responses. Versatility corresponds to the objective of supporting both visitor-facing and staff-oriented use cases through differentiated interaction modes.
Against these criteria, the IMAF components are assessed through a logical walkthrough of the scenario-based demonstrations presented in Section 5. Reliability is demonstrated through the use of the Heritage and Archive layer (Layer 3), which enables relation-based retrieval from the knowledge graph when generating interpretive content. Rather than producing outputs through unconstrained text generation, the system grounds its responses in structured institutional records, thereby contributing to factual consistency in public-facing museum communication. Alignment is illustrated by the system’s incorporation of institutional and operational constraints (Layer 1) into task execution—most notably, the recommendation to use a digital surrogate for an artifact under repatriation review (Section 5.2) and the rerouting of visitor guidance away from overcrowded galleries in Scenario 1 and the recommendation of an alternative gallery due to a climate-control malfunction in Scenario 2. These behaviors confirm that governance, formalized in the bifurcated Context module (Layer 1), is actively integrated into the reasoning of the Planning component (Layer 3) rather than applied as a post hoc filter. Adaptability is demonstrated in Scenario 1, in which real-time occupancy data obtained through the Operational Context (Layer 1) dynamically reshapes tour routing decisions, showing that the framework responds to situational variability through the coordination of the Planning module (Layer 3). Versatility is reflected in the contrasting visitor-oriented and staff-oriented scenarios, in which the Brain (Layer 2) adjusts its persona and interaction mode via the Dual Interface while relying on the same underlying hierarchical architecture, suggesting that a unified governance structure can support differentiated interaction modes.
The alignment between design objectives, evaluation criteria, and demonstrated system behavior across both scenarios indicates that the IMAF demonstrates internal coherence and goal-consistent behavior as a conceptual artifact, providing support for its architectural soundness in museum environments. The objective–criterion traceability of the evaluation is summarized in Table 6, which maps each design objective to its corresponding evaluation criterion and architectural mechanism.

6.2. Technical Feasibility and Agentic Orchestration

The technical feasibility of the IMAF is examined by mapping its conceptual components to an implementable agentic orchestration structure. This study adopts a CrewAI-style orchestration pattern as a reference model to illustrate how the framework could be instantiated using contemporary role-based agent systems [40]. This choice is motivated by the conceptual compatibility between the IMAF and agent architectures that separate reasoning, task specification, and tool-mediated action into distinct but coordinated elements. The code presented below is therefore provided as CrewAI-style pseudo-code intended to clarify one plausible orchestration path and to examine technical feasibility at a conceptual level, rather than as a fully executable implementation.
A.
Hierarchical Configuration: Context and Foundation Resources
The configuration phase establishes the system’s boundary conditions by defining the institutional_context and operational_context as inputs. These modules formalize high-level museum mandates—such as ethical interpretation policies and real-time facility status—into structured instructions that the agent must prioritize. Simultaneously, resources such as memory and tools are initialized, including the HeritageKnowledgeGraphAPI for relational data retrieval and a hub of action-oriented tools for external system integration. This configuration ensures that the governance layer is established as a prerequisite, providing a regulatory framework for the subsequent reasoning and action cycles.
Systems 14 00954 i001
In this configuration, institutional policies and operational conditions are provided as explicit inputs to the agent’s reasoning process. The knowledge graph resource is accessed through the Internal or External API Hub as part of the Heritage and Archive layer, rather than being treated as a separate cognitive component.
B.
Orchestration Logic: Governance-Centric Execution
The IMAF_Hierarchical_Orchestrator class serves as the functional embodiment of the Brain, coordinating the interaction between governance and execution. Upon initialization, the Brain adopts a specific persona via the Dual Interface, ensuring that its reasoning perspective is aligned with the user’s institutional role. The core of the orchestration logic resides in the Task object, where governance is explicitly injected into the agentic loop. By assigning self.governance_layer as a context parameter, the Planning module is technically mandated to validate its reasoning steps against institutional and operational constraints before proceeding to tool invocation. This hierarchical orchestration ensures that the Action layer operates only within the safety and policy boundaries defined by the top-level governance, thereby achieving domain-constrained autonomy.
Systems 14 00954 i002
Here, is_critical_task() denotes an abstract judgment of whether a task requires additional human oversight under the museum’s governance conditions, while human_in_the_loop_review() represents the corresponding approval step. These procedures are kept schematic in the pseudo-code because their detailed criteria and workflows would depend on actual institutional requirements and implementation settings. In a real-world deployment, task validation and domain-specific reasoning could be further strengthened through museum-specific model adaptation (e.g., fine-tuning or instruction tuning), retrieval-augmented generation grounded in the Heritage and Archive layer, and self-reflection or verification procedures within the Planning component. Because these are general enhancement strategies for improving agent performance and reliability rather than defining architectural elements of the IMAF itself, they are not exhaustively represented in the present pseudo-code, whose purpose is to illustrate the orchestration logic at the architectural level.
Through these examples, the IMAF is shown to be technically realizable as an orchestrated agent system in which reasoning processes are guided by contextual constraints and supported by structured knowledge and tool-based execution. As illustrated in the pseudo-code examples, key inputs to the framework—such as institutional policies, operational conditions, and task instructions—can be expressed in natural language. This design enables the IMAF not only to incorporate existing textual resources, including policy documents and operational guidelines, but also to accept newly provided instructions and situational inputs without requiring extensive formalization or re-encoding.
It should be noted that the pseudo-code presented above simplifies several components that would require substantial technical infrastructure in a real deployment. In particular, the Heritage and Archive layer would require access to a dedicated graph database system equipped with a query language interface to support structured relation traversal and provenance-aware retrieval consistent with CIDOC CRM–aligned ontologies. Similarly, the Internal API Hub would necessitate formal API contracts with CMSs and IoT platforms. Each of these requirements, however, corresponds to well-established technical solutions that are already in use within the digital heritage and information systems communities. Taken together, the architectural components of the IMAF map onto available technologies and implementation patterns, suggesting that a technically plausible instantiation can be achieved within the current state of the art.

7. Discussion

7.1. Theoretical Implications: Structuring Domain-Constrained Autonomy

This study extends the generalized agent architecture proposed by Weng [1] to the specific institutional context of cultural heritage organizations. The primary theoretical contribution lies in the reconfiguration of the agentic cognitive loop into a hierarchical structure, where institutional governance and operational constraints are treated as integral systemic regulators. Whereas general-purpose agents typically prioritize task completion within relatively neutral contexts, the IMAF introduces a bifurcated Context module (Layer 1) as the primary regulatory layer. By distinguishing between Institutional Context (ethical and policy-related constraints) and Operational Context (spatiotemporal and physical conditions), the framework establishes a systematic approach for embedding organizational values directly into the reasoning cycle of the Planning module (Layer 3). Furthermore, rather than merely appending a context component, the IMAF demonstrates how the synergistic interaction between its constituent layers facilitates the achievement of organizational goals, ensuring that all agentic actions remain strictly aligned with institutional mandates and real-time operational requirements.
The study also suggests an alternative perspective on memory design for domain-specific agents. Instead of relying exclusively on probabilistic vector-based retrieval, the IMAF incorporates a Heritage and Archive layer (Layer 3) implemented through structured knowledge resources, such as knowledge graphs. As demonstrated in Section 5, this configuration enables relation-based access to historical information, allowing the agent to connect artifacts and interpretive narratives in a coherent manner. This approach emphasizes the necessity of grounding agentic behavior in structured knowledge to ensure factual consistency in knowledge-intensive environments where interpretive coherence is a central requirement.

7.2. Practical Implications: Supporting Engagement and Institutional Operations

From a practical perspective, the IMAF provides a conceptual guideline for cultural heritage institutions seeking to adopt LLM-based systems while maintaining institutional authority. Digital transformation in museums is often hindered by concerns regarding interpretive control and ethical responsibility. By formalizing these concerns within the governance layer, the framework offers a mechanism for integrating agentic AI into museum workflows without displacing existing organizational structures.
The evaluation and technical feasibility analysis indicate that the efficacy of agentic AI in museum environments is not contingent upon the total automation of institutional workflows. Instead, the IMAF functions as a strategic intermediary layer that aligns user requests with pre-defined institutional policies and real-time situational constraints. This alignment is critical for maintaining a balance between personalized user experiences and institutional authority. For visitor-facing services, it ensures that high-level personalization remains consistent with authoritative curatorial narratives. For staff-oriented tasks, the framework can help structure recommendations so that subsequent actions remain subject to established ethical and procedural review.
Beyond immediate operational support, the framework suggests a pathway for implementing advanced AI capabilities for institutions with limited technical resources. By emphasizing structured governance and reusable architectural components, the IMAF enables museums to conceptualize agent-based services in a standardized manner, rather than as ad hoc technical solutions. Within such settings, the framework can be interpreted as a form of professional support system, assisting staff in navigating complex policy requirements and operational conditions while preserving human oversight in critical decision-making processes.
Viewed through a socio-technical lens, the IMAF frames agentic AI not as a substitute for institutional work but as an architectural means to coordinate human decision-making, governance constraints, and technical execution. In this sense, the framework formalizes how AI-based autonomy can be exercised within institutional accountability, enabling AI-enabled services to remain aligned with museum authority while adapting to situational operational demands.

7.3. Ethical Implications: Responsible Agentic AI in Cultural Heritage

The deployment of agentic AI in museums raises ethical issues that extend beyond general concerns about algorithmic bias and data privacy. Because museums mediate cultural memory, public interpretation, and institutional authority, agentic systems may affect how sensitive histories, contested provenance, and culturally specific meanings are represented. In this context, risks include the simplification of culturally sensitive narratives, the amplification of uncertain or incomplete provenance information, inappropriate reuse of copyrighted or restricted digital assets, and the personalization of visitor services based on sensitive behavioral data. These risks are particularly important because museum audiences may interpret AI-generated outputs as institutionally authorized knowledge.
The IMAF addresses these concerns at the architectural level by treating ethical constraints as part of the Institutional Context rather than as post hoc content moderation. Categories such as Ethics and Sensitivity, Rights and Licensing, and Interpretation and Voice can define boundaries for permissible reasoning and action, while the Heritage and Archive layer grounds outputs in structured and provenance-aware knowledge resources. However, these mechanisms should be understood as design-level safeguards rather than guarantees of ethical correctness. In practical deployment, museums would still need human review procedures, source-audit mechanisms, privacy-preserving data practices, and clear accountability rules for decisions involving sensitive cultural materials or public-facing interpretation.

7.4. Limitations and Future Research

Despite its contributions, this study has several limitations. First, the IMAF presupposes a baseline level of digital infrastructure, including structured collection data, interoperable system interfaces, and access to operational context information. Such conditions may not be present in smaller regional museums, community-based institutions, or organizations with limited technological resources. In these environments, the absence of formalized knowledge graphs, integrated APIs, or real-time facility data may constrain the practical applicability of the framework. Future research should therefore explore lightweight or modular adaptations of the IMAF that accommodate varying levels of digital maturity across museum contexts.
Second, the evaluation conducted in this study is limited to a goal-oriented, scenario-based qualitative walkthrough performed in an artificial environment. While this approach is consistent with ex ante design science evaluation for conceptual artifacts [39], it does not substitute for empirical validation in operational settings. Accordingly, the contribution of the present study lies in proposing and analytically examining a museum-oriented architectural artifact through ex ante evaluation. Empirical assessment of its practical value in deployed museum settings remains an important task for future research. Future validation may involve expert review by museum professionals, small-scale pilot testing in a limited gallery or staff workflow, and user studies with visitors and staff.
Third, as with other AI-driven systems, the deployment of the IMAF can raise concerns related to algorithmic bias, data privacy, and the risk of erroneous or misleading outputs [41,42]. These risks are particularly salient in museum contexts, where public trust depends on factual reliability and ethical stewardship. One possible direction for future work is the systematic integration of human-in-the-loop mechanisms, enabling museum professionals to review and intervene in critical situations of agent operation [43]. Such arrangements would allow the framework to support accountable decision-making while preserving institutional oversight and authority.

Funding

This study was funded by 2025 Hongik University Research Fund.

Data Availability Statement

Data are contained within the article.

Acknowledgments

During the preparation of this manuscript, the author used OpenAI’s ChatGPT 5.2 and Google’s Gemini 3 to improve the texts and readability, and to generate sequence diagrams via programming. The author has reviewed and edited the content after using these tools and takes full responsibility for the final content of the publication.

Conflicts of Interest

The author declares that there is no conflict of interest.

References

  1. Weng, L. LLM-Powered Autonomous Agents; Lil’Log. 2023. Available online: https://lilianweng.github.io/posts/2023-06-23-agent/ (accessed on 20 February 2026).
  2. Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; Cao, Y. ReAct: Synergizing Reasoning and Acting in Language Models. arXiv 2022, arXiv:2210.03629. [Google Scholar]
  3. Hevner, A.R.; March, S.T.; Park, J.; Ram, S. Design science in information systems research. MIS Q. 2004, 28, 75–105. [Google Scholar] [CrossRef]
  4. Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, L.; Polosukhin, I. Attention Is All You Need. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Long Beach, CA, USA, 4–9 December 2017. [Google Scholar]
  5. Brown, T.B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. Language Models are Few-Shot Learners. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Virtual, 6–12 December 2020. [Google Scholar]
  6. Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. Training language models to follow instructions with human feedback. arXiv 2022, arXiv:2203.02155. [Google Scholar]
  7. Bommasani, R.; Hudson, D.A.; Adeli, E.; Altman, R.; Arora, S.; von Arx, S.; Bernstein, M.S.; Bohg, J.; Bosselut, A.; Brunskill, E.; et al. On the Opportunities and Risks of Foundation Models. arXiv 2021, arXiv:2108.07258. [Google Scholar]
  8. Zhao, W.X.; Zhou, K.; Li, J.; Tang, T.; Wang, X.; Hou, Y.; Min, Y.; Zhang, B.; Zhang, J.; Dong, Z.; et al. A Survey of Large Language Models. arXiv 2023, arXiv:2303.18223. [Google Scholar]
  9. Park, J.S.; O’BRien, J.; Cai, C.J.; Morris, M.R.; Liang, P.; Bernstein, M.S. Generative Agents: Interactive Simulacra of Human Behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST), San Francisco, CA, USA, 29 October–1 November 2023. [Google Scholar]
  10. Wang, L.; Ma, C.; Feng, X.; Zhang, Z.; Yang, H.; Zhang, J.; Chen, Z.; Tang, J.; Chen, X.; Lin, Y.; et al. A survey on LLM based autonomous agents. arXiv 2023, arXiv:2308.11432. [Google Scholar]
  11. Wu, Q.; Bansal, G.; Zhang, J.; Wu, Y.; Li, B.; Zhu, E.; Jiang, L.; Zhang, X.; Zhang, S.; Liu, J.; et al. Autogen: Enabling next-gen LLM applications via multi-agent conversations. In In Proceedings of the First Conference on Language Modeling, Philadelphia, PA, USA, 7–9 October 2024. [Google Scholar]
  12. Parry, R. Recoding the Museum: Digital Heritage and the Technologies of Change; Routledge: Abingdon, UK, 2007. [Google Scholar]
  13. Tallon, L.; Walker, K. (Eds.) Digital Technologies and the Museum Experience: Handheld Guides and Other Media; AltaMira Press: Lanham, MD, USA, 2008. [Google Scholar]
  14. Marty, P.F. My lost museum: User expectations and motivations for creating personal digital collections on museum websites. Libr. Inf. Sci. Res. 2011, 33, 211–219. [Google Scholar] [CrossRef]
  15. Reusens, M.; Adams, A.; Baesens, B. Large Language Models to make museum archive collections more accessible. AI Soc. 2025, 40, 4485–4497. [Google Scholar] [CrossRef]
  16. Weaver, W.N.; Ruhfel, B.R.; Lough, K.J.; Smith, S.A. Herbarium specimen label transcription reimagined with large language models: Capabilities, productivity, and risks. Am. J. Bot. 2023, 110, e16256. [Google Scholar] [CrossRef] [PubMed]
  17. Wang, H.; Matviienko, A. Experiencing Art Museum with Generative AI Chatbot. In Proceedings of the 2025 ACM International Conference on Interactive Media Experiences, Niterói, Brazil, 3–6 June 2025. [Google Scholar]
  18. Ho, H.P.; Ramesh, V.; Zaloudek, I.; Rikhtehgar, D.J.; Wang, S. Enhancing Visitor Engagement in Interactive Art Exhibitions with Visual-Enhanced Conversational Agents. In Proceedings of the 30th International Conference on Intelligent User Interfaces (IUI ‘25), Cagliari, Italy, 24–27 March 2025; pp. 660–671. [Google Scholar]
  19. Ariya, P.; Khanchai, S.; Intawong, K.; Puritat, K. Enhancing textile heritage engagement through generative AI-based virtual assistants in VR museums. Comput. Educ. X Real. 2025, 7, 100112. [Google Scholar] [CrossRef]
  20. Vasic, I.; Fill, H.-G.; Quattrini, R.; Pierdicca, R. LLM-Aided Museum Guide: Personalized Tours Based on User Preferences. In Extended Reality. XR Salento 2024; De Paolis, L.T., Arpaia, P., Eds.; Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2024; pp. 249–262. [Google Scholar]
  21. Trichopoulos, G.; Konstantakis, M.; Alexandridis, G.; Caridakis, G. LLMs as Recommendation Systems in Museums. Electronics 2023, 12, 3829. [Google Scholar] [CrossRef]
  22. Ferrato, A. Integrating Indoor Positioning, Recommendation, and Personalization to Enhance Museum Visitor Experiences. In Proceedings of the 33rd ACM Conference on User Modeling, Adaptation and Personalization (UMAP ‘25), New York, NY, USA, 16–19 June 2025; pp. 388–392. [Google Scholar]
  23. Su, M.; Liu, C.; Zhang, J.; Shuang, W.; Fan, M. SimViews: An Interactive Multi-Agent System Simulating Visitor-to-Visitor Conversational Patterns to Present Diverse Perspectives of Artifacts in Virtual Museums. In Proceedings of the 33rd ACM International Conference on Multimedia, Dublin, Ireland, 27–31 October 2025. [Google Scholar]
  24. Aragon Diaz, L.L.; D’Haro, L.F. ASSIST: A Multi-Agentic Framework for Human Computer Interaction in Cultural Heritage settings. In Proceedings of the 12th International Conference on Information Management and Big Data, Lima, Peru, 29–31 October 2025. [Google Scholar]
  25. Pacioni, E.; Coman, A.C.; Calvaresi, D.; Manzo, G.; Schumacher, M. Agent-Based Hybrid AI Models and Technologies: A Systematic Literature Review. IEEE Access 2026, 14, 21148–21166. [Google Scholar] [CrossRef]
  26. Dwivedi, Y.K.; Helal, M.Y.I.; Elgendy, I.A.; Albashrawi, M.A.; Hughes, L.; Shawosh, M.; Dutot, V.; Jeon, I. Artificial intelligence agents and agentic systems in hospitality and tourism: Challenges, opportunities and research agenda. Int. J. Contemp. Hosp. Manag. 2026, 38, 27–52. [Google Scholar]
  27. Rieder, A.; Pappas, I.O.; Griffith, T.L. Trajectories of artificial intelligence: Visions from the art world. J. Inf. Technol. Case Appl. Res. 2025, 27, 199–206. [Google Scholar] [CrossRef]
  28. Abdelfattah, A.M.H.; Atef, D.G. From Recognition to Reliability: A Framework for Trustworthy Multimodal AI in Cultural Heritage (The MuseePal Case). IADIS Int. J. WWW/Internet 2025, 23, 81–92. [Google Scholar] [CrossRef]
  29. CIDOC. Definition of the CIDOC Conceptual Reference Model (Version 7.1.3). 2024. Available online: https://cidoc-crm.org/Version/version-7.1.3 (accessed on 20 February 2026).
  30. Hogan, A.; Blomqvist, E.; Cochez, M.; D’amato, C.; Melo, G.D.; Gutierrez, C.; Kirrane, S.; Gayo, J.E.L.; Navigli, R.; Neumaier, S.; et al. Knowledge Graphs. ACM Comput. Surv. 2021, 54, 1–37. [Google Scholar] [CrossRef]
  31. Hyvönen, E. Digital Humanities on the Semantic Web: Sampo Model and Portal Series. Semant. Web. 2023, 14, 729–744. [Google Scholar] [CrossRef]
  32. Agostino, D.; Costantini, C. A measurement framework for assessing the digital transformation of cultural institutions: The Italian case. Meditari Account. Res. 2022, 30, 1141–1168. [Google Scholar] [CrossRef]
  33. Mason, M. The contribution of design thinking to museum digital transformation in post-pandemic times. Multimodal Technol. Interact. 2022, 6, 79. [Google Scholar] [CrossRef]
  34. Dal Falco, F.; Vassos, S. Museum experience design: A modern storytelling methodology. Des. J. 2017, 20, S3975–S3983. [Google Scholar] [CrossRef]
  35. Hosseini, S.; Seilani, H. The role of agentic AI in shaping a smart future: A systematic review. Array 2025, 26, 100399. [Google Scholar] [CrossRef]
  36. ICOM. Museum Definition; International Council of Museums: Paris, France, 2022. [Google Scholar]
  37. Peffers, K.; Tuunanen, T.; Rothenberger, M.A.; Chatterjee, S. A design science research methodology for information systems research. J. Manag. Inf. Syst. 2007, 24, 45–77. [Google Scholar] [CrossRef]
  38. Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Ichter, B.; Xia, F.; Chi, E.; Le, Q.V.; Zhou, D. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In Proceedings of the 36th International Conference on Advances in Neural Information Processing Systems (NeurIPS 2022), New Orleans, LA, USA, 9–28 December 2022; Volume 35, pp. 24824–24837. [Google Scholar]
  39. Venable, J.; Pries-Heje, J.; Baskerville, R. FEDS: A framework for evaluation in design science research. Eur. J. Inf. Syst. 2016, 25, 77–89. [Google Scholar] [CrossRef]
  40. Moura, J. CrewAI: Framework for orchestrating role-playing autonomous AI agents. 2025. Available online: https://github.com/crewAIInc/crewAI (accessed on 20 February 2026).
  41. Bender, E.M.; Gebru, T.; McMillan-Major, A.; Shmitchell, S. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, Virtual Event, 3–10 March 2021; pp. 610–623. [Google Scholar]
  42. Weidinger, L.; Mellor, J.; Rauh, M.; Griffin, C.; Uesato, J.; Huang, P.-S.; Cheng, M.; Glaese, M.; Balle, B.; Kasirzadeh, A.; et al. Ethical and social risks of harm from Language Models. arXiv 2021, arXiv:2112.04359. [Google Scholar]
  43. Wu, X.; Xiao, L.; Sun, Y.; Zhang, J.; Ma, T.; He, L. A Survey of Human-in-the-loop for Machine Learning. Future Gener. Comput. Syst. 2022, 135, 364–381. [Google Scholar] [CrossRef]
Figure 1. The hierarchical architecture of the IMAF, showing the vertical control loop and horizontal interaction axis.
Figure 1. The hierarchical architecture of the IMAF, showing the vertical control loop and horizontal interaction axis.
Systems 14 00954 g001
Figure 2. An example knowledge graph.
Figure 2. An example knowledge graph.
Systems 14 00954 g002
Figure 3. Sequence Diagram for Scenario 1: Visitor Engagement.
Figure 3. Sequence Diagram for Scenario 1: Visitor Engagement.
Systems 14 00954 g003
Figure 4. Sequence Diagram for Scenario 2: Staff Support.
Figure 4. Sequence Diagram for Scenario 2: Staff Support.
Systems 14 00954 g004
Table 1. Comparative positioning of prior research streams.
Table 1. Comparative positioning of prior research streams.
Research StreamRepresentative StudiesMain ContributionRemaining Limitation
General LLM-based autonomous agent architecturesReAct [2]; Weng [1]; Generative Agents [9]; AutoGen [11]; Wang et al. [10]Reasoning–acting loop; memory; planning; reflection; tool use; multi-agent coordinationDomain-general; limited treatment of institutional authority, cultural sensitivity, public accountability, and operational constraints
Museum LLM applications for documentation and accessReusens et al. [15]; Weaver et al. [16]Collection description; transcription; archival access; question answeringTask-specific; focused on information processing rather than autonomous, multi-step, governance-aware behavior
Visitor-facing museum AI and conversational guidanceWang and Matviienko [17]; Ho et al. [18]; Ariya et al. [19]; Vasic et al. [20]Artwork interpretation; chatbot guidance; personalized narratives; virtual toursVisitor-centric; limited support for staff workflows, institutional decision support, and policy-constrained execution
Context-aware recommendation and smart museum systemsTrichopoulos et al. [21]; Ferrato [22]Personalization; spatial context; visitor preferences; situational recommendationContext mainly used for service personalization; limited integration of ethical, rights-related, conservation, and institutional constraints
Multi-agent and multimodal systems in cultural heritageSu et al. [23]; Aragon and D’Haro [24]; Abdelfattah and Atef [28]Multi-perspective interpretation; multimodal interaction; RAG-based guidance; agent coordinationPrototype- or concept-oriented; limited specification of a museum-wide architecture linking governance, memory, planning, and execution
Cultural heritage knowledge representation and semantic infrastructureCIDOC CRM [29]; Hogan et al. [30]; Hyvönen [31]Semantic entities and relations; provenance; knowledge graphs; ontology-based retrievalStrong knowledge structuring, but limited guidance on agent planning, action, communication, and constraint enforcement
Museum digital transformation and socio-technical designParry [12]; Marty [14]; Agostino and Costantini [32]; Mason [33]; Dal Falco and Vassos [34]Institutional mission; visitor experience; digital readiness; human-centered designStrong socio-technical perspective, but limited architectural specification for agentic AI and governed autonomy
Table 2. The Intelligent Museum Agent Framework.
Table 2. The Intelligent Museum Agent Framework.
ComponentRoleSub-Components (and Examples)
Brain (LLM-based Core)Orchestrates the system and processes user requests.LLM Core Engine (reasoning unit); Persona and Tone Controller (institutional voice); Dual Interface (visitor and staff portals).
ContextProvides institutional and operational constraints for agent behavior.Institutional Context: interpretation and voice, conservation and handling, rights and licensing, etc. Operational Context: spatio-temporal availability, infrastructure and asset status, capacity and flow logistics, etc.
MemoryStores and retrieves information required for reasoning and personalization.Perceptual Buffer (sensory input); Active Interaction Context (short-term interaction history); Heritage and Archive layer (knowledge graph and exhibition records).
Planning (Reasoning for complex tasks)Decomposes complex requests and constructs execution workflows.Task decomposition; Chain-of-Thought (CoT) prompting; Self-evaluation and refinement.
Action (Tool integration)Executes plans through interaction with internal and external systems.Code Interpreter (data analysis); Internal API Hub (e.g., CMS, Internet-of-Things (IoT), ticketing); External API Hub (e.g., databases, weather, social platforms).
Table 3. Representative categories of institutional context.
Table 3. Representative categories of institutional context.
CategoryDescriptionPractical Example
Interpretation and VoicePrinciples governing the scope and style of explanations to ensure factual accuracy and avoid unauthorized speculation.Restricting speculative claims in provenance descriptions
Conservation and HandlingOperational rules regarding the physical safety and preservation of heritage objects within various environmental settings.Prohibiting activities that risk physical contact or damage
Rights and LicensingLegal and policy-based restrictions related to the reproduction, distribution, and reuse of cultural data.Managing access levels for high-resolution digital assets
Ethics and SensitivityNormative frameworks for addressing sensitive historical or cultural content with appropriate contextual perspectives.Implementing multi-vocal narratives for contested histories
Accessibility and InclusionRequirements for inclusive communication to ensure information is accessible across diverse linguistic and cognitive needs.Generating simplified or multi-language content summaries
Table 4. Representative categories of operational context.
Table 4. Representative categories of operational context.
CategoryDescriptionPractical Example
Spatio-temporal AvailabilityReal-time constraints related to the physical accessibility of spaces and objects within specific operating hours.Adjusting routes based on temporary gallery closures or seasonal hours
Infrastructure and Asset StatusThe functional condition and maintenance state of technical systems, equipment, and digital touchpoints.Redirecting visitors when specific interactive displays or elevators are out of service
Capacity and Flow LogisticsDynamic limits on visitor density and the scheduling of timed-entry activities to ensure orderly movement.Modulating recommendations to avoid overcapacity in small exhibition rooms
Table 5. Tool categories and examples.
Table 5. Tool categories and examples.
Category Functional RoleExamples
Code InterpreterExecutes computational and analytical tasks that support planning and decision-makingVisitor flow analysis; exhibition layout simulation; statistical processing of attendance data
Internal API Hub Provides access to institution-controlled knowledge and operational systemsCMS; gallery occupancy (IoT) sensors; facility management system; ticketing or reservation platform
External API Hub Extends institutional knowledge and situational awareness through external servicesEuropeana API (https://www.europeana.eu/en/), accessed on 20 February 2026; The Metropolitan Museum of Art Collection API (https://metmuseum.github.io/), accessed on 20 February 2026; Getty Vocabularies (https://www.getty.edu/research/tools/vocabularies/), accessed on 20 February 2026; OpenWeather API (https://openweathermap.org/api), accessed on 20 February 2026
Table 6. Design Objective–Criterion Traceability Matrix.
Table 6. Design Objective–Criterion Traceability Matrix.
Design Objective
(from Section 4.1)
Evaluation Criterion
(from Section 6.1)
Supporting Architectural Mechanism
Academic integrity & information reliabilityReliabilityHeritage & Archive layer; relation-based retrieval/grounded responses
Governance/policy complianceAlignmentBifurcated Context (Institutional + Operational) as regulator
Situational awareness of operational conditionsAdaptabilityOperational Context guiding Planning (occupancy/facility status)
Role diversity (visitor & staff use cases)VersatilityDual Interface + role-sensitive interaction mode
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Ahn, H.J. An Intelligent Museum Agent Framework (IMAF): A Design Science Research Approach to Agentic AI in Museums. Systems 2026, 14, 954. https://doi.org/10.3390/systems14080954

AMA Style

Ahn HJ. An Intelligent Museum Agent Framework (IMAF): A Design Science Research Approach to Agentic AI in Museums. Systems. 2026; 14(8):954. https://doi.org/10.3390/systems14080954

Chicago/Turabian Style

Ahn, Hyung Jun. 2026. "An Intelligent Museum Agent Framework (IMAF): A Design Science Research Approach to Agentic AI in Museums" Systems 14, no. 8: 954. https://doi.org/10.3390/systems14080954

APA Style

Ahn, H. J. (2026). An Intelligent Museum Agent Framework (IMAF): A Design Science Research Approach to Agentic AI in Museums. Systems, 14(8), 954. https://doi.org/10.3390/systems14080954

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop