1. Introduction
The heritage museum is not a neutral information repository. As a public space that interprets cultural heritage, it shapes visitors’ circulation patterns, attention, and memory through carefully designed routes, stopping points, and viewing areas distributed across different spatial locations. A growing body of research shows that a museum’s spatial configuration and exhibition layout structure the content and sequence of visitors’ experiences, thereby influencing narrative comprehension and overall visit quality [
1]. These findings emphasize that exhibition layout is a direct point of intervention for enhancing the visitor experience in cultural heritage museums.
However, in heritage museums, “layout” means far more than the efficiency of the floor plan or visitor circulation. It constitutes an integrated system uniting exhibits, spatial organization, and historic architectural fabric. At the exhibit level, objects function as both narrative anchors and focal points. Their density, scale, accessibility, and grouping can either reinforce or disperse visitors’ attention, thereby creating opportunities for comparison and meaning-making. At the spatial level, circulation structure, node–corridor patterns, transitional spaces, and sightlines jointly shape the distribution of attention and the temporal rhythm of movement between galleries. These elements can not only establish coherent narrative cues, but also generate fragmented experiential episodes [
1]. Architecturally, a heritage museum differs from the “white cube” because the building itself often participates in interpretation through materiality, light, acoustics, and ritualized circulation paths. At the same time, authenticity requirements, conservation standards, structural load limits, and operational safety constraints restrict the scope of interventions [
2,
3]. Consequently, immersive experience in heritage museums more often emerges from the synergy among exhibit organization, spatial order, and architectural atmosphere/constraints, rather than from digital technology alone.
With the broader shift in museums from “display” to “experience”, this challenge has become increasingly salient. Audiences now place greater value on emotional resonance, participatory interaction, and memorable touchpoints. Contemporary museum research highlights multi-sensory experience and embodied participation, while also noting that the supporting evidence remains scattered across disciplinary traditions and practice contexts [
4]. Research on servicescapes and environmental experience likewise indicates that the physical environment shapes emotional responses, cognitive evaluations, and behavioural intentions, making the built environment central to experience creation [
5]. In heritage museums, immersive experience must balance perceived authenticity and preventive conservation with environmental trade-offs (for example, between visual quality/lighting levels and conservation risk) [
2,
3,
6]. The specificity of heritage implies that “immersive experience” should not be treated as a straightforward technological upgrade; instead, it requires a comprehensive, architecture-informed approach to layout decision-making.
Studies of immersive experience often draw on concepts such as immersion/immersive perception and participation, which were primarily developed in the context of mediated technologies [
7,
8]. In museum and heritage settings, digital tools (e.g., augmented/virtual/mixed reality) can strengthen interpretive narratives and participation, and systematic reviews indicate rapid development in such applications [
9]. Nevertheless, spatial layout remains fundamental: even when immersive technologies are introduced, the structure of the visit still depends on route configuration, visual relations, congestion nodes, and the curatorial logic of object sequencing [
10]. Therefore, the challenge for heritage museums is not only to adopt immersive technologies, but also to coordinate exhibit presentation, spatial order, and architectural constraints through layout decision-making, in order to create a credible and meaningful visiting experience that meets cultural heritage conservation requirements.
Despite these advances, there are still two major limitations at the intersection of heritage museums, exhibition layout, and immersive experience. First, existing studies often decompose the drivers of immersion into isolated themes such as digital technologies, atmosphere creation, or individual spatial variables, while lacking a layout-oriented perspective. As a result, they fail to integrate exhibit deployment, space utilization, and building-imposed constraints under the requirements of authenticity and conservation [
6,
9,
11]. Second, the most existing immersion measurement tools originate from virtual environments or general technology experience scales, making it difficult to capture heritage-specific considerations, including perceived authenticity, conservation boundaries, and the interpretive role of architecture. These factors determine the feasibility and legitimacy of interventions in heritage contexts [
2,
3]. In response to these gaps, there is a need to establish a layout-centred and design-oriented indicator system for intervention. Such a system should be hierarchical, comparable, diagnostic, and optimizable, so that “immersive experience” can be translated into operable spatial and curatorial variables, thereby providing a basis for design decision-making and benchmark evaluation.
To address these shortcomings, this study develops a heritage museum exhibition layout evaluation framework with immersive experience as its core. The framework translates immersive experience into a quantifiable diagnostic indicator system while explicitly incorporating the constraints of authenticity and conservation. Given that layout-related factors shaping immersive experience in heritage museums span architecture, exhibition design, heritage conservation, and interpretation research, and that relevant terms are inconsistent with unclear conceptual boundaries, this study applies the Delphi expert consensus method to generate and refine an indicator system applicable to design evaluation [
12]. After the indicators are established, a transparent multi-level weighting structure (dimension, sub-dimension, indicator) is constructed by applying the Group Analytic Hierarchy Process (GAHP) with pairwise comparisons and consistency testing, thereby providing a defensible priority structure for design decision-making. This method is suitable for group decision contexts and treats the expert panel as a unified decision-making unit [
13,
14]. Finally, Importance–Performance Analysis (IPA) is applied to transform the weighted “importance–fit” results into actionable priorities for layout optimization and resource allocation [
15]. A case study further verifies the applicability of the framework.
Therefore, the objectives and problems of this study are as follows:
Research Objectives
Identify core elements of immersive experience-oriented layout design in heritage museums and translate them into actionable evaluation criteria;
Construct a hierarchical indicator system and determine weights for each level to support benchmarking and layout optimization;
Using the constructed evaluation system to assess heritage museums.
Research Questions
What key elements and operational standards influence immersive experiences in heritage museum exhibition layouts?
How can a hierarchical indicator system be constructed and weighted to support design prioritization and optimization?
How does one use the constructed evaluation system to assess heritage museums?
The introduction clarifies the problem framework and research gaps.
Section 2 systematically reviews immersive experience constructs and layout determinants for heritage museums under authenticity and conservation constraints.
Section 3 details the Delphi–group AHP–IPA workflow and metric development methodology.
Section 4 presents case study findings and diagnostic insights.
Section 5 and
Section 6 discuss implications for layout optimization, research limitations, and future directions.
2. Literature Review
This literature review clarifies the theoretical framework of immersive experience and explains how adjacent concepts are applied within the proposed layout-operation framework, thereby avoiding conceptual overlap. It outlines the key pathways through which exhibition layout, together with architectural and environmental conditions, shapes immersive experience in heritage museums. It also reviews research on digital twins, artificial intelligence (AI), and immersive technologies in domains such as heritage conservation and cultural tourism, while identifying the lack of layout-oriented and design-driven assessment approaches. In response to these methodological gaps, this review proposes the development of a hierarchical, weighted, and diagnosable layout-operation framework. The framework will draw on the Delphi method, the Group Analytic Hierarchy Process (GAHP), and the importance–performance logic.
2.1. The Concept of Immersive Experience
The theory of immersive experience primarily originates from research on virtual environments and human–computer interaction and is articulated through a series of related concepts. A fundamental distinction is commonly made between system-level immersion, namely the objective extent to which an environment envelops the senses and interaction channels, and user-level experiential states, such as presence, engagement, and flow.
Immersion is often defined as an objective condition that facilitates experience by shaping the coherence of sensory stimulation and action possibilities, whereas presence refers to the subjective “being there” experience that arises when attention and interpretive processes resonate with environmental cues [
8]. Flow refers to deep concentration and intrinsic enjoyment that result from a balance between perceived challenges and personal skills, and it is often accompanied by sustained attention and positive memory traces [
16]. However, in heritage museums, immersive experiences rarely derive from technological intensity alone.
2.2. Key Pathways Through Which Exhibition Layout Affects Immersive Experience
Research indicates that spatial configuration and display organization systematically influence visitors’ viewing sequences and movement patterns, thereby shaping the quality and coherence of the museum experience. In heritage museums, layout-related factors most relevant to immersive experience can be summarized into five interrelated pathways that are both designable and readily evaluable.
Circulation networks allocate attention and time across galleries: determining whether the visit unfolds as a coherent whole or as fragmented episodes [
10,
17]. Layout legibility, route continuity, and node–corridor connectivity influence wayfinding difficulty, route choice, and revisit likelihood, thereby affecting the temporal stability of the visiting experience [
1,
10]. Spatial organization also interacts with visitor distribution. Crowding and uneven space utilization can disrupt visiting rhythms and weaken interpretive coherence [
17,
18].
Exhibition layout functions as a narrative mechanism: using spatial sequencing, interpretive nodes, and transitional spaces to coordinate themes and visitors’ cognitive framing [
1,
19]. When spatial sequence aligns with interpretive logic, the “meaning gap” between exhibits can be reduced, supporting cumulative understanding [
19]. Conversely, when spatial sequence conflicts with interpretive logic, even compelling individual exhibits may produce experiential discontinuities for visitors [
1,
10].
Sensory and environmental atmosphere under conservation constraints. Atmosphere extends beyond simple scenography. In heritage museums, atmospheric design is constrained by conservation requirements and building performance. Architectural research indicates that conservation, energy use, and visitor comfort form an interdependent decision space, and environmental settings must reduce risks [
3,
6,
20].
Authenticity, credibility, and sense of place: The judgement of authenticity is the core of the heritage experience; any intervention that conflicts with historical credibility may damage this perception. When introducing immersive or digitally enhanced technologies, this problem is particularly prominent. Although digital reconstruction and immersive scenarios can enhance accessibility and participation, poorly designed implementation schemes may also increase the risks to interpretive reliability and perceived authenticity [
21]. Therefore, when evaluating the immersive layout of a heritage museum, authenticity should be regarded as an operational dimension embedded in spatial interpretation decision-making, rather than an external “constraint”.
Interactivity, participation, and learning support: Participation is influenced by both technology and space. Visibility, reachability, waiting costs, and the spatial distribution of interactive opportunities jointly determine whether interactive experiences are sustained or short-lived. Immersive heritage experiences increasingly integrate on-site and remote participation (for example, virtual reality and augmented reality, or metaverse access), extending the concept of “interaction” from device availability to a comprehensive system covering spatial layout, interpretive frameworks, and operational feasibility [
21,
22].
These five paths reinforce the core argument of this study: the immersive experience of heritage museums arises from the synergy among exhibit organization, spatial sequencing, and the environmental and heritage constraints embodied in the building, rather than simply relying on digital technology.
The 14 secondary dimensions are derived from the literature and further subdivided according to the five layout-operation paths (A–E). They are integrated using Delphi-based coding to ensure clear boundaries and heritage applicability. Our contribution is to integrate the operationalization of layout-operation paths with diagnostic indicators. The definitions and representative sources of all secondary dimensions are summarized in
Table A4.
2.3. Application of Digital Twins, Artificial Intelligence, and Immersive Technology in Heritage Protection and Cultural Tourism
In recent years, the literature on architectural heritage and smart heritage management has rapidly expanded the application of digital twins, artificial intelligence, and immersive technologies in heritage protection and cultural tourism, treating these tools as a means of infrastructural empowerment.
Research on digital twins for heritage buildings reveals a multi-level application pathway, from basic digital representation to adaptive systems integrating 3D data capture, heritage building information modelling (HBIM), the Internet of Things (IoT), artificial intelligence, and immersive interfaces, with a focus on automated asset management, interoperability, and decision support [
23]. At the same time, HBIM-centred research notes that heritage conservation workflows increasingly depend on integrated digital processes, while also emphasizing long-standing obstacles such as modelling complexity, standardization, and cross-tool integration [
24]. A key practical direction is to combine IoT sensing, HBIM-based information management, and digital twin technologies for monitoring and predictive maintenance. Recent studies show that real-time data streams can be linked to digital models, thereby supporting anomaly detection, condition tracking, and scenario-based planning for heritage assets [
25]. Reviews of artificial intelligence applications in cultural heritage also highlight the growing use of machine learning in conservation strategies, predictive maintenance, artefact analysis, and visitor experience enhancement, while stressing the need for ethically guided interdisciplinary integration rather than isolated technology deployment [
22]. Immersive technologies and metaverse-related approaches have moved beyond simple “visualization” and developed into broader cultural tourism and accessibility strategies. Systematic reviews and case studies indicate that metaverse applications can support virtual reconstruction, educational communication, global access, and participatory interaction, while also highlighting unresolved challenges related to authenticity, evaluation methods, and sustainability [
21]. From a museum perspective, these developments suggest that immersive experiences are increasingly shaped by hybrid ecosystems: physical layout and environmental constraints still determine the on-site experience, while the digital layer extends interpretation and interaction beyond architecture.
Despite these advances, there are still key gaps in the interdisciplinary areas addressed in this study, for example, the lack of an operational evaluation system for translating “immersive experience” into actionable layout decisions.
2.4. Constructing an Operationalized Layout Weighting Framework
Existing museum immersion research predominantly follows three pathways:
- (i)
Outcome-oriented experience state surveys;
- (ii)
Technology-centric device/medium evaluations;
- (iii)
Qualitative case studies of narrative and atmospheric qualities. While these methods aid in documenting results, they often struggle to guide which layout levers should be adjusted within heritage constraints.
This study’s methodological requirements encompass two aspects:
Indicator generation and boundary definition under dispersed evidence. When conceptual consensus exists but interdisciplinary operational items remain unstable, Delphi-based expert consensus is widely employed to unify terminology, refine indicators, and control research scope. Structured Delphi processes support systematic project generation and screening, enhancing content validity and reducing redundancy [
12].
Prioritization and optimization require transparent hierarchical weighting methods. To achieve actionable prioritization across dimensions, sub-dimensions, and items, hierarchical weighting approaches are essential. The Analytic Hierarchy Process (AHP) provides a consistent framework for deriving relative priorities through pairwise comparisons[
14]. For group contexts, research on judgement and priority aggregation provides a defensible basis for integrating expert opinions into collective weights—critical for designing frameworks in multi-stakeholder projects [
13]. Therefore, the following are integrated:
- (i)
guided open Delphi method to optimize indicators;
- (ii)
Group AHP to achieve hierarchical weight allocation;
- (iii)
Importance–performance logic for diagnostic priority ranking directly addresses the core research gap: translating immersive experience objectives into an evaluation tool that is operationally implementable, comparable, and possesses optimization potential, thereby adapting it to the constraints and decision-making environment of heritage museums.
3. Methodology
This study constructs a hierarchical indicator system for optimizing exhibition layouts in cultural heritage museums focused on immersive experiences. The research design integrates the following: a two-round Delphi method for indicator extraction and convergence; Importance–Performance Analysis (IPA) for priority diagnosis; and an Analytic Hierarchy Process (AHP) for hierarchical weight allocation and consistency verification [
12,
15,
26,
27]. This integrated logic aligns with the mainstream framework of Multi-Criteria Decision-Making (MCDM): first structuring expert judgments into evaluation indicators, then converting them into stable weights and actionable rankings, often supplemented by robustness testing [
28].
Figure 1 summarizes the application steps of the proposed framework. First, a literature synthesis defines five layout-operational pathways (A–E) and compiles an initial item pool. Second, a guided open Delphi is conducted: Round 1 generates items within the A–E boundaries, and Round 2 rates items on importance and fit to screen indicators using item-level consensus (Median/IQR) and panel-level concordance (Kendall’s W). Third, group AHP (AIJ) is used to derive hierarchical weights for Level 1 and Level 2 dimensions with consistency control on aggregated matrices. Fourth, tertiary indicators are ranked using the CV–entropy approach, and IPA is used to translate importance–fit results into actionable priorities. Finally, the core instrument is applied to a case museum to produce weighted diagnosis and layout-optimization recommendations.
3.1. Expert Panel Formation and Recruitment
This study employed a purposive sampling strategy supplemented by snowball referral to recruit an expert panel with empirical experience in cultural heritage museum exhibition design and immersive visitor experience research [
29,
30]. To ensure interdisciplinary representation, invited experts spanned exhibition curation, architecture/interior design, heritage conservation, museum education, and digital media. Selection criteria adhered to the Delphi method’s general norms:
- (i)
high relevance to the research topic;
- (ii)
≥5 years of professional or research experience;
- (iii)
participation in at least two related projects;
- (iv)
Recent outputs (projects, publications, awards) demonstrating sustained contributions to the field [
12,
26].
The initial candidate list was 15 people; information sheets detailing research objectives, procedures, anonymity safeguards, and time commitments were sent to candidates. Ultimately, 11 experts agreed to participate and complete two rounds of evaluation (see
Table 1). This arrangement aligns with Delphi method principles, emphasizing expert depth over statistical representativeness in professional design decision-making [
31,
32].
3.2. Delphi Round 1: Guided Open-Ended Item Generation
The Delphi method was selected for its effectiveness in structuring and integrating expert judgement across domains characterized by the following: strong evidence heterogeneity, stringent contextual constraints, and high reliance on tacit professional knowledge. It also mitigates leader effects through controlled feedback and anonymity mechanisms [
12,
26,
31]. Thus, the first round employed a guided open-ended format to solicit candidate indicators for exhibition layouts that shape immersive experiences in cultural heritage museums. Experts completed three tasks:
- (a)
Identifying layout-related factors most influencing immersion under conservation and management constraints;
- (b)
proposing quantifiable sub-indicators;
- (c)
suggesting feasible data sources for evaluation.
To enhance content validity and minimize interdisciplinary conceptual drift, the prompt provided a five-dimensional framework (circulation and spatial organization; narrative/contextual immersion; sensory atmosphere; authenticity/sense of place; and interactivity/engagement), requiring experts to provide practical examples. Feedback was coded using open coding and semantic clustering, with redundant or synonymous items merged to form a candidate indicator pool for quantitative convergence testing. This process aligns with Delphi method quality assurance recommendations for coding and integration [
32].
3.3. Delphi Round 2: Importance and Fit Rating
Based on the 39 candidate items formed in Round 1, a structured questionnaire was distributed to the same 11 experts in Round 2. Experts were asked to rate each item on a 5-point Likert scale across two dimensions:
Importance: The importance of the element to the immersive experience (1 = Very Unimportant, 5 = Very Important);
Fit: The compatibility/feasibility of the element within the context of heritage museum exhibitions (1 = Very Unfit/Hard to Implement, 5 = Very Fit/Easy to Implement).
A remarks section was provided for each item, allowing experts to suggest additions, deletions, mergers, or wording revisions.
Upon retrieval of the questionnaires, descriptive statistics (
Mdn,
IQR,
Mean ±
SD,
CV) were calculated. Kendall’s coefficient of
concordance (
W) was used to test the overall agreement of expert ratings [
33]. The consensus criteria and stopping rules were as follows [
32]:
(1) Item-level consensus was determined using Median
(Mdn) and Interquartile Range (
IQR). If an item achieved
Mdn ≥
4 and
IQR ≤
1 in a specific dimension, it was considered to have reached expert consensus. Remaining items were revised based on expert remarks or treated as items requiring further verification. The Coefficient of Variation (
CV) was used solely for describing dispersion and sensitivity analysis, not as a rigid threshold for item screening [
32].
(2) Panel-level agreement and stopping rule Kendall’s W and its significance test on the “importance” dimension served as the primary basis for assessing group agreement. When the importance dimension achieved statistical significance (p < 0.05) and item-level consensus was stable, acceptable agreement was deemed achieved, satisfying the stopping condition. To clarify, Kendall’s W was used as a panel-level concordance check rather than a strict termination threshold, because W tends to be conservative when the expert panel is multidisciplinary and when a relatively large number of items are evaluated simultaneously. Delphi termination therefore followed a combined criterion:
- (a)
statistically significant concordance on the importance dimension;
- (b)
stable item-level consensus (Mdn ≥ 4 and IQR ≤ 1) for a substantial portion of items.
Items that did not meet item-level consensus were not forced into the core indicator set but were retained as provisional modules for subsequent case- and audience-based validation. The “fit” dimension served as a supplementary judgement for subsequent optimization priority analysis and empirical verification. Upon meeting these conditions, no further consultation rounds were conducted. For items failing to reach consensus or containing ambiguity, the research team performed merging, splitting, or wording revisions based on expert opinions, with adjustments detailed in the results section. It is important to note that “fit” ratings were primarily used for joint IPA with importance to locate optimization priorities and did not affect the calculation of group AHP weights.
3.4. AHP Weighting and Consistency Check
Following the Delphi-defined indicator system, the Analytic Hierarchy Process (AHP) was used to derive hierarchical weights. Experts conducted pairwise comparisons within each level using Saaty’s 1–9 scale [
34]. To obtain a collective judgement for group AHP, this study adopted Aggregation of Individual Judgments (AIJs), where individual comparison matrices were aggregated using the geometric mean to form a group matrix at each level [
13,
35]. The priority vector
w was then obtained from the principal eigenvector of each group matrix.
Consistency was examined using the standard AHP indices:
where
is the matrix order,
is the maximum eigenvalue, and
is the random index. A threshold of
CR < 0.10 was adopted as acceptable [
34]. Because the objective is to obtain group-level weights, consistency was enforced on the AIJ-aggregated group matrices; all group matrices satisfied
CR < 0.10 (
Level 1: CR = 0.0226;
Level 2: CR ≤ 0.0356). After consistency was confirmed, global weights were computed from top to bottom and indicators were ranked accordingly.
To test the robustness of the prioritization, priorities were additionally computed using Aggregation of Individual Priorities (AIPs) and compared with the AIJ results. The rankings were highly consistent (Spearman’s = 0.900 at Level 1, < 0.05), indicating that the weight structure is stable and independent of the choice of aggregation approach.
3.5. Case Application and Feasibility Demonstration: Yinxu Museum
To translate Delphi method evaluation values into actionable optimization priorities, this study introduces the Importance–Performance Analysis (IPA) method. This method constructs a two-dimensional matrix of importance and performance values, categorizing evaluation indicators into four quadrants: “Keep up the good work”, “Concentrate here”, “Low priority”, and “Possible overkill” [
15]. This decision support model intuitively translates evaluation metrics into differentiated improvement strategies and logical sequencing.
Finally, the constructed weighted indicator system is applied to the exhibition layout of the Yinxu Museum for empirical analysis. This case-based validation approach aligns with the mainstream paradigm in museum layout and navigation research, which tests evaluation frameworks within real physical spatial configurations and visitor behaviour constraints [
10]. Contextualized assessment through case application verifies the operational feasibility of high-weight indicators under practical demands, such as artefact preservation, security, and visitor flow management, thereby enhancing the explanatory power and external validity of the evaluation system.
4. Results
4.1. Delphi Round 1: Guided Open-Ended Text Consolidation Results
In the first guided open-ended Delphi round, 11 experts proposed exhibition layout elements influencing heritage museum immersive experiences through open-ended responses under five predefined primary frameworks: A. Circulation and spatial organization; B. Narrative and contextual immersion; C. Sensory and atmosphere (light/sound/materials); D. Authenticity and sense of place; E. Interactivity and participation. The research team coded the open-ended texts, performed semantic grouping and synonym merging, forming a candidate framework comprising 5 primary dimensions, 14 secondary dimensions, and 39 candidate tertiary entries. This framework provides an entry pool and terminological boundaries for the second round of quantitative convergence. The comprehensive hierarchical framework, including all 39 candidate tertiary indicators and their semantic definitions, is detailed in
Table A2.
4.2. Delphi Round 2: Convergence Results for Importance and Fit Scores
Round 2 collected 5-point Likert ratings (importance and fit) for 39 candidate items from 11 experts. Item-level consensus was defined as Mdn ≥ 4 and IQR ≤ 1.
Convergence classification. Of the 39 items, 20 (51.3%) reached consensus on both importance and fit (retained), 14 (35.9%) reached consensus on one dimension only (provisional), and 5 (12.8%) did not reach consensus on either dimension (removed).
Panel-level agreement. Kendall’s W indicated weak but significant agreement for importance (
W = 0.185,
χ² = 77.42,
df = 38,
p < 0.001), whereas fit did not reach statistical significance (
W = 0.119,
χ² = 49.75,
df = 38,
p = 0.096). This pattern suggests that experts align more readily on “what matters” than on “what is feasible/applicable” across heritage museum contexts; fit therefore requires further validation through case application and audience-based evidence (see
Table 2).
Notably, all retained core indicators satisfied the item-level consensus thresholds (
Mdn ≥ 4 and
IQR ≤ 1), supporting stable convergence at the item level despite the weak overall
Kendall’s W. The core indicators retained after the second round of Delphi surveys are shown in
Table 3.
4.3. AHP Results: Weights of Level 1 and Level 2 Dimensions
Group AHP produced stable weights with acceptable consistency (CR < 0.10 for all group matrices).
Level 1 weights. B. Narrative and scenography (0.2877) > A. Circulation and spatial organization (0.2281) > C. Sensory and atmosphere (0.1981) > D. Authenticity and sense of place (0.1644) > E. Interactivity and participation (0.1217).
Level 2 global priorities. The top-ranked Level 2 dimensions were B1 narrative signposting (0.1248), A2 spatial rhythm and articulation (0.1073), and B2 scenographic time–space context (0.1065), followed by C1 lighting quality (0.0910) and C2 acoustic environment (0.0723). Overall, the weighting emphasizes a “narrative–circulation–atmosphere” backbone for immersion-oriented exhibition layouts.
The Level 2 dimensions reported in
Table 4 were first abstracted from prior museum experience and exhibition layout studies along the five layout-operational pathways (A–E), and then clarified and boundary-checked through Delphi-based coding and expert consensus. Therefore, the “Top Level 2 dimensions” should be interpreted as empirically weighted representations of established theoretical themes (see
Table A4 for definitions and representative sources).
The AHP weighting results indicate that narrative and contextual design (B, 0.2877) and circulation and spatial organization (A, 0.2281) collectively account for over 50% of the cumulative weight. This demonstrates that, within the heritage museum context, the integrity of narrative logic and the organizational capacity of spatial pathways form the foundation of immersive experiences. Among secondary weights, B1 narrative guidance and A2 spatial rhythm ranked highest, further underscoring the critical role of clear narrative threads and well-paced spatial variations in enhancing visitor engagement. This weight distribution aligns closely with the conclusions drawn by experts during the Delphi phase, collectively establishing an evaluation logic anchored by the “narrative–pathway–atmosphere” framework. It positions interactivity (E) as a supporting element, ensuring experiential depth grounded in authenticity.
4.4. Case Application: Yinxu Museum
The Yinxu Museum is located in Anyang, Henan Province, and is closely associated with Yin Xu, the late Shang Dynasty capital site that was inscribed on the UNESCO World Heritage List in 2006. In 2024, the museum’s new building opened to the public, presenting the Shang civilization through large-scale permanent and thematic exhibitions combined with digital interpretation (see
Figure 2).
Three experts were invited to conduct the on-site scoring in the case application. Selection criteria were as follows:
- (i)
At least 8 years of professional or research experience related to museum exhibition/layout, conservation, or visitor studies;
- (ii)
Demonstrated involvement in at least two heritage museum exhibition projects (planning, design, evaluation, or operation);
- (iii)
Complementary expertise to ensure coverage of spatial/layout, curatorial–interpretive narrative, and visitor/education–experience perspectives;
- (iv)
Familiarity with heritage museum constraints (authenticity requirements, conservation limits, and solemn/ritual spatial order). All three evaluators received the same scoring rubric and performed an on-site walkthrough before completing the independent ratings for observation space.
Using the 20-item core instrument, three experts independently rated each item on a 1–5 performance scale (1 = clearly insufficient; 3 = acceptable/baseline; 5 = excellent). Dimension scores (A–E) were calculated as the mean of item scores within each dimension. An overall performance score was computed by combining dimension scores with the Level 1 AHP weights (
A = 0.2281,
B = 0.2877,
C = 0.1981,
D = 0.1644,
E = 0.1217):
where
is the Level 1 weight and
is the mean score (1–5) for Level 1 dimension k. When item-level weights are used, the equivalent form is as follows:
Table 5 reports the dimension-level results and overall weighted score for the Yinxu Museum case application. The mean overall weighted score was 4.36/5, indicating a generally strong performance under the immersion-oriented layout framework.
At the dimension level, the highest mean scores were observed for C (sensory and atmosphere, 4.60) and D (authenticity and sense of place, 4.53), suggesting that lighting/acoustic conditions and authenticity-related interpretive boundaries were perceived as consistently strong. E (interactivity and participation) showed the lowest mean score (4.00), indicating relatively greater room for improvement in interactive usability, capacity, and inclusive participation conditions.
As shown in
Figure 3,the item level, the highest-rated items (mean ≥ 4.67) included “Object–space resonance”, “Illuminance/uniformity and glare control”, “Colour rendering under conservation constraints”, and “Clarity of replicas/reconstructions”. The lowest-rated items (mean = 3.67) were “Narrative callbacks” and “Waiting time and participation capacity”, which together suggest two pragmatic optimization directions:
- (i)
Improving narrative linkage and cross-referencing between sub-themes and the main storyline;
- (ii)
Strengthening queue/capacity management and interactive throughput during peak visitation.
This study uses the mean importance and mean performance () as the cross-boundary lines for the IPA matrix, mapping the 20 evaluation indicators into four quadrants. By examining the distribution patterns of indicators across each quadrant, we can identify the core strengths, potential weaknesses, and areas for resource allocation optimization in the exhibition design of the Yinxu Site Museum.
In the Strengths Quadrant, C1-1 (Illumination level and glare control), C1-2 (Colour rendering and artefact protection), and C2-2 (Speech intelligibility) stand out most prominently. This indicates highly successful foundational physical environment construction, particularly in lighting quality control, artefact preservation, and clear auditory information transmission. These aspects not only receive high expert recognition but also demonstrate their critical importance, forming the core competitiveness of the current exhibition design.
Within the improvement areas, A2-2 (Visual guidance (landmarks and vistas)), B3-2 (Sense of ritual and commemorativeness), and C2-3 (Directional sound and soundscape design) all exhibit characteristics of high importance but low performance values. This indicates that audiences still hold higher expectations for intuitive spatial orientation, the ceremonial experience at narrative junctures, and the nuanced handling of immersive soundscapes. These areas represent key breakthrough directions for future design optimization and enhancing exhibition depth.
Regarding potential overkill, the performance scores for D1-2 (Clarity of replicas and reconstructions), B2-2 (Object–space correspondence), and D1-3 (Minimization of display interference) significantly exceed their importance averages. This indicates that current exhibition practices may already exceed expert expectations in areas like clear labelling of replicas, physical alignment between exhibits and spaces, and elimination of visual distractions. Resource allocation in these areas could be appropriately streamlined to avoid design redundancy.
Within the low priority zone, indicators such as B1-2 (Narrative coherence of sub-themes), E1-3 (Waiting time and capacity), and D3-3 (Spatial pauses and contemplative atmosphere) received relatively low importance ratings and demonstrated performance below average levels. This suggests that, within the current evaluation context, narrative coherence of sub-themes, waiting times for interactive facilities, and immersive contemplative spaces have not yet become core issues. Given limited resources, these indicators can remain as they are, managed as foundational supporting elements.
Figure 4 compares the Yinxu Museum’s dimension-level performance with the relative strategic importance implied by the Level 1 AHP weights. Strategic importance derived from AHP weights was rescaled to a 3.5–5 range using min–max normalization to enable direct visual comparison with performance scores on the same radar chart.Overall performance is strong across all five dimensions, with particularly high scores in sensory and atmosphere (C) and authenticity and sense of place (D). By contrast, the AHP profile assigns greater relative importance to narrative and scenography (B) and circulation and spatial organization (A) than is reflected in the observed performance, suggesting these two dimensions may warrant prioritized refinement to better align performance with strategic importance. Interactivity and participation (E) shows comparatively lower performance but also the lowest strategic weight, indicating a lower priority for intervention unless stakeholder goals shift toward participation-oriented outcomes.
A one-at-a-time sensitivity analysis was conducted to test robustness of the prioritization to weight perturbations. Using the baseline Level 1 weights derived from group AHP (AIJ), we doubled one Level 1 weight at a time and proportionally rescaled the remaining weights to keep the total equal to 1. The
Table 6 reports the resulting Level 1 ranking changes. Under the baseline, the Level 1 priority order is B > A > C > D > E. As expected, doubling a given dimension moves it upward; however, the overall priority structure remains interpretable and stable in the sense that the top tier consistently involves the core layout-operational dimensions (narrative/scenography, circulation/spatial organization, and sensory/atmosphere), while authenticity and participation increase only when explicitly emphasized by the perturbation.
The one-at-a-time weight-doubling tests indicate that the Level 1 ordering is generally stable and changes in a direction consistent with the imposed perturbation. Under the baseline scenario, the rank order is B > A > C > D > E. When a single dimension is doubled, the doubled dimension rises to the first or second position (e.g., A > B > C > D > E when A is doubled; C > B > A > D > E when C is doubled; D > B > A > C > E when D is doubled), while B remains within the top two positions across all scenarios and E only rises substantially when its weight is doubled (B > E > A > C > D). Overall, the sensitivity results suggest that the prioritization is not driven by a single dimension and remains interpretable under plausible changes in stakeholder emphasis.
5. Discussion
This study develops an “immersive experience-oriented” exhibition layout evaluation system for cultural heritage museums. The indicator set was refined and weighted by combining the Delphi method with the Analytic Hierarchy Process (AHP). Results from Round 2 of the Delphi survey indicate that, among 39 candidate indicators, 20 (51.3%) reached consensus on both “importance” and “applicability” and therefore formed the core set; 14 (35.9%) reached consensus on only one dimension and were retained as an optional extension module; and 5 (12.8%) failed to reach consensus and were excluded. Overall consistency tests suggest a weak but statistically significant agreement among experts on the “importance” dimension (Kendall’s W = 0.185, p < 0.001), while no significant consensus was observed for “applicability” (W = 0.119, p = 0.096). This pattern implies that experts more readily agree on the “core elements”, yet differ substantially when judging “adaptability across heritage environments”. Accordingly, this study treats consensus on “importance” as the basis for framework stability, while using “adaptability/applicability” as a priority dimension to support calibration in subsequent case studies and audience-based evidence.
Weighting results further reinforce the core chain of “narrative–path–atmosphere.” At the main-dimension level, narrative ranks highest for situational immersion (0.2877) and for cyclical/spatial organization (0.2281), followed by sensory/atmospheric factors (0.1981), authenticity/sense of place (0.1644), and interactive/participatory elements (0.1217). This suggests that immersive experience is not produced by a single technical or sensory factor; instead, it relies on the synergy between the organization of narrative meaning and the organization of spatial behaviour, and is strengthened by environmental conditions such as lighting, sound, and special effects. In practical use, the framework can serve as a rapid diagnostic and priority-setting tool for new-build or renovation projects in heritage museums.
The main limitations are threefold:
Adaptive consistency remains insufficient; the weighting still largely reflects expert judgement; and the framework’s utility is demonstrated only through a single-museum case. Future research should incorporate multiple museum samples and visitor data to validate the framework and recalibrate weights, thereby strengthening external validity.
A further limitation of the IPA diagnosis is that it uses “fitness” as a proxy for performance. Traditional Importance–Performance Analysis (IPA) typically measures performance through service outcomes or user satisfaction [
15]. However, this approach has been criticized for conceptual and measurement ambiguity, especially when “importance” and “performance” are operationalized in simplified or substitute forms, which can undermine the validity of attribute classification [
45,
46]. In this study, IPA is treated primarily as a design diagnostic tool under heritage constraints. Its key operational concern is the current degree of realization and the compatibility/feasibility of each indicator in a specific museum context (e.g., authenticity requirements, conservation restrictions, and the spatial order of ritual/commemorative functions). This proxy-based operationalization may influence quadrant placement. If performance were instead assessed through on-site observations (e.g., navigation errors, dwell time, congestion/queuing) or visitor feedback on satisfaction/immersion, some indicators could shift across quadrants. Future work should combine expert applicability assessments with behavioural measures and visitor experience outcomes to triangulate IPA findings, aligning with the post-occupancy evaluation (POE)-oriented approaches increasingly emphasized in museum environments [
39].
Another limitation is that the same expert panel participated in both indicator optimization (Delphi) and weight derivation (AHP), which may introduce confirmation bias and path dependency. Although task separation, aggregated feedback, and AHP consistency checks help mitigate this risk, the final weights should still be interpreted as an expert-derived decision structure. This issue can be particularly salient in group AHP, where the choice of aggregation method and the pattern of inconsistencies may affect group priority rankings. Methodological work has distinguished Aggregating Individual Judgments (AIJs) from Aggregating Individual Priorities (AIPs) and clarified their respective decision-context assumptions [
13,
47]. Recent studies also suggest that embedding Delphi-style feedback (and “nudge” mechanisms) within group AHP can reduce inconsistency during iterative consultation and mitigate judgement bias [
48]. Future research should follow robustness recommendations in multi-criteria decision analysis (MCDA) [
49] by validating or recalibrating weights using independent expert panels and visitor/field evidence (e.g., on-site behavioural observation and visitor experience/satisfaction metrics). Sensitivity analyses, such as elimination tests and systematic weight-perturbation tests, should also be conducted to assess ranking stability.
Compared with previous studies that mainly focussed on visitor experience evaluation or qualitative exhibition analysis, this study clearly regards museum architecture as a spatial system(see
Table 7). Although early studies often explore immersion or narrative effects, few studies provide a layout-operation evaluation framework that can support spatial diagnosis and optimize decision-making. The hierarchical index system proposed in this study fills this gap and transforms experience goals into spatial operation standards, such as process organization, spatial sequence, and spatial utilization efficiency.
From the perspective of architectural science, the weight analysis results show that the organization of narrative structure and dynamic line paths has a greater impact than isolated interactive elements. This shows that the immersive experience of museum architecture is fundamentally constrained by the spatial layout. This finding supports the view that experiential design should be harmonized with architectural layout decision-making, and should not just be regarded as a curatorial issue. At the same time, the recent Multi-Criteria Decision-Making (MCDM) research based on hierarchical analysis (AHP) in other applied engineering fields also emphasizes that the sound priority depends on the reliable evaluation standard system and the weight/stability test of the system [
28].
6. Conclusions
This study builds an immersive experience-oriented heritage museum exhibition layout evaluation framework, and establishes a corresponding weight structure. Through the two-round guided open Delphi method, 20 core entries (accounting for 51.3%) were screened out of 39 candidate entries. The Kendall’s W value shows that the consensus of experts on “importance” is weak but statistically significant (W = 0.185, p < 0.001), while the consensus on “applicability” is insufficient (W = 0.119, p = 0.096), which indicates the need for case study. The feasibility of the framework is further verified by the investigation and audience data. Hierarchical analysis (AHP) reveals the priority of major dimensions: Narrative and situational immersion (0.2877) > Process and spatial organization (0.2281) > Sensory experience and atmosphere (0.1981) > Authenticity and sense of place (0.1644) > Interaction and participation (0.1217), emphasizing that the “narrative–flow–atmosphere” chain is the key foundation of immersive experience design.
The single-hall application of the Yinxu Museum further verified the operational feasibility of the system (weighted total score average score: 4.36/5), and put forward operational priorities at the dimension and project levels. Future research will be extended to multiple museum samples and incorporated visitor data to verify the adaptive dimension and calibrate the weight, so as to enhance the universality and interpretability of the system.
The proposed index system and weight results provide a reproducible benchmark for the subsequent research on the architectural environment on the spatial performance of the museum. Future research can be carried out from the following four directions:
- (1)
External verification and calibration: Use data from multiple heritage museums and subsequent usage data (such as visitor surveys, behaviour mapping, stay time tracking, and congestion records) to test and recalibrate indicator weights to enhance their universality beyond expert judgement.
- (2)
Integration of spatial diagnosis model: This framework is combined with quantitative spatial analysis methods to transform qualitative standards into measurable spatial variables.
- (3)
Optimization and simulation: Integrate weighted indicators into the scenario-based layout optimization model. Combined with crowd simulation and multi-objective decision-making, different design schemes are generated and compared under heritage constraints.
- (4)
Promote to other types of heritage buildings: Apply the index level to other public heritage buildings, and compare the spatial priority changes in different types and user groups.