Next Article in Journal
Study on the Compressive Performance of Fabricated Reinforced Concrete Columns Strengthened with CFRP Sheets: Experimental and Finite Element Analysis
Previous Article in Journal
Physics-Informed Neural Networks for Urban and Building Thermal Environment Modeling: A Review of Evolution, Workflows, and Prospects
Previous Article in Special Issue
Reading Layered Industrial Heritage Through Graphic Documentation: Adaptive Reuse, Morphological Continuity, and Selective Legibility at Cibali
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Knowledge Representation Method for Grotto Buddhist Niches Based on Image Semantics and Ontology

1
School of Geomatics and Urban Spatial Informatics, Beijing University of Civil Engineering and Architecture, Beijing 102616, China
2
Key Laboratory for Architectural Heritage Fine Reconstruction & Health Monitoring, Beijing 100044, China
3
Institute for Digital and Intelligent Preservation and Inheritance of Cultural Relics, Chang’an University, Xi’an 710064, China
4
School of Electronics and Control Engineering, Chang’an University, Xi’an 710018, China
5
Digital Conservation Centre, Yungang Research Institute, Datong 037007, China
*
Author to whom correspondence should be addressed.
Buildings 2026, 16(13), 2563; https://doi.org/10.3390/buildings16132563
Submission received: 14 May 2026 / Revised: 12 June 2026 / Accepted: 23 June 2026 / Published: 26 June 2026

Abstract

Grotto Buddhist Niches are important spatial carriers of Buddhist cave art, containing rich architectural, artistic, and historical information. However, image data of these Buddhist niches are fragmented across multiple scales, including visual features, cultural semantics, and spatial structures, which significantly hinders cross-scale correlative analysis. To address this issue, this paper proposes a multi-scale knowledge representation method based on image semantics and ontology. Specifically, we establish a five-tier semantic description model, comprising the visual feature layer, image data layer, entity layer, cultural semantics layer, and relational layer. Furthermore, using Protégé and the classical Seven-Step Method, we develop a domain ontology named Grotto Buddhist Niche Ontology (GBNOnto) to enable unified semantic modeling of multi-scale information. Based on this ontology, a knowledge graph focusing on cave imagery is constructed, with typical caves such as Cave 38 at the Yungang Grottoes selected as case studies. The resulting graph contains 892 entity nodes and 2621 semantic relations, effectively capturing the complex interconnections among architectural typology, artistic characteristics, and cultural semantics within the selected niche instances. The proposed method enables structured and associative integration of multi-scale information in grotto Buddhist niche images. It thus provides a foundational data infrastructure and modeling framework to support effective management, knowledge retrieval, and semantic reasoning.

1. Introduction

The Buddhist niche is a vital spatial component of grotto architecture. It takes the form of a small subsidiary cave-chamber and contains extensive historical information, including architectural typology, spatial layout, structural elements, ornamental motifs, and decorative art [1]. With the development of Buddhism, the Buddhist niche gradually became the visual focus of grottoes and a core embodiment of cave art, serving as key physical evidence of Buddhist architectural art and spatial organization [2]. The global advancement of digital preservation technologies has accelerated the documentation and conservation of cave temple heritage, producing large amounts of text and image data [3,4,5]. These datasets have significant cultural and academic value and deserve in-depth exploration. Compared with plain text, images can depict niche characteristics more vividly [6] and arouse public interest [7]. Moreover, images have complex associative relations with textual information, providing researchers and the public with a direct visual basis for interpretation [8,9,10]. However, current research and applications related to Buddhist statues in Buddhist niches face significant challenges. Such images and texts are often stored in fragmented ways, without a unified organization or management framework, which prevents effective integration. In addition, the information contained in Buddhist niche images is dimensionally complex. It includes low-level visual attributes (e.g., color and texture) as well as high-level semantic information, such as Buddhist niche typology [11,12,13], decorative patterns [14,15], iconographic themes [16], carving techniques [17], artistic styles [18], and spatial layouts [19]. These visual and semantic cues ultimately refer to real-world physical spatial entities, with intricate logical interconnections among the information elements—complexity that conventional text or image-based archival methods struggle to systematize [20]. This “semantic gap” [21] between low-level visual features and high-level domain knowledge, combined with the disconnect between image content and spatial representation, limits the value of such images for digital preservation, scholarly research, and knowledge dissemination. Thus, there is a strong need for knowledge organization and representation of these images.
Semantic description provides a basis for systematically organizing and representing image-based knowledge [22]. Much research has focused on image semantic description, moving from early generic hierarchical models [23,24] to domain-specific frameworks for cultural heritage objects such as Dunhuang murals [25,26,27,28] and Thangka paintings [29]. These studies map visual features to high-level domain knowledge, and this shift from data-driven to semantic-driven approaches enables refined management and knowledge representation of complex visual entities. However, existing models have limitations when applied to grotto Buddhist niches. First, Buddhist niches have spatial attributes and hierarchical structures that are difficult to capture with current image-centric description systems. Second, most studies focus mainly on image-text associations and lack a unified modeling framework that integrates image semantics, textual descriptions, and spatial structures, which fragments multi-scale information at the representational level.
Knowledge graphs and ontology-based approaches provide a promising framework to address these challenges. Knowledge graphs represent domain entities and their relationships using a node–edge structure, offering strong support for organizing and reasoning over complex semantic relations. In this framework, nodes represent core entities, while edges encode multidimensional relationships, thereby enabling an intuitive and comprehensive representation of cultural heritage knowledge. An ontology, as a shared conceptual model for a specific domain, extends beyond image-level semantic descriptions. By establishing an objective conceptual system, it facilitates semantic interoperability across various descriptive standards and supports the integration of heterogeneous information. For cultural heritage image knowledge, many researchers have adopted ontologies as the knowledge framework for knowledge graphs to ensure semantic consistency and support knowledge organization. The CIDOC Conceptual Reference Model (CIDOC CRM), for example, provides a generic reference model for cultural heritage information and has been widely applied to the integration and management of heterogeneous cultural heritage resources [30,31,32,33,34]. Building upon CIDOC CRM, researchers have further developed domain-oriented ontologies for specific types of cultural heritage images. Examples include ontologies for ancient Chinese cartographic knowledge [35], cultural relic images [36], historical maps [37], art images [38], and Renaissance visual representations [39]. These studies have demonstrated the effectiveness of ontology-based approaches for the semantic organization and retrieval of cultural heritage image knowledge. However, existing cultural heritage ontologies primarily focus on the representation of heritage objects, events, actors, places, and documentary information. They offer limited capability for the fine-grained semantic representation of grotto Buddhist niche images, particularly in terms of niche morphology, sculptural composition, iconographic elements, spatial structures, and cultural-symbolic meanings. As a result, the complex visual, spatial, and semantic characteristics of grotto niche imagery cannot be adequately represented by existing generic ontology frameworks alone.
To address the gap in the fine-grained semantic representation of grotto Buddhist niche images, this study proposes a multi-scale knowledge representation method based on image semantic description and domain ontology modeling. Existing general-purpose cultural heritage ontologies, such as CIDOC CRM, provide an effective framework for representing heritage objects, events, actors, places, and documentary information. However, they provide limited support for the detailed representation of niche morphology, sculptural composition, iconographic elements, spatial structures, and cultural-semantic associations that are essential for grotto Buddhist niche imagery. Building upon a five-level semantic description framework, this study develops the Grotto Buddhist Niche Ontology (GBNOnto), a domain ontology specifically designed for Buddhist niche images. GBNOnto systematically organizes niche forms, sculptural components, iconographic themes, spatial relationships, and cultural-semantic associations into a unified conceptual structure. By integrating visual information, cultural knowledge, and spatial entities, the ontology supports semantic querying, knowledge reasoning, and knowledge graph construction for grotto heritage research. The originality of this study lies not only in the development of a domain ontology dedicated to grotto Buddhist niche images, but also in the integration of multi-scale image semantic description with ontology-driven knowledge representation. Rather than replacing CIDOC CRM, GBNOnto complements existing cultural heritage ontologies by providing a domain-specific semantic model tailored to the knowledge organization requirements of Yungang Grotto niche imagery. In this way, it bridges the gap between general cultural heritage ontologies and the fine-grained semantic representation needs of grotto Buddhist niche images.
The significance of this study is threefold. First, it provides a novel approach for the digital organization and management of grotto Buddhist niches and related cultural heritage resources, supporting precise retrieval and in-depth knowledge discovery beyond traditional archaeological typologies and art-historical descriptions. Second, by combining internationally recognized ontology concepts with the unique artistic styles and cultural contexts of grotto heritage, the proposed framework offers a reference for ontology construction and knowledge organization in other cultural heritage domains. Third, the integration of image semantic description, ontology modeling, and knowledge graph construction establishes a new technical pathway for intelligent image analysis, semantic reasoning, and digital heritage research.
The principal contributions of this paper are as follows:
  • A reusable knowledge representation framework. To bridge the semantic gap between low-level visual features and high-level domain semantics in grotto Buddhist niche images, this paper proposes a multi-scale descriptive model that integrates visual, spatial, and semantic dimensions. This framework provides a reusable way to organize knowledge for related research. It enables systematic representation from image features to cultural knowledge and from explicit data to implicit meanings.
  • A conceptual modeling paradigm. Using Protégé and the standard seven-step ontology engineering method, and following international standards such as CIDOC CRM, we construct a domain-specific ontology called the Grotto Buddhist Niche Ontology (GBNOnto). This ontology formalizes core concepts and logical relations for niche typology, iconographic themes, spatial configurations, and other dimensions. Our modeling approach can serve as a reference for digital knowledge modeling of other cultural heritage types, such as murals and ancient architecture, thus facilitating cross-domain information integration and sharing.
  • An application-oriented knowledge graph for grotto Buddhist niche images. Based on the above models, we build a knowledge graph of Buddhist niche images from the Yungang Grottoes. This graph integrates dispersed image resources, documentary records, spatial structures, and semantic knowledge into a structured network. The resulting knowledge graph improves the standardization and interoperability of cultural heritage information, offering new ways for intelligent retrieval, deep semantic discovery, and visual knowledge exploration.
The remainder of this paper is organized as follows. The methodology for knowledge representation is presented in Section 2. The feasibility of the proposed method is validated through a case study in Section 3. The main conclusions, limitations, and directions for future research are discussed in Section 4.

2. Methodology

2.1. Hierarchical Model for Image Semantic Description

Image semantic description refers to the formal representation of an image’s core concepts, attributes, and associative relations. To ensure comprehensive and deep descriptions, fine-grained deconstruction at multiple levels is required, as different levels of granularity are essential for meeting the various needs of feature analysis. A hierarchical semantic description model plays a key role in ontology design and the discovery of associative relations. Based on the Eakins image semantic hierarchy model [23] and the hierarchy proposed by Wang Xiaoguang et al. [25], and considering the unique characteristics of grotto Buddhist niche images, this study establishes a semantic description hierarchy following a “bottom-up, explicit-to-implicit” logic. This hierarchy is formally represented as a five-tuple structure M = { V , I , E , C , R } , corresponding to five levels: visual features, image data, entity objects, cultural semantics, and associative relationships. The model describes the knowledge representation and spatial cognition processes involved in interpreting Buddhist niche images, as shown in Figure 1.
(1)
Visual Feature Layer. As the foundational layer of the GBNOnto multi-scale semantic framework, the Visual Feature Layer represents the low-level visual characteristics extracted directly from image data. It mainly includes appearance features (e.g., color and texture), geometric features (e.g., shape, contour, and edge), and local structural features (e.g., key points and feature regions). These visual descriptors provide the fundamental data basis for subsequent entity recognition, semantic interpretation, and knowledge representation. By establishing the connection between raw image perception and higher-level semantic understanding, this layer serves as the starting point of the semantic knowledge construction process. (Note that this study focuses on high-level semantic information; the extraction of low-level visual features is not explored in depth here).
(2)
Image Data Layer. This layer uses metadata standards to standardize image descriptions, thus building a unified digital object representation. It includes image titles, data formats, resolution, acquisition time, cave chamber identifiers, and other relevant information, turning raw pixel data into structured digital resources. This layer provides essential support for image management, retrieval, and association.
(3)
Entity Object Layer. This layer maps visual information in an image to entity objects with clear semantic boundaries, such as the Buddhist niche body, niche lintel, principal deity statue, decorative motifs, and structural components. This transformation from low-level visual features to high-level semantic objects is a key step in building entity nodes in the knowledge graph. In addition, the entity representations defined here can provide semantic constraints for recognizing and annotating grotto Buddhist niche images.
(4)
Cultural Semantics Layer. This layer uses domain knowledge to classify entity objects typologically, giving them historical and artistic meanings. The semantic content includes iconographic themes, artistic styles, carving techniques, and religious backgrounds. This layer moves from “entity recognition” to “semantic understanding,” thus giving image data interpretable cultural significance.
(5)
Associative Relations Layer. Using multiple relation types, this layer links the entities and semantics from the previous layers to form a comprehensive knowledge network. The relationship types include spatial, temporal, and semantic relations, enabling unified modeling and associative representation across multi-scale information.

2.2. Ontology Design

Because ontology engineering must address different application needs, a variety of ontology development methodologies have emerged. Among the commonly used approaches are the seven-step method, the KACTUS methodology, the Toronto Virtual Enterprise (TOVE) framework, and the skeletal approach. The latter three are mainly tailored to particular domains and therefore lack broad general applicability. By comparison, the seven-step method proposed by Stanford University is domain-independent and can be applied across multiple fields. Figure 2 illustrates the general workflow of the Seven-Step Method and the relationship between it and the construction of the GBNOnto.
The specific steps are as follows:
Step 1: Determine the domain and scope
Determining the domain and scope of an ontology is crucial for identifying the research subject and clarifying the purpose and knowledge content of the ontology design. When building the conceptual model of GBNOnto, we must define its specific domain and scope to ensure that the resulting model strictly follows the discipline’s meaning and structure. As shown in Figure 1, the domain of GBNOnto includes such as image data and entity object information, and its goal is to provide a structured knowledge representation framework for knowledge graph construction.
Step 2: Reuse existing ontologies
The rationale for reusing existing ontologies lies in enhancing the reusability and maintainability of the ontology, while simultaneously improving semantic interoperability across models and reducing modeling costs [40]. In the construction process of this study, priority was given to reusing and referencing well-established, mature ontologies and standards. CIDOC CRM (see http://www.cidoc-crm.org/cidoc-crm/ accessed on 12 November 2025, prefix: crm) is a widely adopted general ontology in the cultural heritage domain.
Within this framework, E22 Human-Made Object is defined as “all permanent physical objects of any size that are deliberately created by human activity.” In this study, we adopt E22 Human-Made Object as the entity type for grotto Buddhist niches and their associated physical objects. We also emphasize the central role of images as visual symbols in the semantic representation of cultural heritage. In CIDOC CRM, E38 Image represents the digital image itself, while E90 Symbolic Object describes the visual content in the image that carries symbolic meaning. Accordingly, GBNOnto instantiates digital image instances as E38 Image and links them, via object properties such as P62 depicts, to the specific visual elements identified in the image, which are typed as E90 Symbolic Object. Then, through properties such as P138 represents, these symbolic objects connect to the real-world physical entities or abstract concepts they signify. This multi-layered mapping ensures that every aspect—from the image’s physical existence to its visual content and the real-world entities it denotes—receives precise semantic representation in the ontology.
GeoSPARQL extends the SPARQL query language (see http://www.opengis.net/ont/geosparql accessed on 18 October 2025, prefix: geo) and provides support for storing, querying, and analyzing geospatial data in the Semantic Web. In GeoSPARQL, the class geo: SpatialObject is defined as any entity with a spatial representation. Within the GBNOnto model, geo: SpatialObject is interpreted as “any spatial entity having shape, location, or extent,” serving as a generic class for describing objects in geographic space. As shown in Figure 3, this study maps visual semantic units in grotto Buddhist niche images to spatial objects by treating E22 Human-Made Object as equivalent to geo: SpatialObject.
In addition, this study uses relevant classes and properties from the Friend of a Friend (FOAF) ontology (see http://xmlns.com/foaf/0.1/ accessed on 5 December 2025), specifically foaf: Agent, foaf: Person, and foaf: Group, to describe social entities in the images, such as individuals and groups. For digital image attributes like title, description, and source, this study uses the Dublin Core Metadata Initiative (DCMI) Metadata Terms standard (see https://www.dublincore.org/specifications/dublin-core/dcmiterms/ accessed on 26 September 2025). We map terms from its core set, including dc: title, dc: description, dc: source, dc: creator, dc: date, dc: type, and dc: identifier, to ensure standardized and interoperable image metadata. The relationship between them is shown in Figure 3.
Step 3: Enumerate important terms
The purpose of enumerating terms is to collect the core vocabulary used in the domain and to ensure that the ontology is comprehensive within its scope. Based on the hierarchical model of image semantic description, this paper systematically identifies and lists the specialized terms for grotto Buddhist niches. This terminology draws on professional research from the fields of grotto Buddhist niches and Buddhist statuary, ensuring that the terms are professional, valid, and applicable. The details are presented in Figure 1.
Step 4: Define the classes and the class hierarchy
Defining classes and class hierarchies aims to systematically categorize and organize domain-specific concepts. Classes, which can also be seen as domain entities, represent a set of objects that share common attributes and meanings. They exist independently of other entities. In GBNOnto, we abstract Buddhist niche image data, visible objects, and implicit semantic elements to build a core class system that includes persons, time periods, events, locations, and physical entities. These classes form a hierarchy through inheritance and create semantic associations via object properties.
As shown in Figure 4a, digital images of grotto Buddhist niches contain many identifiable visual objects—such as the whole Cave, Buddhist Niche, Wall Surface—which also exist in the real world and are therefore defined as geo: SpatialObject. Figure 4b shows that a Buddhist niche can be decomposed into structural components, including the Niche Lintel, Niche Pillars, and Niche Beam. In this study, we abstract these components as subclasses of E90 Symbolic Object, such as Buddhist Niche, Structural Component and Primary Buddha Statue. Further subdivisions, such as persons and events, are given in Table 1.
Step 5: Define the properties of classes and the relation between classes
Data properties describe the specific information of a class and form the basis of structured information representation. Each property has a name, a domain, and a range. These identify the property, specify the class it applies to, and define the allowed value types. Given the characteristics of grotto Buddhist niche images, this study defines a unified set of data properties for core entities such as Buddhist niches and principal deity statues. Although different entities may vary in spatial form and artistic style, they share common descriptive features. Thus, we define a set of basic attributes, including name, identifier, height, width, depth, type, and preservation status. Table 2 gives partial definitions of data properties for Buddhist niches and the primary Buddha statue.
Object properties are used to define the various relations between classes. In GBNOnto, the classes exhibit rich associative characteristics, which can be summarized from three perspectives: spatial relations, temporal relations, and semantic relations, as shown in Figure 5.
(1)
Spatial Relations
Spatial relations are employed to describe the topological, directional, and other aspects of relations among spatial entities [41]. In this study, they are categorized into image spatial relations and scene spatial relations. Image spatial relations refer to the spatial relations among entities within a single image—such as components, statues, and decorative motifs—serving to express the internal spatial organization characteristics of the image. Scene spatial relations, by contrast, transcend the scale of a single image and are used to describe the layout relations of Buddhist niches across cave walls and within the spatial units of a cave chamber. In terms of concrete representation, these relations are primarily described through topological relations (Figure 6), relative positional relations (Figure 7), and wall-niche relations [42] (Figure 8).
(2)
Temporal Relations
Temporal relations are used to describe the occurrence times of entities related to grotto Buddhist niches, as well as the relative temporal relations among them. In this study, temporal relations are classified into two categories: quantitative time and qualitative time. Quantitative time focuses on numerical, precise expression, encompassing the acquisition and generation times of image data, as well as the construction periods with exact chronological records. Qualitative time expresses the temporal relations among entity objects within grotto Buddhist niche images, such as the sequential order of niche excavation. Specific details are presented in Table 3.
(3)
Semantic Relations
Semantic relations are employed to describe the conceptual and cultural-semantic connections among various visual objects and their attributes within Buddhist niche images. In this study, semantic relations are employed not only to describe the structural organization and attributive characteristics of different objects within the images of Buddhist niches, but also to reveal their interrelated significance in terms of artistic form, historical context and cultural connotations. Following the knowledge representation principles of the Web Ontology Language (OWL) and considering the semantic characteristics of Yungang Grotto niche images, this study organizes five categories of semantic relations: hierarchical relations, attribute relations, whole-part relations, instance relations, and logical relations. Among these, hierarchical relations and instance relations are derived from the ontology modeling mechanisms of OWL, while the remaining relation types are established according to the structural features and cultural semantics of grotto imagery. Hierarchical relations describe superclass–subclass relationships among concepts and are used to organize domain knowledge into a structured taxonomy. In OWL, such relations are represented through the subClassOf property. For example, the concept Buddhist Niche Form can be further classified into Round-Arched Niche, Caisson-Shaped Niche, Tent-Shaped Niche, and Canopied Niche, thereby supporting the inheritance and organization of domain knowledge. Attribute relations are used to describe the characteristic information inherent to an object itself. For example, a Primary Buddha Statue can be characterized by attributes such as “iconographic theme,” “posture,” and “dimensions.” Whole-part relations serve to express the structural composition of a niche. For instance, a Buddhist Niche is composed of components such as the niche lintel and niche pillars. Instance relations connect abstract concepts to specific objects, mapping a given class of concepts to concrete instances of niches or statues. For example, “the round-arched niche on the north wall of Cave 38 of the Yungang Grottoes” serves as a concrete instance of the concept Buddhist Niche. Logical relations reflect causal, functional, or symbolic connections among objects. These include, for instance, the relationship between a principal deity and its donors, the correspondence between decorative motifs and Buddhist symbolic meanings, as well as the mapping links between objects and external knowledge such as documentary records, historical events, and chronological periods.
Step 6: Define the facets of the properties
Defining property facets involves specifying the measurement characteristics associated with class attributes. For instance, the height and depth of a niche are typically represented as numerical measurements in centimeters, while descriptive features of a niche or a Primary Buddha Statue are generally encoded as textual strings. Within the CIDOC CRM framework, factual values are modeled using E59 Primitive Value, which can further include subclasses such as E60 Number and E62 String.
Step 7: Create instances
The last step of ontology design is to create class instances. To define an instance, we first determine its class and then assign attribute values that match the properties of that class. For example, Figure 9 shows the Buddhist niche on the west wall of Cave 38 of the Yungang Grottoes. The Buddhist niche has a clear structure, including a principal deity statue, donor figures, an inscription, and other decorative elements. Inside the Buddhist niche, the principal deity is a seated Buddha statue, flanked by several donor figures kneeling with palms joined. Above the Buddhist niche is a nine-square folding pattern, and niche pillars are present on both sides of the wall. The cave also contains the Wu Tian’en Statue Inscription, which clearly records the donor Wu Tian’en and his information, as well as the date and motivation for the creation of the statue.

2.3. Knowledge Acquisition

Knowledge acquisition is a critical step in building a knowledge graph. Its main task is to extract structured knowledge from multi-source heterogeneous data and then standardize and unify its representation and organization. Given the characteristics of grotto Buddhist niche images, the knowledge acquisition process includes three parts: knowledge extraction, entity alignment, and knowledge storage.

2.3.1. Knowledge Extraction

Knowledge extraction identifies and extracts knowledge units (such as entities, relations, and attributes) from semi-structured or unstructured text and is a key step in building knowledge graphs [43]. In natural language processing, knowledge extraction is also called named entity recognition (NER), which finds named entities with clear meanings from text, such as persons, time and locations. In this study, we employed a fine-tuned Universal Information Extraction (UIE) pre-trained model to automatically generate candidate entities and relations from grotto-related textual sources, including archaeological reports, descriptive annotations of Yungang Grotto Buddhist niches, and iconographic records. The UIE model was applied to identify key entity types within the domain of Buddhist niches, such as niche typologies, statues, and decorative patterns. For example, “round-arched niche” is recognized as an entity of the niche typology class, and “Shakyamuni Buddha” is identified as an instance of the Primary Buddha Statue class.
Through named entity recognition, we obtained a set of candidate entities. However, entities alone are not sufficient to build a structured knowledge system. We also needed to identify semantic relations among entities and organize them into knowledge triples. The UIE model generated candidate triples <Entity1, Relation, Entity2>, such as <Cave 38, includes, round-arched niche>, <Cave 38, includes, caisson-shaped niche>, and <round-arched niche, includes, niche lintel>. Specific examples are shown in Table 4. Knowledge extraction also includes extracting entity attributes that describe specific characteristics of Buddhist niches and related entities. Attribute extraction mainly finds attribute names and their values, for example: <round-arched niche, height, 1440 cm> and <round-arched niche, width, 820 cm>.
After the automatic extraction stage, all candidate entities, relations, and attributes underwent a thorough manual verification and correction process. This work was carried out under the guidance of domain experts from the Yungang Grottoes Research Institute, whose deep knowledge of the site and its iconography was essential for ensuring the accuracy and precision of the final knowledge graph. Every extracted item was reviewed, with spurious or irrelevant results removed, duplicate entities merged, mislabeled relations corrected, and missing elements supplemented. The verified entities and relations were then normalized to the GBNOnto schema layer: entities were mapped to the appropriate ontology classes, and relations to the corresponding object properties.

2.3.2. Entity Alignment

Entity alignment aligns and integrates knowledge extraction results, mainly using entity disambiguation and attribute fusion. Entity disambiguation checks whether different names from different sources refer to the same real-world object, resolving ambiguity when multiple expressions denote the same entity. For example, “Shakyamuni and Prabhūtaratna Two Buddhas”, “Two Buddha Statues”, and “Two Buddhas Sitting Side by Side” refer to the same entity, so we unify them under the standard name “Two Buddhas Sitting Side by Side.” However, sometimes the same name refers to different objects. For instance, a cave may have multiple “round-arched niches”. To avoid confusion, we assign each entity a unique identifier, e.g., we name the round-arched niche on the north wall of Cave 38 as “YG-38-BB-YGK-01” and the one on the east wall as “YG-38-DB-YGK-01”. Attribute fusion handles issues like inconsistent attribute names and measurement units. For example, “niche height” and “height” mean the same attribute and should be normalized. Also, units should be standardized, e.g., convert “460 cm” and “4.6 m” to the same unit.

2.3.3. Knowledge Storage

Knowledge storage converts scattered, real-world information into structured data suitable for computer processing. To this end, we use formal description languages including XML, RDF, RDFS, and OWL. Among them, OWL excels at complex semantic modeling. It supports fine-grained description of resources and their relations using a predefined vocabulary. For storage, graph databases have become the standard choice for knowledge graphs. They organize data as nodes and edges, which naturally reflects how the data are connected. Compared with traditional relational databases, graph databases are more efficient at multi-hop queries and complex relation analysis. Given these advantages, we select Neo4j [44]—a widely used graph database—to store and manage our knowledge triples.
The specific implementation procedure is as follows. First, we organize the extracted entities and attributes into a CSV file. Second, we use Neo4j’s LOAD CSV command to import the data. Each row in the CSV file becomes a node, and each column becomes an attribute of that node. Third, based on the entity and relation information in the knowledge triples, we apply the MERGE statement to create or match target nodes and build relations between them. Finally, after importing and mapping all triples, the knowledge data is stored persistently, which provides basic support for subsequent queries and analysis.

3. Results

3.1. Data Collection

As shown in Figure 10, the Yungang Grottoes are located at the southern foot of Wuzhou Mountain in Datong City, Shanxi Province, China. First excavated during the Northern Wei dynasty, they have a history of over 1500 years and stand as an important physical testament to early Buddhist statuary art in China. In December 2001, the Yungang Grottoes were inscribed on the UNESCO World Heritage List as a site of outstanding universal value.
The Yungang Grottoes are renowned for their grand scale and exquisite carvings, housing approximately 1100 Buddhist niches of various types, schematic diagrams of some of which are shown in Figure 11. They represent a significant example of Chinese cave art. A substantial body of archaeological reports, monographs, and academic papers has been accumulated on the grottoes, providing rich and reliable data for related research. Given their outstanding academic value and solid documentation, this study selects the Yungang Grottoes as the research object and chooses representative Buddhist niche instances for analysis.
Data acquisition in this study proceeds in two steps. First, we gathered Buddhist niche images from on-site photography, image collections, and publicly accessible digital archives. Second, we obtained corresponding textual descriptions from authoritative sources, including historical documents, archaeological reports, academic monographs, and journal articles—such as The Dictionary of the Yungang Grottoes (ISBN 9787534449963), The Construction of the Yungang Grottoes (ISBN 9787501050840), and Studies on Late Northern Dynasties Cave Temples (ISBN 9787501014095). After systematic collation and selection, we compiled a raw dataset of grotto Buddhist niche images and texts.

3.2. Ontology Model

This study employed the Protégé 5.5.0 platform, the domain ontology model for grotto Buddhist niches was constructed and uniformly expressed using the Web Ontology Language (OWL) [45] This ontology comprises 9 first-level classes, 32 s-level classes, 47 object properties, and 25 data properties, thereby establishing a relatively systematic framework of domain concepts and semantic relationships. As illustrated in Figure 12 and Figure 13 the ontology model exhibits sound structuring and normative rigor in terms of class hierarchy organization, property constraints, and entity associations. It is capable of describing, with considerable completeness, the semantic relations among niche entities and their related elements, thereby providing foundational support for subsequent knowledge graph construction and semantic applications.
To ensure the logical consistency and quality of GBNOnto, a description logic (DL) reasoning check was conducted in Protégé 5.5.0 using the HermiT reasoner. The reasoner was used to verify class satisfiability, property consistency, domain and range definitions, and the logical coherence of the class hierarchy. Based on the reasoning results, the ontology structure and property definitions were checked and refined. The final reasoning results showed that no inconsistent or unsatisfiable classes were detected, and the class hierarchy, object properties, and data properties were logically coherent under the current ontology constraints. This validation step provides ontology-engineering quality assurance for GBNOnto before its application in knowledge graph construction.

3.3. Knowledge Graph Construction

The knowledge graph of Yungang Grotto Buddhist niche images, constructed based on the GBNOnto ontology, comprises a total of 20 classes. These cover domain elements such as locations, grotto temples, cave chambers, Buddhist niche, typologies, structural components, principal deity statues, subsidiary statues, inscriptions, pedestals, and decorative motifs, forming a comprehensive conceptual architecture. As shown in Figure 14, each class contains a different number of entity instances. For example, the “Typology” class has six entities, the “Decoration” class has 64, and the “Component” class has 6. The core “Buddhist niche” class includes 11 entities, such as the “caisson-shaped niche on the west wall of Cave 38”. Each Buddhist niche entity is described by five core attributes: name, identifier, location, typology, and spatial hierarchy. For the “caisson-shaped niche on the west wall of Yungang Cave 38”, the name is “Buddhist niche on the west wall of Cave 38”, the identifier is “YG-38-WW-LXK-01”, the typology is “caisson-shaped niche”, the location is “west wall” and the affiliated cave chamber is “Cave 38”. The “Primary Buddha Statue” class contains 14 entities. Each statue entity is described by attributes: height, shoulder width, body posture, hairstyle, facial expression, hand gesture, attire, ornaments, and preservation condition. These attributes capture the statue’s morphological features and state of preservation comprehensively.
The visualization of the ultimately constructed knowledge graph of Yungang Grotto niche images in Neo4j is presented in Figure 14. This knowledge graph comprises a total of 892 entity nodes and 2621 semantic relations. Furthermore, manual completion was performed for spatial topological relations that were either implicit or missing in the textual descriptions. For example, although the text does not explicitly describe the adjacency relation between the “niche lintel” and the “niche pillar,” based on common knowledge of cave architecture and with reference to the corresponding images, the niche lintel is adjacent to the niche pillar. Such information, which is visually evident but omitted in the text, cannot be automatically extracted by the aforementioned model alone. Therefore, this study supplemented the spatial topological relations among components through manual verification, thereby further enhancing the completeness and consistency of the knowledge graph.
Approximately 13% of the relations were manually constructed. To ensure transparency, all manually added relations were explicitly labeled and distinguished from automatically extracted ones within the knowledge graph. Each manual annotation was systematically documented, including the corresponding triple, justification, evidence source. To improve consistency and reproducibility, manually added relations were checked against the original images, archaeological drawings, and relevant textual descriptions. Relations lacking sufficient visual or textual support were excluded. This procedure helps reduce subjectivity and ensures that the manually supplemented spatial relations remain traceable and verifiable.

3.4. Application of Knowledge Graph

Knowledge retrieval
To assess how well the constructed knowledge graph supports multi-scale knowledge representation and organization for grotto Buddhist niche images, we used Cypher queries to retrieve knowledge from the graph.
Taking Cave 38 of the Yungang Grottoes as a case study, we retrieved semantic information from the images with the following query:
MATCH (n: Cave {cave_number:‘Cave 38’})-[r*0..5]-> (result)
RETURN n, r, result
This query starts from entity objects and traverses their subordinate niche and statue constituent elements, thereby achieving a global retrieval of image-level entities. The retrieval results (as shown in Region 1 of Figure 15) reveal the hierarchical structure of entities within the cave chamber, including Buddhist niches, Primary Buddha Statues, subsidiary figures, and decorative elements. Regions 2 and 3 of Figure 15 present detailed attribute information. Such structured information can provide researchers with comprehensive visual and cultural semantic references, thereby enabling a systematic understanding of the semantics of Buddhist niche images.
Relation retrieval
Based on our image knowledge representation, this study further validates the effectiveness of the knowledge graph in representing spatial relations among entities. Taking Cave 38 of the Yungang Grottoes as a case study, Cypher query statements are employed to retrieve the spatial relations among Buddhist niche structural components, the Primary Buddha Statue, and subsidiary figures. The query results are presented in Figure 16.
MATCH (n: Niche {Name: ‘Yungang Cave 38 North Wall Arched Niche’})-[r:has_spatial_relation*0..5]->(result)
RETURN n, r, result

4. Discussions and Conclusions

4.1. Discussions

This study explores the knowledge representation of grotto Buddhist niche images at both methodological and applicational levels, its significance is primarily reflected in the following aspects. At the methodological level, the proposed five-tier semantic description model effectively improves the granularity and semantic depth of information representation for grotto Buddhist niche images. Unlike approaches that focus solely on image content annotation or single-object recognition, this model not only captures the external morphology, compositional features, and structural components of Buddhist niches, but also links them to historical contexts, iconographic themes, spatial locations, and cultural attributes. In doing so, it transforms “visible information” into “interpretable knowledge”. Nevertheless, the semantic description process still depends to some extent on manual annotation and domain knowledge accumulation, and its level of automation needs further improvement.
Second, regarding ontology construction, the domain ontology GBNOnto—developed using the seven-step method—systematically abstracts and formally models the concepts, attributes, and relations associated with grotto Buddhist niches. This effectively reduces inconsistencies in terminology, hierarchies, and semantic expressions across different data sources. In particular, when handling mapping relations between Buddhist niches and cave, between components and wholes, and between image objects and cultural entities, the ontology provides essential semantic constraints that enable multi-source knowledge to be organized and extended within a unified framework. Nevertheless, in complex scenarios, knowledge extraction, entity alignment, and relation recognition may still be affected by terminological ambiguity, semantic incompleteness, and variations in image quality. Finally, constructing a knowledge graph for Cave 38 of the Yungang Grottoes further validates the feasibility of our method. The entity nodes and semantic relations in the graph show that grotto Buddhist niche images are not isolated visual objects but composite knowledge units tightly linked to architectural space, iconographic content, artistic styles, and historical-cultural contexts. However, this study is based mainly on a single cave chamber. The generalizability of the domain ontology and knowledge graph still needs to be tested across more cave chambers and different grotto temple complexes.

4.2. Conclusions

This study addresses the structured knowledge requirements of grotto Buddhist niches by proposing a multi-scale knowledge representation method that integrates image semantic description and a domain ontology. Through the construction of the domain ontology model GBNOnto, the conceptual system and logical relations pertaining to niche typology, iconographic themes, decorative patterns, and chronological periods are formalized. Based on this ontology, a knowledge graph is further constructed, enabling the systematic representation and cross-level association of niche images from visual features to cultural semantics. The validation results using Cave 38 of the Yungang Grottoes as a case study demonstrate that the proposed method supports semantic querying and knowledge organization oriented toward cultural connotations, thereby providing an effective pathway for the in-depth interpretation and knowledge-based servicing of grotto Buddhist niche images.
Future research can be extended in two directions. First, the conceptual system and hierarchical structure of GBNOnto will be further refined by incorporating more grotto heritage data from different regions and historical periods, including representative Chinese grotto sites such as the Longmen Grottoes and Mogao Grottoes, thereby improving its generalizability, scalability, and cross-regional applicability. In particular, a standardized glossary or dictionary for Chinese grotto heritage should be established to support the unification of entity names, attribute values, and domain terminology, thereby reducing ambiguity in semantic modeling. Future work will focus on improving the automation of knowledge extraction and knowledge graph construction. Although manual entity disambiguation, attribute disambiguation, and verification help ensure the accuracy and reliability of domain knowledge graphs, this process remains labor-intensive and time-consuming. Therefore, automated knowledge extraction methods, semantic rule definition, and intelligent annotation techniques should be introduced to reduce manual intervention while maintaining knowledge quality. These efforts will promote grotto heritage knowledge management toward higher efficiency, greater precision, and stronger automation.

Author Contributions

L.W. and M.H. conceived the presented idea. L.W. designed the experiments and was the primary writer of the manuscript. M.H., H.S. and J.L. reviewed and edited the manuscript. B.Z. and B.Y. assisted in data collection. B.N. provided research resources. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by National Natural Science Foundation of China (Grant No. 42171444, Project Name: Knowledge modeling of large-and-complex cultural relics for virtual restoration).

Data Availability Statement

The complete dataset generated and analyzed during the current study is not publicly available due to restrictions related to unpublished grotto heritage records and expert annotations. Further information or representative examples may be available from the author upon reasonable request and with permission from the relevant data holders.

Conflicts of Interest

The authors declare that there are no conflicts of interest.

References

  1. Yang, H.; Wang, Q. The Architectural Form and Structural Characteristics of Dunhuang Mogao Grottoes. J. Northwest Univ. (Nat. Sci. Ed.) 2022, 52, 199–212. [Google Scholar] [CrossRef]
  2. Pei, Q.; Wang, G.; Ran, W. An Analysis of the Evolution of Architectural Forms and the Variation of Decorative Arts for Grotto Niche Space. J. Dunhuang Stud. 2022, 114–136. [Google Scholar]
  3. Dou, J.; Qin, J.; Jin, Z.; Li, Z. Knowledge Graph Based on Domain Ontology and Natural Language Processing Technology for Chinese Intangible Cultural Heritage. J. Vis. Lang. Comput. 2018, 48, 19–28. [Google Scholar] [CrossRef]
  4. Siliutina, I.; Tytar, O.; Barbash, M.; Petrenko, N.; Yepyk, L. Cultural Preservation and Digital Heritage: Challenges and Opportunities. Rev. Amazon. Investig. 2024, 14, 262–273. [Google Scholar] [CrossRef]
  5. Buragohain, D.; Meng, Y.; Deng, C.; Li, Q.; Chaudhary, S. Digitalizing Cultural Heritage through Metaverse Applications: Challenges, Opportunities, and Strategies. Herit. Sci. 2024, 12, 295. [Google Scholar] [CrossRef]
  6. Wang, Y. Combinatorial Analysis of Images in the Third Period of the Yungang Grottoes. Dunhuang Res. 2020, 48–62. [Google Scholar] [CrossRef]
  7. Yasser, A.M.; Clawson, K.; Bowerman, C. Saving Cultural Heritage with Digital Make-Believe: Machine Learning and Digital Techniques to the Rescue. In Proceedings of the HCI 2017: Digital Make Believe—Proceedings of the 31st International BCS Human Computer Interaction Conference, HCI 2017, London, UK, 11–13 July 2017. [Google Scholar] [CrossRef]
  8. Li, N. A Study of Niches and Capitals in the Caves of the Northern Dynasties and Their Origins. Cave Temples Stud. 2017, 347–382. [Google Scholar]
  9. Hideo, O.; Xu, X.; Meng, H.; Wang, Y. Early Buddhist Statues at the Yungang Grottoes—Focusing on the Niches in the Five Tan Yao Grottoes. Acad. J. Jinyang 2023, 29–50. [Google Scholar] [CrossRef]
  10. Chen, T. Research on Semantic Model of Image Knowledge Reuse from the Perspective of Digital Humanities. Libr. J. 2023, 42, 106–115. [Google Scholar] [CrossRef]
  11. Song, L. An Overview of the Architectural Styles of the Buddha Niches at the Mogao Caves. Silk Road 2010, 11–13. [Google Scholar]
  12. Zhang, H. Characteristics and Layout of Cave Forms in the Yungang Grottoes. Southeast Cult. 2003, 40–43. [Google Scholar]
  13. Wu, Y. A Brief Analysis of the Architectural Features of the Late-Period Caves at the Yungang Grottoes. In The Proceedings of the First Shanxi Museum Youth Forum; Shanxi Museum Association: Taiyuan, China, 2023; pp. 95–106. [Google Scholar]
  14. Xin, Y. A Study of the Pictorial Language of Honeysuckle Ornamental Patterns in the Yungang Grottoes. New Horiz. 2024, 46–48. [Google Scholar]
  15. Zhao, X. A Brief Discussion of the Decorative Motifs of the Yungang Grottoes. Identif. Apprec. Cult. Relics 2022, 158–161. [Google Scholar] [CrossRef]
  16. Cai, Y. Inter textuality and Inter visuality: Tracing the origins of the Astagada-mahapurusa iconography in Nor thern Zhou MaijishanGrotto Cave 4. J. Nanjing Univ. Arts (Fine Arts Des.) 2025, 159–166+210. [Google Scholar]
  17. Liu, L. The Carving Techniques and Artistic Characteristics of Offering Bodhisattva Statues in Cave 8 of Yungang Grottoes. J. Shanxi Datong Univ. (Soc. Sci. Ed.) 2024, 38, 62–66. [Google Scholar]
  18. Zhang, X. The Influence of Gandhāra Art on the Stylistic Characteristics of the Yungang Grottoes. Identif. Apprec. Cult. Relics 2025, 118–121. [Google Scholar] [CrossRef]
  19. Huo, J. The layout of the sculptures and the architectural features of the Yungang Grottoes. Collections 2025, 28–30. [Google Scholar]
  20. Zhang, J.; Ren, T. A Conceptual Model for Ancient Chinese Ceramics Based on Metadata and Ontology: A Case Study of Collections in the Nankai University Museum. J. Cult. Herit. 2024, 66, 20–36. [Google Scholar] [CrossRef]
  21. Kwan, P.; Kameyama, K.; Gao, J.; Toraichi, K. Content-Based Image Retrieval of Cultural Heritage Symbols by Interaction of Visual Perspectives. Int. J. Pattern Recognit. Artif. Intell. 2011, 25, 643–673. [Google Scholar] [CrossRef]
  22. Giunchiglia, F.; Bagchi, M.; Das, S. From Knowledge Organization to Knowledge Representation and Back. Ann. Libr. Inf. Stud. 2024, 71, 137–149. [Google Scholar] [CrossRef]
  23. Eakins, J. Automatic Image Content Retrieval—Are We Getting Anywhere? In Proceedings of the Third International Conference on Electronic Library and Visual Information Research, Milton Keynes, UK, 30 April–20 May 1996. [Google Scholar]
  24. Sengupta, A. Panofsky’s Iconology. In Studies in Iconology: Humanistic Themes in the Art of the Renaissance; Routledge: Oxfordshire, UK, 1939. [Google Scholar]
  25. Wang, X.; Xu, L.; Li, G. Semantic Description Framework Research on Dunhuang Fresco Digital Images. J. Libr. Sci. China 2014, 40, 50–59. [Google Scholar] [CrossRef]
  26. Wang, X.; Song, N.; Zhang, L.; Jiang, Y. Understanding Subjects Contained in Dunhuang Mural Images for Deep Semantic Annotation. J. Doc. 2017, 74, 333–353. [Google Scholar] [CrossRef]
  27. Zhang, Z.; Cao, Y. Dunhuang_Faces: A Dataset Enhanced through Image Processing for Cultural Heritage and Machine Learning. Trait. Signal 2025, 42, 569–581. [Google Scholar] [CrossRef]
  28. Wang, X.; Gong, Y.; Myers, D.; Wang, S. Arches Dunhuang: Heritage Inventory System for Conservation of Grotto Resources on the Gansu Section of the Silk Road in China. Int. Arch. Photogramm. Remote Sens. Spat. Inf. Sci. 2021, 46, 837–843. [Google Scholar] [CrossRef]
  29. Zhao, S.; Hou, X. A Study on a Descriptive Framework for the Semantic Information of Thangka Images. Knowl. Manag. Forum 2015, 2015, 57–62. [Google Scholar] [CrossRef]
  30. Tan, G.; Hou, X.; Zhuang, W. Research on Semantic Organization of Multimedia Resources of Intangible Cultural Heritage. Res. Libr. Sci. 2017, 42–52. [Google Scholar] [CrossRef]
  31. Fan, T.; Wang, H. Research of Chinese Intangible Cultural Heritage Knowledge Graph Construction and Attribute Value Extraction with Graph Attention Network. Inf. Process. Manag. 2022, 59, 102753. [Google Scholar] [CrossRef]
  32. Fan, T.; Wang, H.; Hodel, T. CICHMKG: A Large-Scale and Comprehensive Chinese Intangible Cultural Heritage Multimodal Knowledge Graph. Herit. Sci. 2023, 11, 115. [Google Scholar] [CrossRef]
  33. Li, J.; Qi, J.; Han, S.C.; Holden, E.-J. MUSEKG: A Knowledge Graph Over Museum Collections. arXiv 2025, arXiv:2511.16014. [Google Scholar]
  34. Dai, T. A Study on Metadata Standards for Digital Images of Museum Artefacts Based on CIDOC CRM: A Case Study of the Design of the Metadata System for Artefact Images at the National Museum of China. Chin. Mus. 2020, 131–136. [Google Scholar]
  35. He, B.; Zhao, Y.; He, J.; Ma, Z. Ontology Modelling of Ancient Map Information Through a Cognition-Practical Model: A Case Study of the Yangshi Lei Archives. Geomat. Inf. Sci. Wuhan Univ. 2024, 49, 546–561. [Google Scholar] [CrossRef]
  36. Gao, J.; Fu, J. A Fine-Grained Knowledge Representation Method of Cultural Heritage Image Resources Based on Knowledge Element. Inf. Sci. 2022, 40, 16–24. [Google Scholar] [CrossRef]
  37. Qi, X.; Yang, H. Historical Maps Knowledge Organization: Need, Framework and Practice. J. Libr. Sci. China 2024, 50, 82–95. [Google Scholar] [CrossRef]
  38. Zhong, Y.; Xia, C. A Preliminary Study on the Construction of Knowledge Graph of Art Image. Libr. Trib. 2022, 42, 109–118. [Google Scholar]
  39. Carboni, N.; De Luca, L. An Ontological Approach to the Description of Visual and Iconographical Representations. Heritage 2019, 2, 1191–1210. [Google Scholar] [CrossRef]
  40. Simperl, E. Reusing Ontologies on the Semantic Web: A Feasibility Study. Data Knowl. Eng. 2009, 68, 905–925. [Google Scholar] [CrossRef]
  41. Lu, H.; Hu, Z. An Ontological Modelling Method for Cultural Landscape Genes of Traditional Settlements. J. Geo-Inf. Sci. 2024, 26, 1407–1425. [Google Scholar] [CrossRef]
  42. Gao, S. A Study of the Spatial Layout of Relief Stupa Carvings in the Yungang Grottoes. Southeast Cult. 2025, 113–123. [Google Scholar]
  43. Al-Moslmi, T.; Gallofré Ocaña, M.; Opdahl, A.L.; Veres, C. Named Entity Extraction for Knowledge Graphs: A Literature Overview. IEEE Access 2020, 8, 32862–32881. [Google Scholar] [CrossRef]
  44. Miller, J.J. Graph Database Applications and Concepts with Neo4j. In Proceedings of the Southern Association for Information Systems Conference, Atlanta, GA, USA, 23–24 March 2013. [Google Scholar]
  45. Gennari, J.H.; Musen, M.A.; Fergerson, R.W.; Grosso, W.E.; Crubézy, M.; Eriksson, H.; Noy, N.F.; Tu, S.W. The Evolution of Protégé: An Environment for Knowledge-Based Systems Development. Int. J. Hum.-Comput. Stud. 2003, 58, 89–123. [Google Scholar] [CrossRef]
Figure 1. Hierarchical Model for Image Semantic Description (The top part of the figure illustrates the process by which the model is transformed into a knowledge network).
Figure 1. Hierarchical Model for Image Semantic Description (The top part of the figure illustrates the process by which the model is transformed into a knowledge network).
Buildings 16 02563 g001
Figure 2. Stanford Seven-Step Ontology Development Method and the relationship with the construction of GBNOnto. Note: The left side illustrates the Stanford Seven-Step Ontology Development Method, while the right side shows how each step is implemented in the construction of GBNOnto, including domain coverage, reuse of existing ontologies, a five-tiered knowledge framework, and ontology model instantiation.
Figure 2. Stanford Seven-Step Ontology Development Method and the relationship with the construction of GBNOnto. Note: The left side illustrates the Stanford Seven-Step Ontology Development Method, while the right side shows how each step is implemented in the construction of GBNOnto, including domain coverage, reuse of existing ontologies, a five-tiered knowledge framework, and ontology model instantiation.
Buildings 16 02563 g002
Figure 3. The relations between the class to be extended and the existing ontology. Note: GBNOnto is constructed by reusing and extending concepts from CIDOC CRM, DC, FOAF, and GeoSPARQL. The figure highlights subclass mappings, equivalent classes, and semantic relationships that support the representation of visual, spatial, and cultural-semantic knowledge in grotto Buddhist niche images.
Figure 3. The relations between the class to be extended and the existing ontology. Note: GBNOnto is constructed by reusing and extending concepts from CIDOC CRM, DC, FOAF, and GeoSPARQL. The figure highlights subclass mappings, equivalent classes, and semantic relationships that support the representation of visual, spatial, and cultural-semantic knowledge in grotto Buddhist niche images.
Buildings 16 02563 g003
Figure 4. Schematic diagram of the spatial structure of the Buddhist niche. (a) The Buddhist Niche on the south and west walls of Cave 5 at the Yungang Grottoes. (b) The round arched niche on the west wall of Cave 5 at the Yungang Grottoes and the Primary Buddha Statue.
Figure 4. Schematic diagram of the spatial structure of the Buddhist niche. (a) The Buddhist Niche on the south and west walls of Cave 5 at the Yungang Grottoes. (b) The round arched niche on the west wall of Cave 5 at the Yungang Grottoes and the Primary Buddha Statue.
Buildings 16 02563 g004
Figure 5. Schematic drawing the relations among entities associated with Buddhist niches in Grotto temples. Note: This figure illustrates the Buddhist niches, iconographic themes, such as the round-arched niches, the crossed-ankle Maitreya bodhisattva, the meditating bodhisattva, the Seven Buddhas, and the scene of Two Buddhas Seated Side by Side), as well as their spatial relations (contains, adjacent to, disjoint), temporal relations (earlier than), and semantic relations on the north and east walls of Cave 38 of the Yungang Grottoes. Nodes represent entities, and directed edges denote the relations among them.
Figure 5. Schematic drawing the relations among entities associated with Buddhist niches in Grotto temples. Note: This figure illustrates the Buddhist niches, iconographic themes, such as the round-arched niches, the crossed-ankle Maitreya bodhisattva, the meditating bodhisattva, the Seven Buddhas, and the scene of Two Buddhas Seated Side by Side), as well as their spatial relations (contains, adjacent to, disjoint), temporal relations (earlier than), and semantic relations on the north and east walls of Cave 38 of the Yungang Grottoes. Nodes represent entities, and directed edges denote the relations among them.
Buildings 16 02563 g005
Figure 6. The description of topological relations.
Figure 6. The description of topological relations.
Buildings 16 02563 g006
Figure 7. The description of relative position relations.
Figure 7. The description of relative position relations.
Buildings 16 02563 g007
Figure 8. The Spatial relations between the Buddhist niches and statues on the wall. Note: (a) The “Composite relations” refers to the deliberate design and simultaneous carving of multiple Buddhist niches to convey a unified thematic concept. This relation is characterized by four features: regularity in the overall spatial layout, central symmetry in the spatial arrangement, shared spatial boundaries, and shared decorative elements. (b) The “Yielding relations” occurs when the carving of later Buddhist niches is deliberately adjusted to preserve a certain amount of space (for example, by adjusting the area indicated by the red dotted line), thereby avoiding damage to the original Buddhist niches. (c) The “Disruptive relations” refers to instances where a later-carved niche figure damages a pre-existing one (for example, the carving of niche B disrupted the original carving of niche A).
Figure 8. The Spatial relations between the Buddhist niches and statues on the wall. Note: (a) The “Composite relations” refers to the deliberate design and simultaneous carving of multiple Buddhist niches to convey a unified thematic concept. This relation is characterized by four features: regularity in the overall spatial layout, central symmetry in the spatial arrangement, shared spatial boundaries, and shared decorative elements. (b) The “Yielding relations” occurs when the carving of later Buddhist niches is deliberately adjusted to preserve a certain amount of space (for example, by adjusting the area indicated by the red dotted line), thereby avoiding damage to the original Buddhist niches. (c) The “Disruptive relations” refers to instances where a later-carved niche figure damages a pre-existing one (for example, the carving of niche B disrupted the original carving of niche A).
Buildings 16 02563 g008
Figure 9. The conceptual graphs of the Yungang Cave 38 West Wall Niche an instance (The inscription titled 《吴氏忠伟为亡息冠军将军华口侯吴天恩造像并窟》 is carved above the entrance of Cave 38. This 300-character stone record states that the grotto was excavated and constructed by Wu Zhongwei to pray for the spiritual blessing of her late son Wu Tian’en).
Figure 9. The conceptual graphs of the Yungang Cave 38 West Wall Niche an instance (The inscription titled 《吴氏忠伟为亡息冠军将军华口侯吴天恩造像并窟》 is carved above the entrance of Cave 38. This 300-character stone record states that the grotto was excavated and constructed by Wu Zhongwei to pray for the spiritual blessing of her late son Wu Tian’en).
Buildings 16 02563 g009
Figure 10. The geographical location of the experiment area.
Figure 10. The geographical location of the experiment area.
Buildings 16 02563 g010
Figure 11. Schematic diagram of the Buddhist Niche in a cave.
Figure 11. Schematic diagram of the Buddhist Niche in a cave.
Buildings 16 02563 g011
Figure 12. The screen shot of the GBNOnto in the Protégé environment.
Figure 12. The screen shot of the GBNOnto in the Protégé environment.
Buildings 16 02563 g012
Figure 13. The main entities, properties, and relations defined in the GBNOnto by Protégé.
Figure 13. The main entities, properties, and relations defined in the GBNOnto by Protégé.
Buildings 16 02563 g013
Figure 14. Visual interface of knowledge graph in Neo4j (*(892) denotes the total number of entities, and *(2621) represents the total number of relational triples in the knowledge graph).
Figure 14. Visual interface of knowledge graph in Neo4j (*(892) denotes the total number of entities, and *(2621) represents the total number of relational triples in the knowledge graph).
Buildings 16 02563 g014
Figure 15. Search results of the knowledge graph of Yungang Cave 38 (Area1 is a capture of the knowledge graph showing entity names and relations; Area2 and Area3 show examples of property values of the entities).
Figure 15. Search results of the knowledge graph of Yungang Cave 38 (Area1 is a capture of the knowledge graph showing entity names and relations; Area2 and Area3 show examples of property values of the entities).
Buildings 16 02563 g015
Figure 16. The topological relations among the structural component elements of the Buddhist Niche (The asterisk (*) followed by a number indicates the quantity of graph elements. * (39) = total entities; * (10) = total relationships).
Figure 16. The topological relations among the structural component elements of the Buddhist Niche (The asterisk (*) followed by a number indicates the quantity of graph elements. * (39) = total entities; * (10) = total relationships).
Buildings 16 02563 g016
Table 1. The list of main classes and the class hierarchy.
Table 1. The list of main classes and the class hierarchy.
ClassSubclassDescription
E90 Symbolic ObjectBuddhist NicheA small rock-cut cell within a grotto, designed to house a Buddha statue, mural, and decorative components, serves as a basic unit of Buddhist worship space
Structural ComponentArchitectural components of a niche, including the niche lintel, wall surface, pedestal, pillar base, etc., used to support and decorate the niche space
Primary Buddha StatueThe main Buddha statue in the center of a niche, usually representing a specific Buddhist theme, religious symbolism, and carving style
Other StatuesAuxiliary Buddha or protector deity statues surrounding the principal deity, used to express relationships of protection, offering, or religious ritual
Decorative PatternsDecorative elements on the niche and wall surfaces, such as motifs, lotus petals, rosette flowers, cloud patterns, etc., reflecting artistic style and religious symbolism
InscriptionNiche inscriptions, dedicatory records, signatures, and carved texts recording information about donors, carvers, and production dates
foaf: AgentPersonHistorical figures, artisans, or religious persons involved in the design, carving, painting, or inscription of niches
GroupTeams of artisans or religious groups that collaboratively completed niche construction, mural painting, or decoration
OrganizationOrganizational units such as temples, monastic communities, or workshops responsible for niche construction, management, and maintenance
E53 PlaceLocationSpecific spatial position of a niche
Spatial HierarchyDividable spatial units inside a cave or niche, e.g., upper space, lower space, etc.
Wall SurfaceWall area of a cave chamber where a niche is located
CaveThe cave chamber entity to which a niche belongs, used to determine spatial affiliation
E5 EventE7 ActivityActivities related to niche construction, restoration, enshrinement, and religious rituals
E65 CreationCreative events such as Buddha carving, mural painting, decorative pattern engraving, and inscription writing
E55 TypeNiche TypologyClassification of niche forms, e.g., round-arched niche, caisson-shaped niche, square niche, etc.
Decorative PatternsTypes of decorative patterns on niches and wall surfaces, including lotus motifs, cloud patterns, rosette flowers, etc.
Themes of Buddhist SculptureIconographic themes of niche Buddha statues, such as Shakyamuni, Maitreya, Avalokitesvara, Manjusri, etc.
Buddhist narrativeReligious stories, sutra illustrations, or historical legends represented by niches and statues
Artistic styleCarving or painting styles of different periods or regions, e.g., high-relief style of the late Northern Wei, realistic style of the early Tang, etc.
Function TypeReligious functions of a niche, such as enshrinement, preaching, pilgrimage, or worship
Engraving TechniquesCarving and engraving techniques for statues and decorations
Disease TypeTypes of damage, exfoliation, weathering, cracks, etc., used for recording preservation status
DynastyHistorical dynasty corresponding to a niche or statue, e.g., Northern Wei, Tang, Song, etc.
E2 Temporal EntityE52 Time SpanTime span during which a niche or statue existed, was built, or was produced
E4 PeriodSpecific historical period used to classify niche artistic styles and production contexts
E73 Information ObjectE38 ImageDigital image of a niche
E102 Evidence SourceLiterature, archaeological reports, archival materials, etc., that support knowledge representation of niches
Geo: SpatialObject/Equivalent class to E22 Human-Made Object, representing spatial entities
E22 Human-Made ObjectCavePhysical cave chamber inside a grotto, e.g., Cave 38, Northern Wei cave, etc.
Wall SurfaceWall area of a cave chamber where a niche is placed, used to place Buddha statues and decorations
Grotto TempleThe overall architectural complex of cave temples, where niches are located
Table 2. The properties of class Buddhist Niche and class Primary Buddha statue.
Table 2. The properties of class Buddhist Niche and class Primary Buddha statue.
LabelDomainRangeExample
NameGBNOnto: Buddhist NicheE62StringCave 38
GBNOnto: CaveE62StringThe Round-arched niche on the north wall
GBNOnto: Primary Buddha StatueE62StringShakyamuni Buddha
IDGBNOnto: Buddhist NicheE62StringYG-38-BB-YGK-01
GBNOnto: CaveE62StringYG-38
Alternative NameGBNOnto: Buddhist NicheE62StringLu-style niche, also known as Lintel-arched niche
TypeGBNOnto: Buddhist NicheE62StringRound-arched niche
GBNOnto: Primary Buddha StatueE62StringTwo Buddhas seated together
GBNOnto: DecorationE62StringHoneysuckle pattern
FunctionGBNOnto: Buddhist NicheE62StringRitual worship
PeriodGBNOnto: Buddhist NicheE62StringEarly period; Late period
HeightGBNOnto: Buddhist NicheE60Number850 cm
WidthGBNOnto: Buddhist NicheE60Number120 cm
DepthGBNOnto: Buddhist NicheE60Number60 cm
DressGBNOnto: Primary Buddha StatueE62StringFull-shoulder monastic robe
HairstyleGBNOnto: Primary Buddha StatueE62StringHigh spiral uṣṇīṣa
Facial expressionGBNOnto: Primary Buddha StatueE62StringPlump and rounded countenance
Body postureGBNOnto: Primary Buddha StatueE62StringThe sitting posture
Hand postureGBNOnto: Primary Buddha StatueE62StringDhyana Mudra
Preservation stateGBNOnto: Buddhist NicheE62StringWeathered
GBNOnto: Primary Buddha StatueE62StringWeathered
Table 3. Temporal Relations.
Table 3. Temporal Relations.
NameDefinitionDiagram
Earlier thanA occurs before B Buildings 16 02563 i001
Later thanA occurs after BBuildings 16 02563 i002
Simultaneous withA and B occur simultaneouslyBuildings 16 02563 i003
DuringA occurs during BBuildings 16 02563 i004
B occurs during ABuildings 16 02563 i005
Table 4. Representative RDF Triples Extracted from Cave 38 of the Yungang Grottoes.
Table 4. Representative RDF Triples Extracted from Cave 38 of the Yungang Grottoes.
Entity1RelationEntity2Description
Cave 38hasDynastyNorthern Wei DynastyCave 38 is located in the western section of the central grotto cluster of Yungang Grottoes and is a representative cave of the Northern Wei period, belonging to the overall Yungang Grotto system.
Cave 38hasShaperectangular planCave 38 has a rectangular floor plan with a truncated pyramid ceiling, reflecting the spatial layout and architectural form of grottoes at that time.
Buddhist NichehasPartNiche Lintel; Niche Pillars; Niche BeamThe Buddhist niche is composed of architectural components such as the niche lintel, niche pillars, and niche beam, reflecting the structural organization and decorative framework of the niche.
Primary Buddha StatuehasGesturepreaching/seated postureThe seated Buddha sits in a cross-legged posture and displays the preaching gesture, demonstrating the religious pose and function of the statue.
Buddhist NicheincludesDonor FiguresFlanking the main niche are rows of donor figures, reflecting the connection between Buddhist sculpture and social belief.
Buddhist NichehasDecorativePatternlotus motifThe niche is decorated with lotus motifs and honeysuckle patterns, enriching the aesthetic details of the space.
Wu Tian’enhasRoleGeneralWu Tian’en held the role of a general, as recorded in the inscription, indicating his social status and involvement in the construction activities.
InscriptionrecordsConstruction ActivityInscriptions within the niche record the creation of statues and cave excavation activities, providing direct evidence for historical verification.
Buddhist NichehasLocationWest WallThe Buddhist niche is located on the west wall, and its sculptural layout reflects the spatial organization of the cave interior.
Primary Buddha StatuehasIconographyTwo Buddhas Seated Side by SideThe primary statue depicts the iconography of two Buddhas seated side by side, reflecting specific religious themes in Buddhist art.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Wan, L.; Hou, M.; Li, J.; Zhao, B.; Yang, B.; Shi, H.; Ning, B. Knowledge Representation Method for Grotto Buddhist Niches Based on Image Semantics and Ontology. Buildings 2026, 16, 2563. https://doi.org/10.3390/buildings16132563

AMA Style

Wan L, Hou M, Li J, Zhao B, Yang B, Shi H, Ning B. Knowledge Representation Method for Grotto Buddhist Niches Based on Image Semantics and Ontology. Buildings. 2026; 16(13):2563. https://doi.org/10.3390/buildings16132563

Chicago/Turabian Style

Wan, Li, Miaole Hou, Jinru Li, Beibei Zhao, Bingyu Yang, Haoyue Shi, and Bo Ning. 2026. "Knowledge Representation Method for Grotto Buddhist Niches Based on Image Semantics and Ontology" Buildings 16, no. 13: 2563. https://doi.org/10.3390/buildings16132563

APA Style

Wan, L., Hou, M., Li, J., Zhao, B., Yang, B., Shi, H., & Ning, B. (2026). Knowledge Representation Method for Grotto Buddhist Niches Based on Image Semantics and Ontology. Buildings, 16(13), 2563. https://doi.org/10.3390/buildings16132563

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop