1. Introduction
Addressing data integration from both conceptual and methodological perspectives represents one of the central challenges in the development of information systems for the cultural heritage domain [
1,
2]. The rapid growth of digital documentation has significantly increased the availability of heterogeneous datasets describing historical sites, objects, conservation states, and associated documentation [
3,
4]. Ensuring that these datasets adhere to FAIR (Findable, Accessible, Interoperable, Reusable) data principles is essential to maximize their long-term usability and cross-domain integration [
5].
While this expansion offers unprecedented opportunities for research and conservation planning, it also introduces substantial challenges related to interoperability, data integration, and semantic consistency. A notable example of efforts to address these challenges is the digital ecosystem developed for the multidisciplinary study of Notre-Dame de Paris [
6]. In the context of the cathedral’s restoration, De Luca and his team designed a digital infrastructure capable of managing and correlating data originating from multiple domains, including architectural surveys, historical documentation, material analyses, and conservation studies. This project illustrates how carefully designed digital frameworks can support the integration of diverse datasets while preserving the specificity of each disciplinary contribution. Within this line of research, De Luca identifies four major gaps that continue to affect digital cultural heritage practices: the semantic gap, referring to the fragmentation and insufficient semantic enrichment of digital resources; the memory gap, related to the difficulty of formalizing implicit knowledge and research protocols; the data correlation gap, which concerns the limited capability to automatically link heterogeneous datasets; and the technological gap, associated with the difficulty of connecting collaborative digital analysis with the physical heritage context.
From a complementary perspective, Bruseker et al. [
7] emphasize the importance of formal ontologies—particularly CIDOC-CRM [
8]—as a foundation for sustainable cultural heritage data management. Ontologies provide the conceptual clarity necessary to mediate between domain expertise and computational systems, enabling long-term interoperability, semantic consistency, and inferential capabilities. This approach has been operationalized in mapping models such as the Swiss Art Research Data Model (SRDM) [
9], which translates CIDOC-CRM into structured mapping guidelines and practical data transformation workflows [
10].
A further layer of complexity arises from the fact that cultural heritage documentation is typically produced across multiple institutional contexts and disciplinary perspectives. As a result, documentation practices vary widely in structure, terminology, and conceptual frameworks, making the integration of heterogeneous datasets particularly challenging for digital heritage infrastructures. These issues become even more pronounced in the field of cultural heritage risk management, where documentation processes often involve multiple actors and professional perspectives. Risk assessment typically requires the collaboration of emergency responders, conservation specialists, heritage authorities, and administrative bodies, each of whom documents risk scenarios according to their own operational priorities. For example, emergency responders may emphasize factors such as fire propagation, structural vulnerability, or accessibility in emergency situations, whereas heritage professionals may priorities conservation conditions, material fragility, or historical significance. Although this plurality of perspectives represents a valuable source of knowledge, it also introduces significant challenges for data integration.
Recent developments in digital cultural heritage research have increasingly explored the use of semantic infrastructures, knowledge graphs, and digital twins as mechanisms for integrating heterogeneous information and supporting advanced analytical processes. Digital twin approaches have been proposed as a means of linking physical heritage assets with continuously updated digital representations, enabling more dynamic forms of monitoring, interpretation, and decision support [
11,
12]. Similarly, recent work on semantic data modelling and knowledge graph construction has highlighted the importance of formal ontologies and reusable semantic frameworks for ensuring long-term interoperability, transparency, and sustainability of cultural heritage information systems [
10]. These developments demonstrate the growing maturity of semantic technologies within the cultural heritage domain. However, their application to cultural heritage risk assessment and preventive conservation remains comparatively limited. In particular, the integration of risk evaluation records, environmental threats, and conservation-oriented decision-making processes within interoperable semantic infrastructures is still insufficiently explored, especially at the territorial scale. This gap motivates the present study.
While CIDOC-CRM, X3ML-based transformations, and Linked Open Data infrastructures have been extensively applied to cultural heritage documentation, their application to cultural heritage risk assessment remains comparatively underexplored. In particular, the semantic formalization of structured risk assessment methodologies and the representation of risk scenarios within an event-centric ontological framework and their integration with national heritage knowledge infrastructures have received limited attention. The originality of the present study lies in the development of a reproducible semantic workflow that bridges cultural heritage risk assessment practices and interoperable knowledge graph technologies, enabling the integration of ABC-based risk evaluation data within a CIDOC-CRM-compliant environment.
This study proposes the adoption of the CIDOC Conceptual Reference Model (CIDOC-CRM) for the semantic integration of cultural heritage risk assessment data [
1]. By providing a formal ontology capable of modelling events, actors, objects, conditions, and their interrelations, CIDOC-CRM enables heterogeneous documentation practices to be aligned within a shared conceptual framework. Such an approach makes it possible to reconcile different institutional viewpoints while preserving their specificity, thereby supporting more coherent and interoperable cultural heritage risk management systems. Furthermore, the ontological nature of CIDOC-CRM not only enables data harmonization but also reasoning and knowledge inference. By formally representing actors, events, interventions, and conservation states within an explicit semantic structure, the model allows implicit relationships between heritage assets, risk factors, and evaluation processes to be analyzed through advanced queries and analytical reasoning, thus contributing to more informed decision-making processes and to the development of integrated, data-driven approaches to cultural heritage risk management. In this context, the present study proposes a semantic workflow for integrating cultural heritage risk assessment data within a CIDOC-CRM-compliant knowledge infrastructure.
3. Results
3.1. Generation of the Cultural Heritage Risk Knowledge Graph
The semantic workflow described in the previous section was applied to the dataset of cultural heritage sites located in the municipality of Ravenna. Through the ETL pipeline, the harmonized XML records derived from the ABC risk assessment forms were transformed into RDF triples compliant with the CIDOC-CRM ontology.
The transformation process resulted in the generation of a semantic knowledge graph representing the relationships between cultural heritage sites, risk scenarios, actors involved in risk processes, and evaluation data. Each cultural heritage asset is represented as an entity within the graph and connected to documentary information, risk events, and measurement activities through explicitly defined semantic relationships.
The resulting dataset (
Table 1) includes 134,611 RDF triples and 18,954 unique entities, structured across 295 cultural heritage sites, but currently contains only 36 actors and 86 events, indicating a relatively sparse CIDOC CRM event–actor representation within the knowledge graph.
These entities include cultural heritage assets, risk events affecting the assets, agents responsible for the risk conditions, and documentary resources describing the evaluation process. The RDF graphs produced through the X3ML transformation were successfully imported into the ResearchSpace environment, where they can be explored and queried through semantic interfaces and SPARQL queries.
The successful deployment of the resulting RDF graphs within the ResearchSpace environment demonstrates not only the feasibility of the proposed ETL workflow, but also its capacity to operationalize heterogeneous risk assessment records as semantically structured and queryable knowledge. By transforming fragmented documentation into an interconnected graph of cultural assets, risk events, actors, measurements, and documentary resources, the workflow establishes the conditions for advanced analytical exploration, cross-dataset integration, and decision-support activities. Furthermore, the use of persistent identifiers and ontology-based mappings enables the federation of local risk assessment data with external cultural heritage knowledge infrastructures, extending the analytical potential of the resulting dataset beyond the boundaries of the original project. An additional contribution of the workflow lies in its reproducibility. By explicitly formalizing the semantic relationships between source records and ontology classes through the 3M mapping environment and the X3ML transformation engine, the methodology transforms domain knowledge and documentation logic into reusable and transparent semantic rules. This approach not only facilitates the consistent generation of RDF knowledge graphs from heterogeneous datasets but also contributes to reducing dependence on implicit expertise embedded in traditional documentation practices, thereby supporting more sustainable and interoperable cultural heritage risk management infrastructures.
3.2. Event-Based Representation of Cultural Heritage Risk
One of the central outcomes of the proposed modelling approach is the event-based representation of cultural heritage risk (
Figure 2). Within the CIDOC-CRM framework, risk is not treated as a static property associated with a cultural asset, but rather as an event that may affect or transform the asset.
This modelling strategy enables a structured representation of the relationships between cultural assets, hazards, and evaluation processes. By explicitly modelling risk as a network of interconnected events, actors, and measurements, the knowledge graph supports more advanced analytical queries, such as the identification of recurring risk patterns, the comparison of risk scenarios across different sites, or the analysis of risk agents affecting specific categories of heritage assets.
3.3. Interoperability with External Cultural Heritage Datasets
Another significant result of the workflow concerns the semantic interoperability achieved through the integration of authoritative identifiers and external reference datasets. During the harmonization phase, each cultural heritage site record was enriched with the Codice Catalogo Nazionale, corresponding to the identifier assigned within the Italian national catalogue of cultural heritage.
Through this identifier, the knowledge graph establishes semantic links with the national dataset published via the ArCo ontology, which represents the official RDF implementation of the Italian cultural heritage catalogue. The mapping schema therefore enables the federation of the local risk dataset with the national cultural heritage knowledge graph, allowing cross-dataset queries and the integration of additional descriptive information about the cultural assets. In practical terms, this interoperability is implemented through CIDOC-CRM property alignments such as the mapping of P2 has type to E55 Type, where the typology values are directly aligned with the corresponding controlled IRI terms provided by ArCo.
This interoperability mechanism demonstrates how local risk documentation systems can be integrated within broader Linked Open Data ecosystems, enhancing the reusability and contextualization of risk assessment data. By linking the Ravenna dataset to external cultural heritage resources, the proposed workflow contributes to building a more interconnected digital infrastructure for cultural heritage documentation and risk management.
3.4. Example Semantic Query and Knowledge Retrieval
To demonstrate the analytical potential of the resulting knowledge graph, semantic queries can be used to retrieve and correlate information concerning cultural heritage assets, risk events, risk agents, and associated evaluation data. Unlike traditional record-based documentation systems, the semantic structure of the graph enables the exploration of explicit relationships between entities and supports cross-dataset analysis. The following example illustrates a typical risk-analysis query executed within the ResearchSpace environment.
Query 1—Risk event and agent retrieval (comparative analysis)
This component retrieves modification events associated with a given heritage asset and identifies the involved risk agents, enabling the comparative analysis of actor–event relationships across the dataset as illustrated in
Figure 3.
Query 2—Risk magnitude estimation and prioritization
This component computes an aggregated risk score per event based on multiple measurement dimensions. The resulting values can be ordered to support the prioritization of events or assets according to their estimated risk level.
Figure 4 illustrates the decision-support view. The graph ranks the risk events associated with a selected heritage asset according to their aggregated risk magnitude, with color coding highlighting relative severity (green to red). This enables conservation professionals to quickly identify the most critical risk factors and priorities interventions, inspections, or resource allocation based on quantitative evidence derived from the knowledge graph.
Query 3—Uncertainty assessment (decision reliability)
This component extracts minimum and maximum uncertainty values associated with risk measurements, allowing the evaluation of confidence intervals and supporting uncertainty-aware decision-making as represented in
Figure 5.
The provided examples illustrate how semantically modelled relationships between heritage sites, risk events, agents, and measurements can be queried across the entire dataset. Such queries support comparative risk analysis, facilitate the identification of recurring patterns and critical risk factors, and contribute to evidence-based decision-making processes for preventive conservation and territorial risk management.
The examples presented above demonstrate not only the technical feasibility of the semantic workflow, but also its potential to support interoperable knowledge management, semantic querying, and decision-support activities in cultural heritage risk assessment. These aspects are further discussed in the following section.
The full SPARQL query implementation used in this study is publicly available [
36].
4. Discussion
4.1. Bridging the Semantic Gap in Cultural Heritage Risk Data
The results obtained through the implementation of the proposed workflow demonstrate the potential of ontology-based infrastructures to address the semantic fragmentation that characterizes cultural heritage risk documentation. As highlighted in recent scholarship, the extensive production of digital resources in the heritage domain does not automatically translate into interoperable knowledge systems. In many cases, documentation remains fragmented across institutions, disciplinary traditions, and data formats, limiting the possibility of integrating and analyzing information at broader spatial or institutional scales.
The semantic workflow proposed in this study contributes to addressing this issue by transforming heterogeneous risk assessment records into a structured knowledge graph based on the CIDOC-CRM ontology. Through the formal mapping of entities such as cultural heritage assets, risk events, actors, and evaluative measurements, the approach enables previously disconnected documentation practices to be integrated within a shared conceptual framework. In this sense, the workflow does not simply harmonize data structures; it introduces a semantic layer that makes the relationships between heritage assets, hazards, and evaluation processes explicit and machine-readable
The event-based modelling of risk plays a particularly important role in this process. By conceptualizing risk scenarios as events affecting cultural assets, the model captures the dynamic and processual nature of risk conditions. This representation allows heterogeneous risk factors and evaluation processes to be integrated within a coherent analytical structure, enabling the comparison of risk scenarios across multiple sites and facilitating the exploration of patterns that would remain difficult to identify within traditional documentation systems.
Compared with conventional relational database approaches, the proposed semantic workflow offers several advantages for the integration and analysis of cultural heritage risk data. Relational databases are highly effective for storing structured information, but they generally rely on predefined schemas and custom integration procedures when heterogeneous datasets need to be combined. In contrast, the ontology-based approach adopted in this study provides a shared conceptual framework capable of accommodating different documentation practices, vocabularies, and institutional perspectives. Furthermore, semantic querying through RDF and SPARQL enables the exploration of explicit relationships between cultural assets, risk events, actors, and evaluation processes, supporting analytical questions that are often difficult to formulate within conventional relational structures. The integration of authoritative identifiers and Linked Open Data resources further enhances interoperability by facilitating the federation of local datasets with external knowledge infrastructures.
4.2. Formalizing Documentation Practices and Addressing the Memory Gap
Beyond semantic interoperability, the proposed methodology also contributes to addressing what De Luca [
6] describes as the memory gap in digital cultural heritage practices. In many heritage documentation workflows, key elements of interpretation, classification, and evaluation remain embedded in implicit professional knowledge rather than being formally represented within the data structure. This reliance on tacit expertise can hinder the reproducibility of documentation processes and complicate the long-term reuse of digital resources.
The semantic modelling strategy adopted in this study seeks to reduce this dependence on implicit knowledge by explicitly representing documentation logic within an ontological framework. Through the mapping schema defined in the 3M editor and executed via the X3ML engine, the conceptual relationships underlying the risk assessment process become formally encoded within the transformation rules that generate the knowledge graph. In this way, the methodological assumptions embedded in the ABC risk assessment workflow are translated into explicit semantic structures that can be shared, reused, and extended in other contexts.
This explicit modelling of documentation practices has important implications for the sustainability of digital heritage infrastructures. By transforming procedural knowledge into formal semantic representations, the workflow supports the development of reproducible and transparent data integration pipelines. Such an approach contributes to reducing the dependence on highly personalised expertise and facilitates collaborative data production across institutions and disciplinary communities.
4.3. Metadata Quality and Limitations of the Current Dataset
The experimentation conducted on the Ravenna dataset has also highlighted several limitations related to the quality and consistency of the underlying metadata. Inconsistencies in typological classifications within the national cultural heritage catalogue have occasionally generated redundancy or ambiguity in the semantic typing of cultural heritage sites. Similar issues emerged in the geocoding process, where non-standardized address expressions or incomplete location data prevented the automatic generation of geographic coordinates for some records.
These issues underline an important methodological consideration: while semantic infrastructures can significantly enhance interoperability and knowledge integration, their effectiveness ultimately depends on the quality and consistency of the source data. Semantic modelling and ontology-based mapping can mitigate some forms of heterogeneity, but they cannot fully compensate for structural inconsistencies or incomplete metadata at the source level.
Consequently, the implementation of semantic integration workflows should be accompanied by systematic metadata governance strategies, including the adoption of controlled vocabularies, standardized documentation practices, and continuous data quality validation. In this sense, the experience described in this study suggests that the development of interoperable semantic infrastructures must be understood not only as a technological challenge but also as an organizational and institutional one, requiring coordinated efforts in data curation and standardization across cultural heritage institutions.
The present study should be regarded as a proof-of-concept implementation aimed at demonstrating the feasibility of integrating cultural heritage risk assessment data within a CIDOC-CRM-compliant semantic infrastructure. The primary objective was the design, implementation, and validation of a semantic workflow capable of transforming heterogeneous risk assessment records into interoperable knowledge graphs. Consequently, the research focused on data harmonization, ontology-based modelling, and semantic integration rather than on the evaluation of user interaction or system performance. Formal usability assessments, user-centered evaluations, and systematic benchmarking of query performance were therefore beyond the scope of the present work. Future developments will focus on the operational deployment of the infrastructure and its validation through real-world use cases involving cultural heritage professionals and institutional stakeholders, with particular attention to usability, scalability, query performance, and decision-support effectiveness.
Although the workflow was developed within the context of the SIRIUS and RESTART projects and relies on the ABC Method, ArCo, and ICCD documentation standards, its overall architecture is not intrinsically tied to these specific frameworks. The event-centric semantic model provided by CIDOC-CRM enables the representation of risk assessment processes at a conceptual level, allowing alternative methodologies to be accommodated through the remapping of evaluation entities, measurement procedures, and domain-specific vocabularies. Similarly, while ArCo and ICCD were adopted because they represent the official Italian cultural heritage information infrastructure, the proposed workflow is based on modular semantic components and ontology mappings that could be adapted to other national or institutional cataloguing systems, provided that appropriate semantic correspondences and persistent identifiers are available.
From an organizational perspective, the workflow is particularly suited to contexts where heterogeneous documentation sources need to be integrated and analyzed across multiple cultural heritage assets, such as regional heritage authorities, national heritage agencies, research institutions, and multi-site conservation programmes. Smaller institutions, including museums and local heritage organizations, could also benefit from the approach through shared infrastructures or consortium-based implementations. However, its adoption requires expertise in ontology modelling, semantic technologies, and data governance, highlighting the importance of interdisciplinary collaboration between domain experts, heritage professionals, and information specialists.
5. Conclusions
This study presented a semantic workflow for the integration and operationalization of cultural heritage risk assessment data within a CIDOC-CRM-compliant environment. By combining data harmonization and ontology-based semantic modelling, the proposed approach enables the conversion of heterogeneous ABC risk assessment records into interoperable RDF knowledge graphs.
The implementation carried out on the Ravenna dataset demonstrates the feasibility and scalability of the workflow. Through the integration of controlled vocabularies, project-specific thesauri, and formal ontology mapping, the system allows cultural heritage sites, risk events, agents, and evaluation processes to be represented within a coherent semantic structure. The resulting knowledge graph supports advanced querying, cross-dataset comparison, and the integration of risk documentation within broader cultural heritage data ecosystems.
The contribution of this study does not lie in the individual technologies employed, but in their integration within a coherent workflow specifically designed for cultural heritage risk assessment and semantic interoperability. Beyond the technical implementation, the study contributes to the broader discourse on semantic infrastructures for cultural heritage by addressing two key challenges identified in recent scholarship: the semantic gap and the memory gap in digital heritage documentation. By formalizing risk assessment practices within a transparent ontological framework, the proposed workflow transforms fragmented documentation records into structured, interoperable knowledge and reduces the reliance on implicit domain expertise embedded in traditional documentation processes.
At the same time, the experimentation highlights the importance of metadata quality and governance as critical factors for the success of semantic integration initiatives. While ontology-based infrastructures can significantly enhance interoperability, their effectiveness ultimately depends on the consistency and reliability of the underlying documentation practices.
Future developments may extend the proposed framework to larger territorial datasets and integrate additional analytical tools for risk evaluation and decision support. In this perspective, the adoption of ontology-based semantic infrastructures may play an increasingly important role in supporting coordinated, data-driven approaches to cultural heritage risk management at the territorial scale. In particular, such infrastructures should be further enhanced through the integration of a description logic layer, extending the existing RDF model into a well-defined OWL-DL graph, thereby enabling more robust semantic reasoning, consistency checking, and advanced inferencing capabilities. Additionally, the incorporation of Large Language Models (LLMs) may further support the reasoning process by enabling natural language interaction with the knowledge base, facilitating semantic querying, and assisting in the interpretation and contextualization of heterogeneous data sources. When combined with formal ontology-based reasoning, LLMs can contribute to hybrid approaches that enhance decision-making processes, support knowledge discovery, and improve the accessibility of complex cultural heritage information systems for both experts and non-expert users.
From a theoretical perspective, the study currently contributes to ongoing discussions on semantic interoperability in cultural heritage by demonstrating how risk assessment data—traditionally fragmented across disciplinary and institutional boundaries—can be formalized within a shared ontological framework. The proposed workflow extends existing applications of CIDOC-CRM beyond documentation and inventory management, showing its potential to support the representation of risk processes, evaluation activities, and decision-support knowledge structures. From a practical perspective, the workflow provides cultural heritage institutions with a reproducible framework for integrating heterogeneous risk assessment datasets, improving data consistency, facilitating information exchange, and supporting evidence-based conservation planning. By enabling semantic querying and interoperability with external knowledge infrastructures, the approach may contribute to more coordinated and informed risk management strategies at both site and territorial scales.
Although developed within the Italian cultural heritage context, the proposed workflow is based on modular semantic components and can be adapted to alternative documentation frameworks and cataloguing systems. Future work will focus on operational deployment, performance evaluation, and user-centered validation.