Next Article in Journal
MSFA: Multi-Strategy Fusion Algorithm for Data Cleaning and Its Application in Offshore Marine Environmental Monitoring
Previous Article in Journal
Cognitive Big Data Architecture for Daily Operational Jamming Transition Detection with Low-Latency Inference in Infrastructure-Constrained Financial Markets: The MERI Framework
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
This is an early access version, the complete PDF, HTML, and XML versions will be available soon.
Article

A Model for Metadata Organisation and Management for Compilation of Specialised Datasets from Big Data

Department of Computational Linguistics, Institute for Bulgarian Language, Bulgarian Academy of Sciences, 52 Shipchenski Prohod Blvd., Sofia 1113, Bulgaria
*
Author to whom correspondence should be addressed.
Big Data Cogn. Comput. 2026, 10(7), 241; https://doi.org/10.3390/bdcc10070241
Submission received: 26 May 2026 / Revised: 1 July 2026 / Accepted: 7 July 2026 / Published: 17 July 2026
(This article belongs to the Section Big Data)

Abstract

The paper presents a model for the design and management of metadata that enables the efficient compilation of specialised datasets from large, heterogeneous data collections. The metadata are represented as a typed property graph that facilitates the FAIR principles in data compilation: Findable, Accessible, Interoperable, Reusable. The representation is general and independent of the modality and format of the data. The graph-based design of the metadata supports the incremental extension of categories and relationships without requiring the migration of existing data. Its feasibility is demonstrated through the use of a graph database, in which the metadata for 689,645 Bulgarian textual data units are combined with a web-based filtering interface. Metadata retrieval is implemented through Cypher queries executed as graph traversals, enabling the extraction of thematic and application-oriented data subsets based on combinations of selection criteria. The application validates the suitability of the graph-based metadata design for compiling specialised datasets for training and fine-tuning large language models and other NLP applications.
Keywords: metadata; metadata design; graph database; big data; large language models metadata; metadata design; graph database; big data; large language models

Share and Cite

MDPI and ACS Style

Koeva, S.; Stoyanova, I. A Model for Metadata Organisation and Management for Compilation of Specialised Datasets from Big Data. Big Data Cogn. Comput. 2026, 10, 241. https://doi.org/10.3390/bdcc10070241

AMA Style

Koeva S, Stoyanova I. A Model for Metadata Organisation and Management for Compilation of Specialised Datasets from Big Data. Big Data and Cognitive Computing. 2026; 10(7):241. https://doi.org/10.3390/bdcc10070241

Chicago/Turabian Style

Koeva, Svetla, and Ivelina Stoyanova. 2026. "A Model for Metadata Organisation and Management for Compilation of Specialised Datasets from Big Data" Big Data and Cognitive Computing 10, no. 7: 241. https://doi.org/10.3390/bdcc10070241

APA Style

Koeva, S., & Stoyanova, I. (2026). A Model for Metadata Organisation and Management for Compilation of Specialised Datasets from Big Data. Big Data and Cognitive Computing, 10(7), 241. https://doi.org/10.3390/bdcc10070241

Article Metrics

Back to TopTop