Journal Description
Data
Data
is a peer-reviewed, open access journal on data in science, with the aim of enhancing data transparency and reusability. The journal publishes in two sections: a section on the collection, treatment and analysis methods of data in science; a section publishing descriptions of scientific and scholarly datasets (one dataset per paper). The journal is published monthly online by MDPI.
- Open Access— free for readers, with article processing charges (APC) paid by authors or their institutions.
- High Visibility: indexed within Scopus, ESCI (Web of Science), Ei Compendex, dblp, Inspec, RePEc, and other databases.
- Journal Rank: JCR - Q2 (Multidisciplinary Sciences) / CiteScore - Q1 (Information Systems and Management)
- Rapid Publication: manuscripts are peer-reviewed and a first decision is provided to authors approximately 19.2 days after submission; acceptance to publication is undertaken in 3.8 days (median values for papers published in this journal in the first half of 2026).
- Recognition of Reviewers: reviewers who provide timely, thorough peer-review reports receive vouchers entitling them to a discount on the APC of their next publication in any MDPI journal, in appreciation of the work done.
- Journal Cluster of Information Systems and Technology: Analytics, Applied System Innovation, Cryptography, Data, Digital, Informatics, Information, Journal of Cybersecurity and Privacy and Multimedia.
Impact Factor:
2.4 (2025);
5-Year Impact Factor:
2.5 (2025)
Latest Articles
A Device-Level IoT Network Traffic Dataset with Distributed Capture and Non-IID Characteristics
Data 2026, 11(8), 207; https://doi.org/10.3390/data11080207 - 14 Aug 2026
Abstract
The development of intrusion detection and network security solutions for securing Internet of Things (IoT) networks is constrained by the limited availability of representative network security datasets. Many existing datasets rely on centralised traffic collection and do not capture the non-Independent and Identically
[...] Read more.
The development of intrusion detection and network security solutions for securing Internet of Things (IoT) networks is constrained by the limited availability of representative network security datasets. Many existing datasets rely on centralised traffic collection and do not capture the non-Independent and Identically Distributed (non-IID) characteristics inherent to edge environments. To address this limitation, this work presents a device-level IoT network dataset generated using the open-source Gotham testbed, a virtualised smart city environment. Network traffic is collected in a distributed manner at the interfaces of 78 heterogeneous IoT devices operating across multiple protocols, including MQTT, CoAP, and RTSP. The dataset comprises over 31.8 million packet-level records, each described by 22 features. It includes both benign traffic and multiple attack classes, namely Network Scanning, Brute Force, Infection, Denial of Service (DoS), and Command and Control (C&C) Communication. Ground-truth labels are assigned using a deterministic process based on orchestration logs. The dataset preserves device-level traffic distributions and captures non-IID characteristics without artificial partitioning. It is publicly available and can be used to support reproducible evaluation of intrusion detection approaches and network analysis tasks in both centralised and distributed learning settings.
Full article
(This article belongs to the Section Information Systems and Data Management)
Open AccessData Descriptor
A Multi-Metric NDVI-Derived Dataset of Vegetation Dynamics and Land Surface Phenology in Southern and Central Europe (1982–2022)
by
Caterina Samela, Maria Lanfredi, Rosa Coluzzi and Vito Imbrenda
Data 2026, 11(8), 206; https://doi.org/10.3390/data11080206 - 11 Aug 2026
Abstract
This Data Descriptor presents a value-added suite of derived vegetation phenology products for Southern and Central Europe (10° W–28° E; 35° N–50° N), generated from the PKU GIMMS NDVI v1.2 archive (1982–2022) through a standardized processing workflow including quality screening, temporal compositing, phenological
[...] Read more.
This Data Descriptor presents a value-added suite of derived vegetation phenology products for Southern and Central Europe (10° W–28° E; 35° N–50° N), generated from the PKU GIMMS NDVI v1.2 archive (1982–2022) through a standardized processing workflow including quality screening, temporal compositing, phenological metric extraction, eco-phenological regionalization, and variability analysis. The final collection includes 82 GeoTIFF raster layers and one CSV file, structured into long-term monthly NDVI climatologies, decadal NDVI composites, eco-phenological regionalization, Land Surface Phenology (LSP) metrics with inter-decadal shift layers, and a Phenology Variability Index (PVI). In addition, cluster-level Mann–Kendall and Theil–Sen trend statistics are provided in tabular form. All products are distributed in WGS84 (EPSG:4326) at 0.0833° spatial resolution. The dataset provides ready-to-use vegetation phenology indicators, variability metrics, and spatially consistent climatological products, supporting applications in ecosystem monitoring, climate impact assessment, biodiversity studies, and large-scale environmental analysis across European bioclimatic regions.
Full article
(This article belongs to the Section Spatial Data Science for Environment and Earth)
►▼
Show Figures

Figure 1
Open AccessData Descriptor
Topoclimatic Graph Dataset for Frost Prediction in the Tropical High-Mountain Altiplano Cundiboyacense, Colombia
by
Evelin Calderón Caro, Dario Antonio Castañeda Sánchez, John R. Ballesteros and John W. Branch-Bedoya
Data 2026, 11(8), 205; https://doi.org/10.3390/data11080205 - 11 Aug 2026
Abstract
Frost prediction in tropical high-mountain agricultural regions is difficult because sparse meteorological networks must represent strong terrain-driven microclimatic variability. This article presents a topoclimatic graph dataset for frost prediction in the Altiplano Cundiboyacense, Colombia. The released core dataset contains 23 agricultural weather stations
[...] Read more.
Frost prediction in tropical high-mountain agricultural regions is difficult because sparse meteorological networks must represent strong terrain-driven microclimatic variability. This article presents a topoclimatic graph dataset for frost prediction in the Altiplano Cundiboyacense, Colombia. The released core dataset contains 23 agricultural weather stations and is seasonally focused on recurrent November–February frost periods rather than year-round continuous monitoring. It includes four consistently available meteorological variables at 30 min resolution: air temperature, relative humidity, dew point temperature, and solar radiation. The data were consolidated from multiple operational sources, harmonized to a common temporal grid, subjected to physical and consistency-based quality control, and completed through temporal and spatial reconstruction with traceability labels. The final release also provides binary frost_event and frost_warning_6h labels, point-based topographic descriptors, 1 km buffer-based raster summaries, land-cover proportions, station-level static feature vectors, and graph products including edge lists and adjacency matrices. These data products support graph-based deep learning, multimodal spatiotemporal analysis, and frost early warning experiments in a tropical mountain agroecosystem. The dataset offers a reproducible framework for integrating heterogeneous environmental observations into graph-ready representations while preserving sufficient environmental context for benchmarking frost prediction methods in data-sparse regions.
Full article
(This article belongs to the Topic Applications of Artificial Intelligence Models and Spatiotemporal Data in Agriculture and the Ecological Environment)
►▼
Show Figures

Graphical abstract
Open AccessData Descriptor
Vibration Dataset for Crack Analysis and Detection in a Rotating Bladed System
by
Adolfo Salgado-Ancona, José Billerman Robles-Ocampo, Edwin Neptalí Hernández-Estrada, Andrés López-López, Juvenal Rodríguez-Resendíz and Perla Yazmín Sevilla-Camacho
Data 2026, 11(8), 204; https://doi.org/10.3390/data11080204 - 10 Aug 2026
Abstract
►▼
Show Figures
This work presents conditioned and normalized vibration signal datasets acquired from the spanwise axis of the three blades of a rotating bladed system operating at a constant rotational speed of 240 rpm. The conditioned dataset was obtained using piezoelectric accelerometers mounted at the
[...] Read more.
This work presents conditioned and normalized vibration signal datasets acquired from the spanwise axis of the three blades of a rotating bladed system operating at a constant rotational speed of 240 rpm. The conditioned dataset was obtained using piezoelectric accelerometers mounted at the blade roots. The accelerometer output signals were conditioned and recorded by a dedicated data acquisition system. The signals were acquired under both healthy and damaged operating conditions. Baseline vibration signals were first recorded with all three blades in a healthy condition. Subsequently, cracks were deliberately introduced at three different locations along the blade span—the root, middle, and tip zones. Each crack location was independently evaluated on each of the three blades, resulting in a comprehensive dataset that includes healthy operation and all combinations of blade–damage locations. The datasets enable analysis of the system’s vibratory response and of dynamic information propagation toward the blade root, depending on the crack zone. Their main contribution is to provide reliable experimental data for the development, validation, and benchmarking of vibration-based diagnostic and structural health monitoring techniques. Furthermore, the datasets serve as valuable resources for advancing early crack detection strategies and enhancing the reliability of rotating industrial equipment with blades, such as fans, compressors, turbines, and aerogenerators.
Full article

Figure 1
Open AccessData Descriptor
Global Metadata of the Influence of Cover Crops on Key Soil Hydraulic Properties
by
Sabin Shrestha, Puja Sapkota, Bharat Sharma Acharya, Jason de Koff, Bharat Pokharel and Resham Thapa
Data 2026, 11(8), 203; https://doi.org/10.3390/data11080203 - 7 Aug 2026
Abstract
We present a global metadata comprising results from studies investigating the effects of cover crops (CCs) on six key soil hydraulic properties, namely total porosity, infiltration rate, saturated hydraulic conductivity, water retention at field capacity and permanent wilting points, and available water holding
[...] Read more.
We present a global metadata comprising results from studies investigating the effects of cover crops (CCs) on six key soil hydraulic properties, namely total porosity, infiltration rate, saturated hydraulic conductivity, water retention at field capacity and permanent wilting points, and available water holding capacity. This data repository is the result of a global meta-analysis entitled “Cover Crop Performance and Functional Groups Regulate Improvements in Soil Hydrology: A Global Meta-analysis”. Globally, numerous studies have investigated the role of CCs on soil hydraulic properties, but the results have varied across sites and years. Hence, the objective of the meta-analysis was to synthesize the existing knowledge base to assess the overall effects of CCs on these soil hydraulic properties and evaluate how environmental and management factors moderate these overall CC responses. We searched for peer-reviewed research articles published through 5 October 2024 in the ISI Web of Science database, with reference checking following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines. A total of 146 relevant articles were identified from which data on CC responses were extracted. The metadata consists of 1007 pairwise observations comparing CC vs. no-CC controls across diverse geographic regions worldwide. Moreover, we collected associated metadata for each pairwise comparison that includes a broad set of bibliographic, geographic, soil, climate, and management variables. Categorical variables were grouped into pre-defined factor levels or classes. Missing soil and climate data were filled using publicly available data products. Our data repository can be a valuable resource for the field and modeling community to identify knowledge gaps and guide future research.
Full article
(This article belongs to the Section Spatial Data Science for Environment and Earth)
►▼
Show Figures

Figure 1
Open AccessData Descriptor
MUTra-CDMX: Multisource Urban Traffic Dataset for the Insurgentes Sur Corridor in Mexico City
by
Arturo Rodríguez-Roman, Alicia Martínez-Rebollar, Hugo Estrada Esquivel, Ernesto de la Cruz-Nicolás and Eddie Clemente
Data 2026, 11(8), 202; https://doi.org/10.3390/data11080202 - 6 Aug 2026
Abstract
The growing complexity of urban mobility requires datasets that integrate dynamic traffic observations with meteorological, geometric, and urban-context information. This study presents MUTra-CDMX, a multisource urban traffic dataset covering a 14.72 km section of the Insurgentes Sur corridor in Mexico City. Traffic data
[...] Read more.
The growing complexity of urban mobility requires datasets that integrate dynamic traffic observations with meteorological, geometric, and urban-context information. This study presents MUTra-CDMX, a multisource urban traffic dataset covering a 14.72 km section of the Insurgentes Sur corridor in Mexico City. Traffic data were obtained from TomTom at five-minute intervals for 20 consecutive road segments from 1 November 2024 to 28 February 2025. Hourly meteorological data were retrieved from Meteosource, while segment-level geometry, topology, signalized locations, and nearby points of interest were derived from TomTom metadata and OpenStreetMap. The primary analytical file contains 691,200 segment–timestamp records and 12 variables describing traffic and free-flow conditions, meteorological information, derived operational indicators, and reconstruction status. Of these records, 682,264 are original observations and 8936 are reconstructed segment–timestamp combinations, identified by the Boolean variable is_imputed. Technical validation confirmed complete temporal coverage, preservation of original traffic observations, consistent weather alignment, and reconstruction performance through artificial masking. Predictive utility was evaluated through chronological travel-time forecasting under a leakage-controlled protocol. At the 30 min horizon, XGBoost achieved a mean absolute error of 12.84 s, a root mean squared error of 37.91 s, and a coefficient of determination ( ) of 0.771, outperforming a persistence baseline. MUTra-CDMX supports congestion analysis, imputation studies, spatiotemporal modeling, and travel-time forecasting.
Full article
(This article belongs to the Section Spatial Data Science for Environment and Earth)
►▼
Show Figures

Figure 1
Open AccessArticle
Mapping Data-Driven Governance in Sharing Economy Platforms: Algorithmic Management, Platform Control, and Value-Creation Mechanisms
by
Maria-Francisca Blasco-Lopez, Ramón Alberto Carrasco and Sulaiman Krayem
Data 2026, 11(8), 201; https://doi.org/10.3390/data11080201 - 6 Aug 2026
Abstract
Research on sharing economy platforms has expanded rapidly, yet the literature remains fragmented across studies on platform business models, gig work, algorithmic management, trust, reputation systems, artificial intelligence, and data-driven value creation. This article addresses this fragmentation through a bibliometric and systematic review
[...] Read more.
Research on sharing economy platforms has expanded rapidly, yet the literature remains fragmented across studies on platform business models, gig work, algorithmic management, trust, reputation systems, artificial intelligence, and data-driven value creation. This article addresses this fragmentation through a bibliometric and systematic review of 660 documents retrieved from Scopus and Web of Science covering the period from 2010 to May 2026. A PRISMA-based protocol guided identification, deduplication, screening, eligibility assessment, and final corpus construction. The analysis combined performance indicators, co-citation analysis, keyword co-occurrence mapping, country collaboration analysis, longitudinal thematic evolution, strategic diagrams, and systematic content coding using Bibliometrix/Biblioshiny 5.4.1, VOSviewer 1.6.21, and SciMAT 1.1.04. The results show a marked acceleration of the field after 2020 and identify major research clusters around algorithmic labour and platform control, algorithmic management, trust and reputation, and dynamic pricing. The systematic coding further indicates that algorithmic management, reputation systems, dynamic pricing, surveillance, matching, and AI-enabled mechanisms recur across governance and value-creation processes. The study develops an integrative framework that interprets these patterns through four connected elements: data inputs, algorithmic mechanisms, governance functions, and value outcomes. This framework provides managers and regulators with a basis for assessing transparency, accountability, participant autonomy, value distribution, and the legitimacy of platform governance.
Full article
(This article belongs to the Section Information Systems and Data Management)
►▼
Show Figures

Figure 1
Open AccessArticle
Machine Learning in FinTech for Financial Fraud Data Detection
by
Sanjaikanth E. Vadakkethil Somanathan Pillai and Wen-Chen Hu
Data 2026, 11(8), 200; https://doi.org/10.3390/data11080200 - 6 Aug 2026
Abstract
Financial fraud keeps rising these days. Organizations attempt to stop this trend by using various methods, such as distributing guides on how to avoid scams and frauds and automatically generating alerts when suspicious activities occur. However, this passive approach does not mitigate the
[...] Read more.
Financial fraud keeps rising these days. Organizations attempt to stop this trend by using various methods, such as distributing guides on how to avoid scams and frauds and automatically generating alerts when suspicious activities occur. However, this passive approach does not mitigate the problem, as the trend is worsening, and it is usually too late when victims realize they have been scammed. Therefore, active approaches must be employed before scams reach the victims. A wide variety of preventive methods, such as neural networks and data mining, have been used to detect financial fraud data, but none have proven entirely effective in combating scams. Each method has its pros and cons. This research takes advantage of multiple machine learning techniques, such as k-nearest neighbors (kNN) and decision trees, by utilizing data fusion to detect financial fraud accurately. The data fusion function used here is self-adjusting through learning. During the training phase, the system is repeatedly applied to the dataset until an optimal detection rate is achieved. Experimental results from credit card transactions show that the proposed method outperforms each individual method. Parameter or threshold values for the data fusion are set heuristically. Future research will focus on developing reconfigurable data fusion by automatically adjusting the values.
Full article
(This article belongs to the Special Issue Artificial Intelligence and Data Science for Fintech)
►▼
Show Figures

Figure 1
Open AccessArticle
A Data-Centric Network Traffic Dataset for Anomaly Detection: Construction, Reproducible Pipeline, and Technical Validation
by
Daniel Quirumbay Yagual, Diego Fernández Iglesias, Francisco J. Nóvoa and Daniel Garabato
Data 2026, 11(8), 199; https://doi.org/10.3390/data11080199 - 6 Aug 2026
Abstract
The effectiveness of machine learning and deep learning methods for network anomaly detection depends strongly on the quality and representativeness of the datasets used for training and evaluation. Despite recent advances, many publicly available benchmarks rely on synthetic traffic, outdated attack scenarios, or
[...] Read more.
The effectiveness of machine learning and deep learning methods for network anomaly detection depends strongly on the quality and representativeness of the datasets used for training and evaluation. Despite recent advances, many publicly available benchmarks rely on synthetic traffic, outdated attack scenarios, or limited representation of encrypted communications. This work presents a network traffic dataset derived from operational firewall logs collected in a heterogeneous institutional environment dominated by HTTPS/TLS traffic. A structured data-centric pipeline was implemented, including preprocessing, behavioral feature engineering, unsupervised pseudo-labeling through the EFMS–KMeans algorithm, class balancing using SMOTE, and the generation of model-oriented sequential representations for deep learning analysis. The resulting dataset contains large-scale flow-level records describing volumetric, behavioral, and temporal traffic characteristics while preserving privacy through anonymization procedures. Technical validation was conducted using statistical analysis, entropy-based measurements, clustering quality metrics, and dimensionality reduction techniques, confirming data consistency, structural diversity, and class separability. The dataset is publicly available through the Mendeley Data repository together with metadata and documentation supporting anomaly detection research, encrypted traffic analysis, and the evaluation of machine learning and deep learning approaches in realistic cybersecurity environments.
Full article
(This article belongs to the Topic Data Stream Mining and Processing)
►▼
Show Figures

Graphical abstract
Open AccessArticle
Dengue IgM ELISA Dataset from Trinidad and Tobago: Analytical Evaluation of Borderline Serological Results and Seroprevalence Estimation
by
Angel Justiz-Vaillant, Rodolfo Arozarena Fundora and Sachin Soodeen
Data 2026, 11(8), 198; https://doi.org/10.3390/data11080198 - 6 Aug 2026
Abstract
Background: Borderline dengue immunoglobulin M (IgM) enzyme-linked immunosorbent assay (ELISA) results create diagnostic and epidemiological uncertainty. Collapsing the manufacturer’s three qualitative categories into a binary outcome changes the effective decision rule and may materially alter the apparent positivity proportion. This Data article describes
[...] Read more.
Background: Borderline dengue immunoglobulin M (IgM) enzyme-linked immunosorbent assay (ELISA) results create diagnostic and epidemiological uncertainty. Collapsing the manufacturer’s three qualitative categories into a binary outcome changes the effective decision rule and may materially alter the apparent positivity proportion. This Data article describes an anonymized laboratory dataset from Trinidad and Tobago and evaluates both category-handling uncertainty and test-misclassification uncertainty. Methods: The dataset comprises 161 consecutive serum specimens submitted for routine dengue IgM testing between 1 September 2025 and 28 February 2026 and processed in four analytical batches. The results were analyzed under three prespecified scenarios: borderline classified as negative, borderline excluded, and borderline classified as positive. Exact Clopper–Pearson 95% confidence intervals (CIs) were calculated. Rogan–Gladen adjustment was applied only to the borderline-excluded scenario because the ELISA test validation estimates of sensitivity and specificity were calculated after excluding borderline results. Results: Twenty specimens (12.4%) were positive, 29 (18.0%) borderline, and 112 (69.6%) negative. The apparent IgM positivity was 12.4% (20/161; 95% CI 7.8–18.5%) when borderline results were classified as negative, 15.2% (20/132; 95% CI 9.5–22.4%) when they were excluded, and 30.4% (49/161; 95% CI 23.4–38.2%) when they were classified as positive. For determinate results, the conditional Rogan–Gladen estimates were 13.2% using a sensitivity of 100.0% and specificity of 97.7%, and 15.1% using a sensitivity of 82.2% and specificity of 96.8%. Conclusions: Misclassification adjustment is informative but remains conditional on the transportability of external assay-performance estimates. The dataset supports transparent category-level reanalysis, while the absence of continuous index values, confirmatory testing, and detailed clinical metadata limits numerical cut-off recalibration and patient-level inference.
Full article
(This article belongs to the Section Computational Biology, Bioinformatics, and Biomedical Data Science)
►▼
Show Figures

Graphical abstract
Open AccessData Descriptor
Draft Genome Sequence Data of Multidrug-Resistant Escherichia coli CUK-76 Co-Harboring Class A and Class C β-Lactamases from Wastewater of India
by
Achhada Ujalkaur Avatsingh, Shilpa Sharma, Shilippreet Kour, Anvesha Bhardwaj, Prem Prashant Chaudhary and Nasib Singh
Data 2026, 11(8), 197; https://doi.org/10.3390/data11080197 - 6 Aug 2026
Abstract
The present study was performed to determine the antibiotic resistance genes (ARGs), virulence determinants, and mobile genetic elements in multidrug-resistant Escherichia coli CUK-76 isolated from wastewater in Himachal Pradesh, India. Whole genome sequencing was performed using the Illumina Miseq system, and the draft
[...] Read more.
The present study was performed to determine the antibiotic resistance genes (ARGs), virulence determinants, and mobile genetic elements in multidrug-resistant Escherichia coli CUK-76 isolated from wastewater in Himachal Pradesh, India. Whole genome sequencing was performed using the Illumina Miseq system, and the draft genome sequence was assembled by Unicycler v0.5.1 and annotated by the NCBI Prokaryotic Genome Annotation Pipeline (PGAP v6.10). The bioinformatics-based prediction analysis was performed using ResFinder v4.7.2 and CARD v4.0.1 (antibiotic resistance genes), VirulenceFinder v2.0, VFDB and MGEFinder v1.0.3 (virulence determinants), PlasmidFinder v2.0.1 (plasmid sequences), MLST v2.0 (sequence type), ISFinder and TnCentral v2.0 (insertion sequences and transposons), PathogenFinder2 v0.6.0 (pathogenicity), RAST (subsystems category) and Phigaro (prophage sequences). The draft genome of E. coli CUK-76 strain comprised 4,607,136 bp with a GC content of 51%. Genome annotation revealed 4498 genes of which 4289 were protein-coding genes, 78 RNA genes, and 131 pseudogenes. It was related to sequence type ST949 and its predicted resistome consisted of blaCTX-M-15, blaTEM-1B (class A β-lactamase genes), blaEC-14 (class C β-lactamase gene), aph(6)-Id, aph(3″)-Ib, qnrS1, sul2, tet(A), and dfrA14 genes. Additionally, multiple virulence genes, two plasmid sequences viz. IncFIB(K) and IncFIB(AP001918), insertion sequences, transposons and prophage sequences were detected. The genomic dataset of this strain will be a valuable resource for comparative genomic studies on E. coli.
Full article
(This article belongs to the Special Issue Benchmarking Datasets in Bioinformatics, 3rd Edition)
►▼
Show Figures

Figure 1
Open AccessData Descriptor
The ArchiveGene Corpus: A Synthetic Multi-Layer Benchmark for Genealogical Information Extraction from Uzbek Historical Archival Texts
by
Adilbek Dauletov, Noila Matyakubova, Sevara Allabergenova, Nargisa Ashirmatova, Miyassar Tillayeva, Sevara Yoqubova and Ikrom Islomov
Data 2026, 11(8), 196; https://doi.org/10.3390/data11080196 - 5 Aug 2026
Abstract
►▼
Show Figures
Automatic extraction of genealogical information from historical archival-genealogical documents in Uzbek is an understudied problem for low-resource languages. Multi-layer NLP benchmarks are not sufficient to automatically identify individuals, family relationships, dates, place names, and archival identifiers in such texts. Also, the same people
[...] Read more.
Automatic extraction of genealogical information from historical archival-genealogical documents in Uzbek is an understudied problem for low-resource languages. Multi-layer NLP benchmarks are not sufficient to automatically identify individuals, family relationships, dates, place names, and archival identifiers in such texts. Also, the same people are mentioned in various forms: full name, pronoun (18.8%), initial, surname-name order, indirect expression (9.4%), and title. Existing NER and relation extraction corpora are mainly focused on high-resource languages or general domain texts and do not sufficiently cover the FAMILY_ROLE signals, historical spelling variants, and fond–opis–delos identifiers specific to Uzbek archival-genealogical texts. Proposed resource: We present the ArchiveGene Corpus, a controlled, fully synthetic, and reproducible five-layer resource consisting of 1000 Uzbek archival-genealogical-style documents, divided into 700 training, 150 validation, and 150 test documents. The corpus contains 8366 named entities, 10,625 person mentions, 2000 coreference chains, and 1000 genealogical relation triples. The dataset was generated using a deterministic template-based pipeline and a lexicon of Uzbek names, and is fully reproducible. Inter-annotator agreement values were 0.847 for NER, 0.793 for coreference, and 0.821 for RE, according to Cohen’s κ. Comparative results are presented with four baseline models (rule-based, BiLSTM-CRF, mBERT, and XLM-RoBERTa). The dataset is openly hosted on the Zenodo platform under the CC BY 4.0 license; concept DOI: 10.5281/zenodo.20670360, v1.1.1 version DOI: 10.5281/zenodo. 21429998. Scientific significance: To the best of our knowledge, ArchiveGene is among the first openly released, controlled synthetic resources for Uzbek that integrates named-entity recognition, person-mention detection, coreference resolution, genealogical relation extraction, and final tuple generation within a single annotation framework. The baseline analysis provides three main conclusions: (1) on the clean synthetic test set, the transformer models already reach 100.00 Micro-F1 for NER and 100.00 Macro-F1 for coreference-aware relation extraction, so coreference aggregation adds little on synthetic data (+2.25 for mBERT and +0.04 for XLM-RoBERTa) but its contribution is expected to grow on real archival text; (2) the rule-based and heuristic baselines lag far behind (Macro-F1 50.28 and 70.73) and fail entirely on spouse_of, showing the limits of lexical rules; and (3) a zero-shot evaluation on a real-document pilot reduces NER Micro-F1 from 100.00 to 22.17, indicating that the synthetic corpus is trivially learnable and that real-archival validation is essential.
Full article

Figure 1
Open AccessData Descriptor
A Multi-Class SDN Intrusion Detection Dataset with Synchronized OpenFlow Control-Plane Telemetry
by
Juliana Arévalo-Herrera, Jorge E. Camargo, José Ignacio Martínez Torre, Juan Marcos Ramírez and Tatiana Zona-Ortiz
Data 2026, 11(8), 195; https://doi.org/10.3390/data11080195 - 5 Aug 2026
Abstract
Software-Defined Networking (SDN) separates the control and data planes, introducing a logically centralized controller that is itself a high-value attack target. Despite growing interest in SDN intrusion detection, publicly available datasets either restrict evaluation to binary normal-vs-DDoS classification or lack control-plane telemetry, leaving
[...] Read more.
Software-Defined Networking (SDN) separates the control and data planes, introducing a logically centralized controller that is itself a high-value attack target. Despite growing interest in SDN intrusion detection, publicly available datasets either restrict evaluation to binary normal-vs-DDoS classification or lack control-plane telemetry, leaving multi-class detection of SDN-architectural attacks without a dedicated benchmark. This work presents LAN-SDN-NIDS, a publicly available, multi-class flow-level dataset of 1,125,059 records generated in a fully containerized Containernet/OpenDaylight testbed across five standard network topologies. Each flow record combines 29 traffic-level features with 11 control-plane-aware metrics—including Packet-In and Flow-Mod counts and first-seen delay. The dataset covers five attack classes in two categories: three that exploit SDN control-plane mechanisms (link fabrication, host injection, and port hijack) alongside DDoS and port scan, plus normal traffic. An XGBoost classifier trained on the full feature set achieved a macro F1 of 0.94; an ablation study showed that removing OpenFlow features causes link fabrication F1 to collapse from 0.97 to 0.19, indicating that control-plane telemetry is decisive for detecting SDN-architectural attacks under the conditions evaluated. A UMAP embedding is consistent with class separability, except for a structural overlap between host injection and normal traffic attributable to their shared ARP protocol.
Full article
(This article belongs to the Section Information Systems and Data Management)
►▼
Show Figures

Figure 1
Open AccessArticle
An Update on the Top 100 Most-Cited Articles in Diabetes Research: A Bibliometric Analysis
by
Reza Fahimi, Abdulaziz Alrubayyi, Thomas M. Barber and Olalekan A. Uthman
Data 2026, 11(8), 194; https://doi.org/10.3390/data11080194 - 5 Aug 2026
Abstract
►▼
Show Figures
Background: Diabetes mellitus remains one of the most extensively researched conditions in global health, with a rapidly growing body of literature addressing its prevention, management, and treatment. Characterising highly cited diabetes-related research helps identify research trends, shifts in scientific priorities, and factors associated
[...] Read more.
Background: Diabetes mellitus remains one of the most extensively researched conditions in global health, with a rapidly growing body of literature addressing its prevention, management, and treatment. Characterising highly cited diabetes-related research helps identify research trends, shifts in scientific priorities, and factors associated with citation impact in diabetes research. Objective: This study aimed to perform a bibliometric analysis of the top 100 most-cited diabetes-related records published between 2011 and 2022, in order to describe major scientific contributions and citation patterns in this field. Methods: The search was conducted in April 2023 using the Web of Science Core Collection, Science Citation Index Expanded. Records containing the term “diabetes” were retrieved using the topic field and ranked according to total Web of Science citation count. The top 300 records were screened, and the 100 eligible diabetes-related records were selected for detailed analysis. Extracted data included title, citation count, citation density, PubMed Central (PMC) citations, patent citations, h-index, journal, year of publication, country of origin, study type, funding sources, and authorship patterns. Descriptive and correlational analyses were performed using IBM SPSS Statistics, version 27.0.1.0. Results: The 100 most-cited records accumulated 174,857 citations, with individual citation counts ranging from 832 to 5898. Most records were original research articles (68%), followed by reviews (27%). The most frequently represented journals were Diabetes Care (n = 21), New England Journal of Medicine (n = 20), and The Lancet (n = 9). The United States contributed 45% of the records. A strong positive correlation was observed between total Web of Science citations and PMC citations using both Pearson correlation (r = 0.863) and Spearman rank correlation (ρ = 0.834). The association between total citations and first-author h-index was weak and not statistically significant. Patent citation analysis showed that 41 records were cited in at least one patent. Most records acknowledged funding, with both public-sector and pharmaceutical funders represented among the most frequently reported sources. Conclusions: This study characterises highly cited diabetes-related research published from 2011 to 2022. The findings suggest that clinical trials, reviews, guidelines, collaborative research, patent citation activity, and publication in high-visibility journals were common features of highly cited diabetes-related records, although citation-based indicators should be interpreted with caution.
Full article

Graphical abstract
Open AccessReview
Decoding Dental Insurance Claims Data for Oral Health Research and Policy: Challenges, Validation, and Opportunities in Biomedical Informatics—A Scoping Review
by
Deepti Virupakshappa, Rajashekhara Bhari Sharanesha, Alwaleed Abushanan, Sara Alghamdi, Maram Alagla and Faisal Alotaibi
Data 2026, 11(8), 193; https://doi.org/10.3390/data11080193 - 4 Aug 2026
Abstract
Background: Dental insurance claims data are vital for research in oral health, epidemiology, and policy. However, issues like data quality, coding standards, validity, interoperability, and analytical approaches hinder their use. This review outlines these challenges. Methods: Following PRISMA-ScR and Arksey-O’Malley, we searched Web
[...] Read more.
Background: Dental insurance claims data are vital for research in oral health, epidemiology, and policy. However, issues like data quality, coding standards, validity, interoperability, and analytical approaches hinder their use. This review outlines these challenges. Methods: Following PRISMA-ScR and Arksey-O’Malley, we searched Web of Science, Scopus, and PubMed through April 2026 for peer-reviewed studies on dental insurance data issues. Two reviewers screened and extracted data, identifying key challenges and implications. Results: Out of 563 records, 389 remained after deduplication; 45 studies met criteria. Data sources included Medicaid, Medicare, insurers, and national systems from various countries. Six main challenges emerged: (1) coding errors and lack of standardization; (2) data validity and quality concerns; (3) interoperability and linkage barriers; (4) fraud detection issues; (5) analytical limitations; (6) policy insights on disparities. Validation showed variable accuracy, with diagnosis codes more reliable than procedure codes. Conclusions: Challenges limit data use in research and policy. Standardized coding, validation, interoperability, transparency, and causal inference are essential for leveraging these data to improve oral health research and policies.
Full article
(This article belongs to the Section Computational Biology, Bioinformatics, and Biomedical Data Science)
►▼
Show Figures

Figure 1
Open AccessCommunication
Omics Strategies and Big Data: Transforming Precision Medicine and Systems Biology
by
Ana Checa-Ros, Owahabanun-Joshua Okojie and Luis D’Marco
Data 2026, 11(8), 192; https://doi.org/10.3390/data11080192 - 1 Aug 2026
Abstract
The convergence of multi-omics strategies and big data analytics is transforming modern healthcare by shifting medicine from a reactive discipline to a proactive, personalized system. This review explores how integrating genomics, transcriptomics, proteomics, and metabolomics provides a holistic view of complex biological systems.
[...] Read more.
The convergence of multi-omics strategies and big data analytics is transforming modern healthcare by shifting medicine from a reactive discipline to a proactive, personalized system. This review explores how integrating genomics, transcriptomics, proteomics, and metabolomics provides a holistic view of complex biological systems. While current literature heavily documents theoretical models, a persistent gap remains in translating these high-dimensional architectures into validated healthcare workflows. We highlight the critical role of big data infrastructure, specifically cloud computing and artificial intelligence (AI), in processing massive datasets and we benchmark our approach against existing reviews to emphasize the path toward routine clinical deployment. Through concrete case studies in oncology and metabolic disorders, we illustrate the potential clinical utility of these technologies in advancing precision medicine. Finally, we address persistent challenges including data heterogeneity, privacy concerns, and computational bottlenecks, and discuss future directions required to translate multi-omics insights into routine clinical practice.
Full article
(This article belongs to the Section Computational Biology, Bioinformatics, and Biomedical Data Science)
►▼
Show Figures

Graphical abstract
Open AccessData Descriptor
Dataset on Agrometeorological Parameters in the Souss-Massa Plain
by
Hamza Ait-Ichou, Mohammed Hssaisoune, Abdelwahed Chaaou, Mohammed El Hafyani, Asma Abou Ali, Adnane Chakir, Yassine Ait-Brahim, Khaoula Bakas, Amine Saddik, Ilham Elhaid, Soufiane Taia, Said El Hachemy, Aya Rais, Adnane Labbaci, Salwa Belaqziz, Abdellaali Tairi, Safae Ijlil, Houria Abahous, Elhousna Faouzi, Ismail Ait Lahssaine, Rachid El Moumen, Moussa Ait El Kadi, Fatima Abdelfadel, Sofyan Sbahi, Sokaina Tadoumant, Brahim Meskour, Soumia Gouahi, Chaima Aglagal, Hamza Ait Moh, Hassan Mosaid and Lhoussaine Bouchaouadd
Show full author list
remove
Hide full author list
Data 2026, 11(8), 191; https://doi.org/10.3390/data11080191 - 1 Aug 2026
Abstract
The Eddy Covariance station provides observations of agrometeorological variables and surface energy fluxes, collected from 2019 to 2022, in a citrus orchard located in the Souss-Massa plain, Morocco. The present dataset comprises measurements recorded via a set of aboveground and subsurface sensors. The
[...] Read more.
The Eddy Covariance station provides observations of agrometeorological variables and surface energy fluxes, collected from 2019 to 2022, in a citrus orchard located in the Souss-Massa plain, Morocco. The present dataset comprises measurements recorded via a set of aboveground and subsurface sensors. The aboveground setup consistently measures air temperature, relative humidity, wind speed, net radiation, and precipitation. Additionally, the subsurface setup continuously tracks soil temperature, moisture, and electrical conductivity at depths from 5 to 80 cm, along with soil heat flux. Moreover, these setups enable the measurement of turbulent fluxes (sensible and latent heat). Given the limited availability of long-term agrometeorological data in semi-arid regions of the Mediterranean, this paper addresses a critical data gap by providing a reliable agrometeorological dataset. The latter consists of two types of data: 30 min interval files and high-frequency files (20 Hz, i.e., one measurement every 50 ms). The processing of this data involved Card Convert, MATLAB EC-Pack, and Excel, with data quality control performed by removing outliers and excluding nighttime fluxes. The dataset is organized in a table and provided in a .csv format with standard metadata. It is designed for a wide range of applications, including evapotranspiration modeling, satellite product validation, agroclimatic monitoring, determining crop irrigation requirements, precision irrigation planning, and water management. Additionally, the dataset can be reused for crop and hydrological model calibration, as well as soil moisture and crop stress prediction using machine learning algorithms.
Full article
(This article belongs to the Section Spatial Data Science for Environment and Earth)
►▼
Show Figures

Figure 1
Open AccessArticle
Performance of Professional Soccer Players: Multilevel Classification Using Machine Learning and SHAP Explainability by Playing Position
by
Boryi A. Becerra-Patiño, Rodrigo Villaseca-Vicuña, Diego Andrés Rada-Perdigón, Juan David Paucar-Uribe, Wilder Geovanny Valencia-Sánchez, José Francisco López-Gil and Rodrigo Yáñez-Sepúlveda
Data 2026, 11(8), 190; https://doi.org/10.3390/data11080190 - 30 Jul 2026
Abstract
Background. Recent advances in data systematization have enabled the development of machine learning models to evaluate performance in elite sports; however, studies are needed to analyze the performance of professional players in relation to their playing position. Objective. To analyze the
[...] Read more.
Background. Recent advances in data systematization have enabled the development of machine learning models to evaluate performance in elite sports; however, studies are needed to analyze the performance of professional players in relation to their playing position. Objective. To analyze the performance of professional soccer players who competed between 2017 and 2024 by applying a multilevel classification approach that integrates different machine learning algorithms. Materials and Methods. We analyzed 9088 player-seasons from professional players during the 2017–2024 seasons. These data were extracted from standardized databases on sports performance analysis belonging to the following leagues: LaLiga (Spain), Premier League (England), Bundesliga (Germany), Serie A (Italy), and Ligue 1 (France). The final sample, distributed by performance level, was as follows: elite players (n = 1818; 20%), mid-level players (n = 2727; 30%), and low-level players (n = 4543; 50%). The average age of the players analyzed was 26.1 ± 4.0 years, distributed across three outfield playing positions: center backs (CB), central midfielders (CM), and strikers (ST); goalkeepers were excluded because the dataset contains no goalkeeper-specific performance metrics. Results. Under a leakage-controlled protocol (the label-defining indicators were excluded from the predictors and the train/test split preceded all preprocessing), a linear model (logistic regression) achieved the best overall performance (mean macro-F1 = 0.729), ahead of ensemble and kernel-based methods; this indicates that, once target leakage is removed, the classification does not require non-linear models. Strikers (ST) were the most separable position (best macro-F1 = 0.818, area under the receiver operating characteristic curve [AUC-ROC] = 0.948) and center backs (CB) the most difficult (macro-F1 = 0.602, AUC-ROC = 0.791), with midfielders (CM) intermediate (macro-F1 = 0.770, AUC-ROC = 0.916). Conclusions. The results confirmed that performance structures differ substantially depending on the position in the field, supporting the use of position-specific analytical strategies. In this context, the combination of position-stratified dimensionality reduction, handling of imbalance, and explainable artificial intelligence allowed for the identification of interpretable performance patterns associated with the profiles of elite, mid-level, and low-level players.
Full article
(This article belongs to the Special Issue Big Data and Data-Driven Research in Sports)
►▼
Show Figures

Figure 1
Open AccessData Descriptor
PR-IHC-40X: Progesterone Receptor Immunohistochemistry Dataset for Breast Cancer Diagnosis
by
Hasanul Bannah, Md Serajun Nabi, Mohammad Faizal Ahmad Fauzi, Sarina Mansor, Wan Siti Halimatul Munirah Wan Ahmad, Aysha Akter Shahazadi, Seow-Fan Chiew, Phaik-Leng Cheah and Lai-Meng Looi
Data 2026, 11(8), 189; https://doi.org/10.3390/data11080189 - 28 Jul 2026
Abstract
The PR-IHC-40X dataset comprises a high-resolution collection of region-of-interest (ROI) images and corresponding ground-truth (GT) annotations for progesterone receptor (PR) immunohistochemistry (IHC) analysis in breast cancer pathology. We obtained 50 glass slides from the University of Malaya Medical Centre (UMMC) and digitized them
[...] Read more.
The PR-IHC-40X dataset comprises a high-resolution collection of region-of-interest (ROI) images and corresponding ground-truth (GT) annotations for progesterone receptor (PR) immunohistochemistry (IHC) analysis in breast cancer pathology. We obtained 50 glass slides from the University of Malaya Medical Centre (UMMC) and digitized them into whole-slide images (WSIs) at 40× magnification using a 3DHistech Pannoramic DESK scanner. Pathologists annotated ROIs on the collaborative Cytomine platform, which formed the basis of dataset extraction. Ground-truth masks were generated in a multi-stage process: binary nuclei masks for segmentation were first created with a StarDist deep learning model and refined by manual correction, while the classification ground truth was first determined using a CNN-based approach and then modified by diaminobenzidine (DAB) intensity thresholding into four expression classes: Strong (red), Moderate (yellow), Weak (green), and Negative (blue). The classification outputs were re-corrected in a loop against the pathologists’ feedback and the manually checked results. There were approximately 32,000 nuclei within 250 ROI images that were manually checked and validated by senior pathologists individually. Each ROI comes with its binary segmentation mask and four-class color annotations, which make it a reliable dataset for deep learning research on nuclei segmentation, PR expression classification, and Allred scoring. To ensure a fair and reproducible evaluation, the dataset is released with a predefined slide-level (patient-wise) partition into training, testing, and evaluation subsets so that no slide contributes regions of interest to more than one subset and data leakage across subsets is avoided.
Full article
(This article belongs to the Section Computational Biology, Bioinformatics, and Biomedical Data Science)
►▼
Show Figures

Figure 1
Open AccessData Descriptor
A Genre Classification Scheme for Metal-Music Corpus Studies: An Eleven-Bucket and Seventeen-Category Encoding of Encyclopaedia Metallum Genre Strings
by
Ignacio Soto-Silva
Data 2026, 11(8), 188; https://doi.org/10.3390/data11080188 - 28 Jul 2026
Abstract
This data descriptor documents a two-level genre classification scheme for metal-music corpus studies derived from Encyclopaedia Metallum’s multi-label genre strings. The scheme assigns each band to one of eleven mutually exclusive primary buckets and, alternatively, to one of seventeen finer-grained sub-categories that separate
[...] Read more.
This data descriptor documents a two-level genre classification scheme for metal-music corpus studies derived from Encyclopaedia Metallum’s multi-label genre strings. The scheme assigns each band to one of eleven mutually exclusive primary buckets and, alternatively, to one of seventeen finer-grained sub-categories that separate closely related sub-styles. Classification is implemented by keyword priority on the lower-cased genre string and is fully deterministic given the published keyword table. The scheme was developed and tested on a corpus of 560 metal bands from Southern Chile (1988–2024), which serves as the development corpus throughout. On a stratified 60-band subset of this corpus, automatic assignments agreed with independent expert human coding at 81.7% (Cohen’s κ = 0.78, 95% bootstrap CI [0.66, 0.88]); an out-of-sample application to a 60-band Norwegian sample, with the keyword tables left unchanged, retained the scheme’s core logic at a 5.0% residual rate. The encoding rules, the keyword tables, and the mapping CSVs are released under CC-BY 4.0 to enable replication and adaptation across other metal-scene corpora.
Full article
(This article belongs to the Section Information Systems and Data Management)
Journal Menu
► ▼ Journal Menu-
- Data Home
- Aims & Scope
- Editorial Board
- Reviewer Board
- Topical Advisory Panel
- Instructions for Authors
- Guidelines for Reviewers
- Special Issues
- Topics
- Sections & Collections
- Article Processing Charge
- Indexing & Archiving
- Editor’s Choice Articles
- Most Cited & Viewed
- Journal Statistics
- Journal History
- Journal Awards
- Conferences
- Editorial Office
- 10th Anniversary
Journal Browser
► ▼ Journal BrowserHighly Accessed Articles
Latest Books
E-Mail Alert
News
Topics
Topic in
AI, Algorithms, BDCC, Computers, Data, Future Internet, Informatics, Information, MAKE, Publications, Smart Cities
Learning to Live with Gen-AI
Topic Editors: Antony Bryant, Paolo Bellavista, Kenji Suzuki, Horacio Saggion, Roberto Montemanni, Andreas Holzinger, Min ChenDeadline: 31 August 2026
Topic in
Geosciences, IJGI, Remote Sensing, Sensors, Data
Advances in Sensor Data Fusion and AI for Environmental Monitoring
Topic Editors: Zhenyu Yu, Mohd Yamani Idna Idris, Yu Li, Aleksandar Dj ValjarevićDeadline: 30 September 2026
Topic in
Aerospace, Applied Sciences, Data, Remote Sensing, Sensors, Universe
Techniques and Science Exploitations for Earth Observation and Planetary Exploration-2nd Edition
Topic Editors: Yu Tao, Siting Xiong, Rui SongDeadline: 30 November 2026
Topic in
Sensors, Energies, Applied Sciences, Electronics, Technologies, Data, Modelling, Mathematics
Industrial Big Data and Artificial Intelligence
Topic Editors: Chun Yin, Jiusi Zhang, Quan Qian, Tenglong HuangDeadline: 20 December 2026
Conferences
Special Issues
Special Issue in
Data
IoT and Big Data Applications in Smart Cities: Recent Advances, Challenges, and Critical Issues
Guest Editors: Diego Gabriel Rossit, Pedro Moreno-BernalDeadline: 31 August 2026
Special Issue in
Data
Natural Language Processing in the Era of Big Data
Guest Editors: Mingyu Wan, Chu Ren HuangDeadline: 30 September 2026
Special Issue in
Data
Mining and Computational Intelligence for E-Learning and Education—4th Edition
Guest Editor: Antonio Sarasa CabezueloDeadline: 30 September 2026
Special Issue in
Data
Artificial Intelligence and Data Science for Fintech
Guest Editor: Kani ChenDeadline: 31 October 2026
Topical Collections
Topical Collection in
Data
Modern Geophysical and Climate Data Analysis: Tools and Methods
Collection Editors: Vladimir Sreckovic, Zoran Mijic



