Next Issue
Volume 11, June
Previous Issue
Volume 11, April
 
 

Data, Volume 11, Issue 5 (May 2026) – 35 articles

Cover Story (view full-size image): This dataset presents empirical research conducted at the Archaeological Site of Ancient Dodona in Greece, focusing on visitor experience, interpretive needs, and attitudes toward digital technologies in cultural heritage environments. Based on questionnaire responses collected in situ from adult visitors, the dataset documents visitor perceptions of existing interpretive material, spatial behaviour within the archaeological site, and interest in digital applications such as augmented reality, digital storytelling, and interactive heritage tools. The dataset contributes to research in digital heritage, visitor studies, cultural tourism, and visitor-centred interpretation, while supporting the future development of immersive and adaptive interpretation strategies for archaeological sites. View this paper
  • Issues are regarded as officially published after their release is announced to the table of contents alert mailing list.
  • You may sign up for e-mail alerts to receive table of contents of newly released issues.
  • PDF is the official format for papers published in both, html and pdf forms. To view the papers in pdf format, click on the "PDF Full-text" link, and use the free Adobe Reader to open them.
Order results
Result details
Section
Select all
Export citation of selected articles as:
13 pages, 1675 KB  
Data Descriptor
Expression of Genes Associated with Epithelial to Mesenchymal Transition in MCF-7 Breast Cancer Cells Treated with Monocarbonyl Analogs of Curcumin C66 and B2BrBC—RT-qPCR Array Dataset
by Radoslav Stojchevski, Sara Velichkovikj, Jane Bogdanov, Katerina Dragarska, Ivana Todorovska, Nikola Hadzi-Petrushev, Mitko Mladenov, Leonid Poretsky and Dimiter Avtanski
Data 2026, 11(5), 125; https://doi.org/10.3390/data11050125 - 21 May 2026
Viewed by 669
Abstract
Curcumin is a polyphenolic bio-compound derived from the rhizomes of the turmeric plant (Curcuma longa) that has proven anti-carcinogenic properties but poor bioavailability. By modifying its chemical structure, the monocarbonyl analogs of curcumin (MACs) possess improved stability, resorption, and circulation. This [...] Read more.
Curcumin is a polyphenolic bio-compound derived from the rhizomes of the turmeric plant (Curcuma longa) that has proven anti-carcinogenic properties but poor bioavailability. By modifying its chemical structure, the monocarbonyl analogs of curcumin (MACs) possess improved stability, resorption, and circulation. This dataset presents RT-qPCR array analysis of 84 genes associated with Epithelial to Mesenchymal Transition (EMT), a key early event in cancer progression and metastasis, in human MCF-7 breast cancer cells. Cells were stimulated toward EMT reprogramming by treatment with a combination of EMT-inducing factors and co-treated with two experimental MACs, C66 or B2BrBC. Gene expression was measured using the human EMT QIAGEN RT2 Profiler kit, and results were obtained from three independent experiments. Gene expression changes are presented as both fold regulation and fold change values, with statistical significance determined by Student’s t-test (p < 0.05). This comprehensive dataset enables investigation into how MACs modulate the EMT transcriptome in breast cancer cells, with potential applications for understanding EMT mechanisms. The raw and processed data are publicly available and can be used for comparative analyses, validation studies, and bioinformatic analyses of EMT-related signaling pathways. Full article
Show Figures

Graphical abstract

12 pages, 53356 KB  
Data Descriptor
High-Temporal-Resolution In Situ Tropical Meteorological Dataset from Santo Domingo, Dominican Republic for Energy and Environmental Applications
by Francisco A. Ramírez-Rivera and Néstor F. Guerrero-Rodríguez
Data 2026, 11(5), 124; https://doi.org/10.3390/data11050124 - 21 May 2026
Viewed by 686
Abstract
High-temporal-resolution meteorological data collected under tropical climatic conditions are scarce yet highly valuable, because they facilitate a more precise characterisation of atmospheric variability. This paper presents a high-temporal-resolution meteorological database collected in Santo Domingo, Dominican Republic, using a Davis Vantage Pro2 Plus weather [...] Read more.
High-temporal-resolution meteorological data collected under tropical climatic conditions are scarce yet highly valuable, because they facilitate a more precise characterisation of atmospheric variability. This paper presents a high-temporal-resolution meteorological database collected in Santo Domingo, Dominican Republic, using a Davis Vantage Pro2 Plus weather station. The database encompasses 170,861 observations covering 35 meteorological parameters, all of which were acquired at a temporal resolution of one minute. The database was prepared through a rigorous process of data organisation and cleaning. In the initial stage, NaN values representing missing data were identified and removed. In the second stage, based on the temporal and statistical characteristics of the database, missing values were reconstructed using an imputation technique. This approach ensures data quality and consistency, which are essential for the reliable utilisation of the database in scientific, energy-related, and engineering applications. The database is a valuable resource for the development of predictive tools based on machine learning and deep learning techniques to forecast solar radiation and wind speed. Additionally, it can serve as a useful input for the modelling and simulation of renewable energy systems. Full article
Show Figures

Figure 1

11 pages, 5044 KB  
Data Descriptor
VaxiGen Database of Tumor Immunogens
by Stanislav Sotirov, Ivan Dimitrov and Irini Doytchinova
Data 2026, 11(5), 123; https://doi.org/10.3390/data11050123 - 20 May 2026
Viewed by 809
Abstract
Peptide-based cancer vaccines have emerged as a prominent focus in contemporary oncological research, as the quest for innovative cancer treatment modalities continues to gain momentum. A pivotal facet of their development is the precise delineation and characterization of immunogenic tumor antigens. In this [...] Read more.
Peptide-based cancer vaccines have emerged as a prominent focus in contemporary oncological research, as the quest for innovative cancer treatment modalities continues to gain momentum. A pivotal facet of their development is the precise delineation and characterization of immunogenic tumor antigens. In this context, VaxiJen stands out as one of the most widely used and cited computational servers for predicting immunogenicity, making it an invaluable tool for in silico antigen prediction. However, the database underpinning VaxiJen’s predictions has not undergone a comprehensive update for over fifteen years. To address this, a systematic search of the PubMed database was conducted to identify scholarly articles reporting data on novel immunogenic proteins and peptides undergoing human testing. The corresponding sequences of these proteins and peptides were subsequently curated from UniProtKB. Therefore, in this study, we introduce an updated dataset encompassing a repertoire of tumor immunogens, comprising 546 full-length human proteins and 212 human tumor peptides, as well as tumor non-immunogens, comprising 548 full-length human proteins and 181 human tumor peptides. The recently compiled VaxiGen tumor dataset is openly accessible. Researchers can conveniently download, search, and process it. This dataset, when paired with a suitable negative dataset, can further serve as a valuable training set, thereby facilitating improved predictions of the potential immunogenicity of hitherto uncharacterized protein or peptide sequences. Full article
Show Figures

Figure 1

24 pages, 1283 KB  
Article
Evaluating the Integrity of LLM-Generated Citations: Prevalence and Risks of Fabricated References in Scientific Literature
by Pablo Picazo-Sanchez and Lara Ortiz-Martin
Data 2026, 11(5), 122; https://doi.org/10.3390/data11050122 - 20 May 2026
Cited by 1 | Viewed by 1288
Abstract
Large Language Models have become important in our lives, and academia is not agnostic to this trend, offering tools like text rephrasing and summarisation. However, this integration raises significant concerns regarding the integrity of science. In this paper, we investigate hallucinations of LLMs [...] Read more.
Large Language Models have become important in our lives, and academia is not agnostic to this trend, offering tools like text rephrasing and summarisation. However, this integration raises significant concerns regarding the integrity of science. In this paper, we investigate hallucinations of LLMs when generating scientific references. Using nine LLMs, we generated a dataset of 74,196 BIBTEX references to quantify and analyse fabricated references, focusing on distinguishing between intrinsic and extrinsic hallucinations. Also, we extracted and analysed 127,063 references from 3541 published papers in 2023 to assess the prevalence of fake bibliographic data. Our manual verification process identified eight instances of fabricated references. While the overall rate is statistically low, the mere existence of fabricated content in the peer-reviewed literature is a critical integrity issue, demonstrating a vulnerability in current academic validation systems. The significance of our finding is not the statistical prevalence but rather the necessity for rigorous, human-validated processes to prevent the injection of spurious citations regardless of their source. Full article
Show Figures

Figure 1

23 pages, 2226 KB  
Data Descriptor
Agricultural Life Cycle Assessment Dataset for Phase 1 Goals, Products, and Scope Definitions
by Rahmah Alhashim and Aavudai Anandhi
Data 2026, 11(5), 121; https://doi.org/10.3390/data11050121 - 20 May 2026
Cited by 1 | Viewed by 762
Abstract
Life cycle assessment (LCA) is widely used to evaluate the environmental impacts of agricultural production systems with four phases. The first phase (Phase 1) is an important phase describing the goal and scope of the entire LCA. However, data on existing and potential [...] Read more.
Life cycle assessment (LCA) is widely used to evaluate the environmental impacts of agricultural production systems with four phases. The first phase (Phase 1) is an important phase describing the goal and scope of the entire LCA. However, data on existing and potential goals and scope are scattered across studies and not available at a single location, making it hard to reuse and compare them. The objective of this study is to create a dataset of Phase 1 information (goals, products, and scopes) for agricultural LCA. The dataset was generated from a systematic review of 184 published agricultural LCA studies, including peer-reviewed journal articles and selected conference papers, published between 1999 and 2025, following PRISMA guidelines. Studies were identified through keyword searches on Google Scholar and screened for relevance and availability. Only studies that clearly reported Phase 1 information were included. Data was collected manually and organized using standard IDs. The dataset has 41 goals, 65 products, and 7 scopes; each was assigned an ID (Goal_ID, Product_ID, Scope_ID, and Stage_ID) to support consistency and traceability. The dataset supports comparisons across studies, assists users in selecting appropriate goals, products, and system boundaries, and can support the development of LCA tools, databases, and decision-support frameworks. Full article
(This article belongs to the Section Information Systems and Data Management)
Show Figures

Figure 1

22 pages, 1421 KB  
Article
Long-Term Repositories—Maintaining Research Databases for Interdisciplinary Projects
by Vincent Feldmar, Tanja Kramm, Constanze Curdt, Dirk Hoffmeister, Olaf Bubenzer and Georg Bareth
Data 2026, 11(5), 120; https://doi.org/10.3390/data11050120 - 19 May 2026
Viewed by 602
Abstract
This paper describes the maintenance and modernisation of custom data repositories that have supported three long-term research projects since their launch in 2007. It highlights an almost complete rewrite of the code base in 2020 and 2021. To enable research data management (RDM) [...] Read more.
This paper describes the maintenance and modernisation of custom data repositories that have supported three long-term research projects since their launch in 2007. It highlights an almost complete rewrite of the code base in 2020 and 2021. To enable research data management (RDM) that adheres to modern and FAIR standards, many features were rebuilt and streamlined for ease of use, with the goal of reducing the friction involved in RDM as much as possible. The update significantly improved the file upload by switching to a fully browser-based solution and completely overhauled the outdated metadata editor with a more interactive version. The ability to search for and find data in the repository has also been enhanced by switching to a flexible, filter-based solution which displays results in real time. It shows that older repositories can be kept in line with the changing landscape of RDM, ensuring that the research data contained therein is not lost. Through these updates and continuing maintenance, these data repositories have stayed available for almost 20 years, even after project funding has ended. Full article
(This article belongs to the Section Information Systems and Data Management)
Show Figures

Figure 1

15 pages, 4382 KB  
Data Descriptor
Genome-Based Characterization of Bacillus velezensis HM1 from Silver Mine Tailings Reveals Potential Metal Resistance and Sulfur Assimilation Traits
by Gustavo Cuaxinque-Flores, Lorena Jacqueline Gómez-Godínez, Marco A. Ramírez-Mosqueda, Jorge David Cadena-Zamudio, Alma Armenta-Medina and José Luis Aguirre-Noyola
Data 2026, 11(5), 119; https://doi.org/10.3390/data11050119 - 15 May 2026
Viewed by 452
Abstract
The genus Bacillus is widely recognized for its metabolic versatility, enabling it to colonize extreme environments, including sites contaminated with metals. In this study, we report the genome of B. velezensis strain HM1, isolated from sulfur-rich mine tailings from silver mining activities in [...] Read more.
The genus Bacillus is widely recognized for its metabolic versatility, enabling it to colonize extreme environments, including sites contaminated with metals. In this study, we report the genome of B. velezensis strain HM1, isolated from sulfur-rich mine tailings from silver mining activities in southwestern Mexico. Isolation was performed by heat treatment followed by selective cultivation in a medium enriched with mine tailings extract (metals and sulfates), resulting in a single dominant morphotype corresponding to strain HM1. Whole-genome sequencing was carried out using the Illumina NovaSeq platform (2 × 250 bp). The assembled genome of strain HM1 has a size of 4,044,128 bp, distributed across 20 contigs, with an N50 of 700,388 bp and an L50 of 3, and an average coverage of 66.8×. The GC content was 46.31%, with an estimated completeness of 99.81% and contamination of 0.01%. Genome analyses indicate that the assembly corresponds to a single chromosome, with no evidence of plasmid replicons. Genome annotation identified 3950 coding sequences (CDSs), 83 tRNAs, 11 rRNAs, 26 ncRNAs, and 4 sORFs. Phylogenomic analysis, together with genomic similarity metrics (ANI > 98.6%, AAI > 98.8%, dDDH > 87%), confirms its classification as Bacillus velezensis. Functionally, the genome encodes multiple genes involved in resistance to metals and metalloids (including ABC transporters, efflux pumps, and biotransformation enzymes), as well as a complete pathway for sulfate assimilation. Collectively, these genomic features reveal a broad repertoire of adaptive strategies employed by strain HM1 to thrive in metal-contaminated environments. Full article
(This article belongs to the Special Issue Benchmarking Datasets in Bioinformatics, 3rd Edition)
Show Figures

Graphical abstract

17 pages, 5716 KB  
Data Descriptor
A Dataset: Experimental Analysis of Outdoor Exposed Four-Year-Old Photovoltaic Modules in Dhaka, Bangladesh
by Md. Sabbir Alam, Ahmed Al Mansur, Shahariar Ahmed Himo, Md. Imamul Islam, Khawza Iftekhar Uddin Ahmed and Md. Fayyaz Khan
Data 2026, 11(5), 118; https://doi.org/10.3390/data11050118 - 14 May 2026
Viewed by 569
Abstract
The long-term performance of photovoltaic (PV) modules significantly affects the reliability and economic viability of solar energy systems, as various environmental and operational factors can gradually degrade module efficiency and reduce energy output. This study investigates the long-term performance degradation analysis of 40 [...] Read more.
The long-term performance of photovoltaic (PV) modules significantly affects the reliability and economic viability of solar energy systems, as various environmental and operational factors can gradually degrade module efficiency and reduce energy output. This study investigates the long-term performance degradation analysis of 40 outdoor photovoltaic (PV) modules exposed for four years on a five-level building in Mirpur, Dhaka, Bangladesh. Electrical parameters, including voltage, current, power, and fill factor, were measured using a PROVA 1011 PV analyzer under IEC60904-1 standard test conditions, and analyzed to evaluate the extent of long-term degradation of PV modules. The image-based analysis identified degradation factors such as dust accumulation, soiling, hotspots, discoloration, micro-cracks, delamination, and corrosion. All test data were normalized to standard conditions (1000 W/m2, 25 °C) for consistency. The measured average maximum power output was 9.85 W, with an average fill factor of 0.713 and a standard deviation of 0.939 for the 40 photovoltaic modules with a rated capacity of 10 W each. The dataset provides valuable insights for researchers and industry professionals to assess long-term PV performance, optimize maintenance strategies, and support solar energy deployment in tropical environments. Additionally, it can aid policymakers in developing regulatory frameworks for improving solar infrastructure resilience. Full article
Show Figures

Figure 1

9 pages, 411 KB  
Data Descriptor
A Rare Earth Elements Database for Peru
by Sergio Ticona, Pablo A. Garcia-Chevesich, Héctor L. Venegas-Quiñones, Guido Salas, Oliver Wanderley Gomez Villagra, Zidane Rooney Pachari Gutierrez, Johanys Trujillo Choque, Madeleine Guillen, Gisella Martínez, Rolando Quispe Aquino, Marcela Huerta, Yezelia Cáceres, Eliseo Zeballos, Mario Nuñez, Cesar Carbajal, Elizabeth Holley and Rod Eggert
Data 2026, 11(5), 117; https://doi.org/10.3390/data11050117 - 13 May 2026
Viewed by 1209
Abstract
The global transition toward low-carbon energy and advanced technologies has intensified demands for rare earth elements (REEs), while available data remain fragmented across government and academic sources. To address this gap, this study compiles and standardizes a publicly accessible geochemical database of REE [...] Read more.
The global transition toward low-carbon energy and advanced technologies has intensified demands for rare earth elements (REEs), while available data remain fragmented across government and academic sources. To address this gap, this study compiles and standardizes a publicly accessible geochemical database of REE concentrations (among other variables) across Peru, motivated by the need for consolidated, evidence-based resource assessment. The dataset integrates over 30,000 records from national agencies and university repositories, classified by sample source (stream sediments, rocks and minerals, deep and surface soils, tailings, and industrial materials). This structure enhances comparability and interpretation across geological and environmental contexts. By providing a centralized, high-resolution dataset for an undercharacterized mineral province, this work offers a resource for exploration, policy development, and sustainable management amid growing global demand and strategic interest in critical minerals. Full article
(This article belongs to the Section Spatial Data Science for Environment and Earth)
Show Figures

Figure 1

11 pages, 13355 KB  
Data Descriptor
A Dataset of Raw Fabric Grayscale Images for Defect Detection
by Ruben Pérez-Llorens, Teresa Albero-Albero and Javier Silvestre-Blanes
Data 2026, 11(5), 116; https://doi.org/10.3390/data11050116 - 12 May 2026
Viewed by 779
Abstract
This article presents RAW-FABRID (RAW FABric Image Dataset), a publicly available annotated dataset for raw fabric defect detection using computer vision techniques. It addresses a major limitation in textile inspection, where reliance on private datasets hinders objective methodological comparisons. RAW-FABRID was acquired using [...] Read more.
This article presents RAW-FABRID (RAW FABric Image Dataset), a publicly available annotated dataset for raw fabric defect detection using computer vision techniques. It addresses a major limitation in textile inspection, where reliance on private datasets hinders objective methodological comparisons. RAW-FABRID was acquired using a custom-built inspection machine equipped with controlled LED illumination and a line-scan camera. The dataset includes grayscale fabric images collected from several manufacturers to ensure variability in textures and patterns. It comprises 709 high-resolution images (1792 × 1024 pixels), including both defect-free and defective samples. To maximize reusability, data are provided in two complementary formats: high-resolution images (cropped to remove peripheral acquisition artifacts) for global analysis, and a patch-based organization following the widely adopted MVTec Anomaly Detection benchmark structure. The latter divides images into 256 × 256 pixel patches for direct machine learning integration. Crucially, the dataset is accompanied by comprehensive metadata (CSV) and precise COCO-formatted annotations (JSON) for both subsets, ensuring full traceability and supporting object detection and semantic segmentation. The dataset is publicly available through Mendeley Data, enabling reproducible research and objective benchmarking of defect detection algorithms. Full article
(This article belongs to the Section Information Systems and Data Management)
Show Figures

Figure 1

14 pages, 3062 KB  
Article
A New Measurement-Based Benchmark Data Set for Radio Spectrum Analysis Applications
by Szilárd László Takács, Lajos Muzsai, Zoltán Németh, Bence Bakos, András Lukács, Csaba Huszty, Péter Vári and András Lapsánszky
Data 2026, 11(5), 115; https://doi.org/10.3390/data11050115 - 11 May 2026
Viewed by 616
Abstract
Radio spectrum is a limited national resource whose efficient utilization is of strategic importance. With the rapid advancement of wireless technologies, maintaining spectrum cleanliness and enabling fast and reliable anomaly detection have become critical challenges. Artificial intelligence (AI)-based approaches have recently shown great [...] Read more.
Radio spectrum is a limited national resource whose efficient utilization is of strategic importance. With the rapid advancement of wireless technologies, maintaining spectrum cleanliness and enabling fast and reliable anomaly detection have become critical challenges. Artificial intelligence (AI)-based approaches have recently shown great promise in addressing these issues; however, their effectiveness strongly depends on the availability of high-quality, representative, and annotated datasets. Generating such datasets is a complex task, further complicated by environmental conditions such as weather. Hungary’s nationwide spectrum monitoring network enables continuous observation of frequency bands, thereby providing the opportunity to construct large-scale and sustainable datasets. This study introduces a measurement methodology designed for the FM sound broadcasting in the VHF band (87.5–108 MHz), presents the resulting dataset, and details the annotation process. The published, openly accessible dataset is expected to serve not only as a valuable reference point but also as a benchmark for the international research community, facilitating the development, validation, and objective comparison of AI-driven spectrum monitoring solutions. Full article
(This article belongs to the Topic Data Stream Mining and Processing)
Show Figures

Figure 1

10 pages, 4057 KB  
Data Descriptor
Dataset for Collaborative Robotics
by Shurook S. Almohamade, John A. Clark and James Law
Data 2026, 11(5), 114; https://doi.org/10.3390/data11050114 - 10 May 2026
Cited by 1 | Viewed by 687
Abstract
This dataset represents the physical interactions collected by the robot’s sensors during a collaborative effort between humans and robots. The experiment was conducted at the Sheffield Robotics laboratory at the University of Sheffield, UK, utilizing a KUKA LBR iiwa 7 R800 serial manipulator. [...] Read more.
This dataset represents the physical interactions collected by the robot’s sensors during a collaborative effort between humans and robots. The experiment was conducted at the Sheffield Robotics laboratory at the University of Sheffield, UK, utilizing a KUKA LBR iiwa 7 R800 serial manipulator. Thirty participants, consisting of 14 males and 16 females, participated, including both students and faculty members. Participants were instructed to guide the robot’s end effector through a two-dimensional maze situated on a horizontal plane. Each participant performed the same task 15 times, resulting in 450 complete interaction sequences. All data files are provided in CSV (comma-separated values) file format, which allows data to be stored in a table-structured format. The complete dataset is publicly available via the Mendeley Data repository (DOI: 10.17632/4fr33dkrjt.3). Full article
Show Figures

Figure 1

13 pages, 5991 KB  
Article
TCM-MS2Link: A Unified AI-Ready Dataset Integrating TCM Herb–Compound Knowledge and MS/MS Spectral Data
by Qianjin Li, Feifan Zhao, Jihang Zhang, Heng Zhou, Lin Guo and Xingchuang Xiong
Data 2026, 11(5), 113; https://doi.org/10.3390/data11050113 - 10 May 2026
Viewed by 605
Abstract
This study presents TCM-MS2Link, a standardized mass spectrometry-based association dataset for traditional Chinese medicine (TCM), serving as an important resource for natural product research in TCM. The dataset adopts a dual-layer “knowledge–data” architecture: the first layer, TCM-MolLink, comprises curated herb–compound association data, constructed [...] Read more.
This study presents TCM-MS2Link, a standardized mass spectrometry-based association dataset for traditional Chinese medicine (TCM), serving as an important resource for natural product research in TCM. The dataset adopts a dual-layer “knowledge–data” architecture: the first layer, TCM-MolLink, comprises curated herb–compound association data, constructed through the integration of multiple heterogeneous databases and rigorous consistency filtering to establish high-confidence relationships between TCM herbs and their chemical constituents; the second layer, MS2-MLReady, is a benchmark dataset for mass spectrometry-based machine learning which, after systematic data cleaning, standardized preprocessing, and well-designed data partitioning, can directly support the training and evaluation of artificial intelligence models. By addressing key limitations in existing public resources, including data fragmentation, inconsistent annotations, and insufficient computational usability, TCM-MS2Link effectively overcomes major bottlenecks in the systematic analysis of TCM components and data-driven research. This study significantly enhances the reliability of herb–compound associations and the modeling readiness of mass spectrometry data, providing a high-quality, standardized, and reusable data foundation for applications such as TCM knowledge base construction and automated spectrum–structure identification, thereby promoting the advancement of TCM informatics and data-driven research. Full article
(This article belongs to the Section Data Science for Chemistry, Energy and Materials)
Show Figures

Figure 1

17 pages, 724 KB  
Article
A Scalable Data Pipeline for Early Detection and Decision Support in Higher Education: YuumCare
by Anabel Pineda-Briseño, María Guadalupe Hernández-Compean, Gabriela Aida Flores-Becerra, María de Jesús Hernández-Quezada and Mayra Manuela De los Santos-Alonso
Data 2026, 11(5), 112; https://doi.org/10.3390/data11050112 - 10 May 2026
Viewed by 1162
Abstract
Early identification of behavioral risk patterns in large student populations remains a challenge in higher education, particularly when support systems depend on voluntary help-seeking. This study presents YuumCare, a structured and scalable framework that operationalizes population-level digital screening through a reproducible data pipeline [...] Read more.
Early identification of behavioral risk patterns in large student populations remains a challenge in higher education, particularly when support systems depend on voluntary help-seeking. This study presents YuumCare, a structured and scalable framework that operationalizes population-level digital screening through a reproducible data pipeline for early detection and decision support. The framework was implemented during the first weeks of the academic term in a public higher education institution in Latin America, where 466 first-year students (38.9% coverage) completed a structured questionnaire capturing indicators of emotional well-being, academic pressure, and help-seeking attitudes. Responses were processed through a structured data pipeline comprising data ingestion, preparation, feature construction, and rule-based classification, transforming distributed self-reported data into standardized features and interpretable institutional signals for consistent analysis at scale. Results show that emotional strain, evaluation-related anxiety, and adaptation difficulties emerge early and frequently co-occur, while most students report low willingness to seek professional support. The classification process indicates that approximately one third of the cohort presents moderate to critical levels of need, providing a structured representation of vulnerability. The proposed approach connects digital screening with institutional decision-making through an interpretable and operational workflow that does not rely on complex infrastructure. Beyond descriptive findings, the study contributes a lightweight and reproducible data framework that supports scalable monitoring and coordinated response under real-world constraints, demonstrating the feasibility of transforming self-reported behavioral data into actionable decision-support signals for population-level monitoring in higher education. Full article
Show Figures

Figure 1

32 pages, 9452 KB  
Article
Intervention to Improve Attitudes Toward Stuttering: A Multi-Site International Replication and Expansion
by Kenneth O. St. Louis, Ben Bolton-Grant, Autumn Cannon, Edna J. Carlo, Sveta Fichman, Shweta Gupta, Krittika Kunda, Hailey M. O’Como, Catherine Porter, Bárbara M. Pratts Pérez, Isabella Reichel, Anne Z. Williams, Salman Abdi, Elizabeth F. Aliveto, Ann Beste-Guldborg, Agata Błachnio, Timothy Flynn, Lejla Junuzović-Žunić, Aneta Przepiórka, Hossein Rezai, Chelsea Roche, Mohyeddin Teimouri Sangani, Michael Azios, Shin Ying Chu, Irena Polewczyk, Cara M. Singer, John A. Tetnowski, Janet S. Tilstra and Katarzyna Węsierskaadd Show full author list remove Hide full author list
Data 2026, 11(5), 111; https://doi.org/10.3390/data11050111 - 8 May 2026
Viewed by 1055
Abstract
Background: Negative public attitudes promote undesirable stereotypes and stigma in stutterers. Method: To mitigate negative attitudes, 403 respondents combined from 16 international samples filled out the Public Opinion Survey of Human Attributes–Stuttering (POSHA–S) before and after interventions to improve attitudes and [...] Read more.
Background: Negative public attitudes promote undesirable stereotypes and stigma in stutterers. Method: To mitigate negative attitudes, 403 respondents combined from 16 international samples filled out the Public Opinion Survey of Human Attributes–Stuttering (POSHA–S) before and after interventions to improve attitudes and were compared to 249 respondents from seven control groups. Investigators aimed (a) to replicate an extreme case of regression to the mean (i.e., “crossover” effect) reported earlier in larger combined samples in which respondents with high pre-scores ended with low post-scores, respondents with low pre-scores finished with high post-scores, and intermediate scorers were unchanged; and (b) to identify individual POSHA–S items related to overall attitude change and among the high and low scorers. Results: As in previous studies, stuttering attitudes improved in the intervention group but not in the control group. Intervention and control respondents demonstrated “crossover” but less than the earlier samples due to lower pre–post correlations. Item contributions to pre–post change and differences among the three change groups were inconsistent; however, high agreement items by respondents were less likely to vary than low agreement items. Conclusion: The “crossover” effect was replicated, and future research should explore its presence in other measures or conditions. Full article
Show Figures

Figure 1

16 pages, 2409 KB  
Data Descriptor
Methodology for Generating, Augmenting, and Validating an Audio Dataset for Classifying Fire and Forest Sounds
by Robert-Nicolae Boştinaru, Sebastian-Alexandru Drǎguşin, Nicu Bizon and Vasile-Gabriel Iana
Data 2026, 11(5), 110; https://doi.org/10.3390/data11050110 - 8 May 2026
Cited by 1 | Viewed by 686
Abstract
This paper proposes a methodological framework for building a binary audio dataset for the automatic classification of fire sounds and forest ambience. Two operational recordings, one for the fire class and one for the forest class, are used strictly as seed data for [...] Read more.
This paper proposes a methodological framework for building a binary audio dataset for the automatic classification of fire sounds and forest ambience. Two operational recordings, one for the fire class and one for the forest class, are used strictly as seed data for controlled segmentation and augmentation. The workflow includes mono conversion at 16 kHz, amplitude normalization, segmentation into 5 s windows with 2 s overlap, low-intensity stochastic augmentation, and the systematic logging of the generated samples. The study also explains why augmented data are appropriate for training and internal validation, while final performance claims must remain reserved for testing on independent, standardized real recordings. Full article
(This article belongs to the Section Information Systems and Data Management)
Show Figures

Figure 1

31 pages, 4120 KB  
Data Descriptor
A Curated Experimental Dataset of UCS and CBR Results from Biopolymer-Based Two-Additive Stabilisation Studies on Fine-Grained Soils
by Abolfazl Baghbani, Delaram Bahrampour, Ahmad Moballegh and Firas Daghistani
Data 2026, 11(5), 109; https://doi.org/10.3390/data11050109 - 8 May 2026
Cited by 1 | Viewed by 912
Abstract
Published laboratory data on soil stabilisation are abundant, yet they remain fragmented across studies and are often difficult to reuse because of inconsistent reporting formats, heterogeneous testing conditions, and incomplete metadata. This article presents a curated experimental dataset compiled from 20 published studies [...] Read more.
Published laboratory data on soil stabilisation are abundant, yet they remain fragmented across studies and are often difficult to reuse because of inconsistent reporting formats, heterogeneous testing conditions, and incomplete metadata. This article presents a curated experimental dataset compiled from 20 published studies on fine-grained soils, comprising 560 records, including 397 unconfined compressive strength (UCS) results and 163 California Bearing Ratio (CBR) results. The dataset is defined by the inclusion of laboratory studies designed around biopolymer-based two-additive stabilisation frameworks, while intentionally retaining untreated and single-additive comparator records reported within the same experimental programmes. This design is a key distinguishing feature of the dataset because it enables analysis of baseline soil behaviour, isolated additive effects, and combined-additive responses within a traceable study context. Across the included studies, the treatment systems cover a wide range of biopolymer- and lignin-related materials, including xanthan gum, guar gum, chitosan, sodium lignosulfonate, and electrolyte lignin stabiliser, together with complementary additives such as cement, lime, fly ash, ground granulated blast-furnace slag, rice husk ash, glass powder, concrete waste, nano-additives, and natural or synthetic fibres. In addition to UCS and CBR outcomes, the dataset preserves key contextual variables required for meaningful secondary reuse, including soil classification, grain-size fractions, Atterberg limits, compaction properties, curing duration, additive identities and dosages, and source-level traceability. The data are distributed as a structured Excel workbook comprising two cleaned outcome-specific sheets (CBR_clean and UCS_clean) and four supporting documentation sheets (StudyInventory, DataDictionary, VocabularyMap, and QC_Log). Record-level identifiers, DOI-linked source fields, inferred-curing flags, and qualified outcome descriptors are retained to support auditability, selective filtering, and reproducible reuse. The resulting dataset provides a practical foundation for comparative assessment of stabilisation strategies, pavement and subgrade engineering studies, meta-analysis, and machine learning applications in geotechnical engineering. Full article
(This article belongs to the Section Information Systems and Data Management)
Show Figures

Graphical abstract

14 pages, 776 KB  
Article
A Data-Driven Approach to Cardiometabolic Risk Stratification: Development of the Adiposity-Fitness Imbalance Index Using a National Chilean Dataset
by Rodrigo Yáñez-Sepúlveda, José Francisco Tornero-Aguilera, Mario Muñoz-López, Edgar Sancho-Haro, Yeny Concha-Cisternas, Exal Garcia-Carrillo, Jacqueline Páez-Herrera, Felipe Montalva-Valenzuela and Eduardo Guzmán-Muñoz
Data 2026, 11(5), 108; https://doi.org/10.3390/data11050108 - 8 May 2026
Viewed by 873
Abstract
The increasing prevalence of adolescent obesity and declining physical fitness highlights the need for integrative, non-invasive tools to identify central-adiposity–related cardiometabolic risk early. This study aimed to develop and analytically evaluate the adiposity–fitness imbalance (AFI) index and to examine its association with an [...] Read more.
The increasing prevalence of adolescent obesity and declining physical fitness highlights the need for integrative, non-invasive tools to identify central-adiposity–related cardiometabolic risk early. This study aimed to develop and analytically evaluate the adiposity–fitness imbalance (AFI) index and to examine its association with an anthropometric proxy of cardiometabolic risk (waist-to-height ratio > 0.50) in a nationally representative sample of Chilean adolescents. This cross-sectional study analyzed data from 7852 students from the Chilean National Physical Fitness Assessment System (SIMCE-EF). The AFI index was calculated as the difference between standardized adiposity and fitness components. Logistic and robust linear regression models were used. Higher standing long jump (OR = 0.69, 95% CI 0.65–0.74), push-ups (OR = 0.76, 95% CI 0.71–0.80), sit-ups (OR = 0.81, 95% CI 0.77–0.85), and VO2max (OR = 0.82, 95% CI 0.75–0.89) were associated with lower odds of elevated WHtR (all p < 0.001), and a small protective association was also observed for flexibility (OR = 0.93, 95% CI 0.88–0.99, p = 0.016). Each one-standard-deviation increase in the AFI index was associated with a substantially higher odds of elevated WHtR (OR = 26.74, 95% CI 22.57–31.68, p < 0.001). In a sensitivity analysis that removed WHtR from the adiposity pillar, to avoid component–outcome overlap, the AFI index remained strongly associated with the outcome (OR per 1 SD = 14.60, 95% CI 12.77–16.70), with internal-validation discrimination of AUC = 0.93. The AFI index may represent a practical and scalable tool for early screening of central-adiposity–related risk in adolescents. Full article
(This article belongs to the Section Information Systems and Data Management)
Show Figures

Figure 1

11 pages, 2604 KB  
Article
FD-TamperBoard: A Tampering Features Dataset of Fuel Dispenser PCBs for Illicit Metering Detection
by Chenbo Pei, Bin Wang, Xingchuang Xiong, Zhanshuo Cao and Zilong Liu
Data 2026, 11(5), 107; https://doi.org/10.3390/data11050107 - 7 May 2026
Viewed by 809
Abstract
With the development of the Internet of Things (IoT) and microelectronics technology, the methods used to tamper with fuel dispensers have become increasingly concealed, posing significant challenges to market supervision and law enforcement. This paper releases a tampering features dataset of assembled printed [...] Read more.
With the development of the Internet of Things (IoT) and microelectronics technology, the methods used to tamper with fuel dispensers have become increasingly concealed, posing significant challenges to market supervision and law enforcement. This paper releases a tampering features dataset of assembled printed circuit boards (PCBs) from fuel dispensers, aiming to provide high-quality data support for automated, computer-vision-based illicit metering detection. The dataset encompasses multi-class tampering features derived from 189 high-resolution images of PCBs seized during real-world law enforcement, covering 5 mainstream brands. To eliminate perspective bias, rigorous lens distortion correction and four-point homography transformation preprocessing were conducted on the images. Additionally, six typical tampering features (e.g., the addition of tampered surface-mount resistors) were manually and precisely annotated, and then cross-checked and confirmed by domain experts. Furthermore, the dataset was benchmarked using multiple generations of You Only Look Once (YOLO) object detection models (Baseline Validation), which have been demonstrated to handle both large and small object detection in high-resolution images. The evaluation results, including confusion matrices and t-distributed Stochastic Neighbor Embedding (t-SNE) feature clustering diagrams, demonstrate the reliability and effectiveness of this dataset for training high-precision fraud detection models. This dataset is intended to support computer vision and anti-fraud research, promoting the automated development of fuel dispenser tampering detection. Full article
Show Figures

Graphical abstract

16 pages, 2666 KB  
Data Descriptor
Multimodal Dataset of In-Home Physiological and Inertial Measurements from Older Heart Failure Patients
by Marcin Kolakowski, Vitomir Djaja-Josko, Jerzy Kolakowski, Irina Georgiana Mocanu, Oana Cramariuc, Ian Perera, Jerzy Gąsowski and Karolina Piotrowicz
Data 2026, 11(5), 106; https://doi.org/10.3390/data11050106 - 7 May 2026
Viewed by 1900
Abstract
Numerous studies have shown that remote monitoring of heart failure patients can reduce hospital readmission rates and mortality. This dataset includes multimodal physiological and inertial signals (acceleration and angular velocity data) recorded with PerHeart—a remote health monitoring platform intended for heart failure patients. [...] Read more.
Numerous studies have shown that remote monitoring of heart failure patients can reduce hospital readmission rates and mortality. This dataset includes multimodal physiological and inertial signals (acceleration and angular velocity data) recorded with PerHeart—a remote health monitoring platform intended for heart failure patients. In the pilot, which took place in Poland, 27 participants’ health was monitored for one month using the platform with commercially available devices (blood pressure meters, pulse oximeters, bathroom scales, thermometers, and glucometers), resulting in over four thousand physiological measurements. Eight adults were additionally monitored for gait and activity analysis using custom wrist sensors with inertial measurement units, yielding 2536 h of movement data collected over 204 days with almost 690,000 steps detected. Full article
(This article belongs to the Special Issue Benchmarking Datasets in Bioinformatics, 3rd Edition)
Show Figures

Graphical abstract

22 pages, 1218 KB  
Article
A Conceptual Framework for Semantic Indexing of Data Sources Based on Structured Peer-to-Peer Model, Hilbert Curve, Hypercube and Data Analysis
by Mohammed Ammari, Fadwa Ammari and Abdelaziz Boumahdi
Data 2026, 11(5), 105; https://doi.org/10.3390/data11050105 - 5 May 2026
Viewed by 504
Abstract
Semantic indexing ensures better organization and optimized searching of heterogeneous, autonomous, and distributed data sources. This approach leverages meaning and context rather than just keywords to better manage the increasing volume, complexity, and heterogeneity of modern data, enabling precise searching, optimized integration, and [...] Read more.
Semantic indexing ensures better organization and optimized searching of heterogeneous, autonomous, and distributed data sources. This approach leverages meaning and context rather than just keywords to better manage the increasing volume, complexity, and heterogeneity of modern data, enabling precise searching, optimized integration, and improved interoperability between domains. Several approaches to semantic indexing are available: ontology-based indexing, machine learning and automated semantic annotation of data sources. However, the main challenge remains scaling up. This article focuses on a conceptual framework designed for scalable semantic indexing of data sources based on a structured peer-to-peer architecture adapted for managing a very large number of nodes, Hilbert curve renowned for its preservation of semantic affinity while scaling, hypercube structure with its efficient diffusion algorithm, semantic annotation of data sources based on keywords, as well as machine learning techniques, in particular, multidimensional data analysis. An illustrative exploratory example of the Meta Skills semantic class is presented to outline the proposed architecture. This study proposes a conceptual and exploratory framework for large-scale semantic indexing of data sources. The proposed approach has not yet been implemented or validated on a large scale; its objective is to provide an initial structured model to serve as a basis for future empirical research. Full article
(This article belongs to the Section Information Systems and Data Management)
Show Figures

Figure 1

17 pages, 808 KB  
Article
Development of Intra-Individual Process Metrics in a Serious-Video Game Intervention for ADHD
by Marina Martin-Moratinos, Marcos Bella-Fernández, Maria Rodrigo-Yanguas, Carlos González-Tardón, Aarón Sújar and Hilario Blasco-Fontecilla
Data 2026, 11(5), 104; https://doi.org/10.3390/data11050104 - 5 May 2026
Viewed by 552
Abstract
(1) Background: Attention-deficit/hyperactivity disorder (ADHD) is characterized by persistent difficulties related to inattention, hyperactivity, and impulsivity, which significantly impair daily functioning. The primary objective of this study is to examine the utility of intra-individual metrics as indicators of dynamic cognitive regulation during the [...] Read more.
(1) Background: Attention-deficit/hyperactivity disorder (ADHD) is characterized by persistent difficulties related to inattention, hyperactivity, and impulsivity, which significantly impair daily functioning. The primary objective of this study is to examine the utility of intra-individual metrics as indicators of dynamic cognitive regulation during the intervention with a serious video game (The Secret Trail of Moon, MOON). (2) Methods: Performance data were collected from participants with ADHD enrolled in a randomized clinical trial. Within the MOON group, intra-individual metrics were derived from repeated gameplay sessions of a continuous performance task. For each participant, simple linear regression models were used to estimate the slope of performance across repeated exposures to the task. Slopes were interpreted as indicators of intra-individual change over time. The within-subject standard deviation was also calculated to observe how much a person’s performance fluctuates between sessions. (3) Results: A total of 76 patients with ADHD participated in the clinical trial and were randomized in a 1:1 ratio (MOON: n = 38, 50% and control: n = 38, 50%). The mean performance index of the MOON group (M = 0.88, SD = 0.09) indicates a generally high level of response accuracy, with moderate inter-individual variability across participants. Notably, moderate intra-individual variability (e.g., RT variability, lapse-related indices) was observed, suggesting fluctuations in attentional control despite stable average performance. The absence of linear improvement should not be interpreted as a lack of intervention effect, but rather as evidence of rapid task familiarization and ceiling effects. (4) Conclusions: Intra-individual variability may be a key metric for understanding attentional control in ecological, game-based environments. In this context, performance variability and attentional stability emerge as more sensitive indicators of cognitive regulation than mean-level changes. Full article
(This article belongs to the Section Information Systems and Data Management)
Show Figures

Figure 1

18 pages, 708 KB  
Article
NSCH-Flourishing-ML: A Curated Dataset and Reproducible Pipeline for Machine Learning Analysis of Child Flourishing
by Miguel Arcos-Argudo, Rodolfo Bojorque, Fernando Pesántez and Kely Nieto-Andrade
Data 2026, 11(5), 103; https://doi.org/10.3390/data11050103 - 3 May 2026
Viewed by 831
Abstract
Large-scale population surveys provide valuable information for studying child well-being, yet their structure often limits the direct application of machine-learning methods. The National Survey of Children’s Health (NSCH) is one of the most comprehensive datasets for monitoring children’s health and development in the [...] Read more.
Large-scale population surveys provide valuable information for studying child well-being, yet their structure often limits the direct application of machine-learning methods. The National Survey of Children’s Health (NSCH) is one of the most comprehensive datasets for monitoring children’s health and development in the United States, but the raw survey files contain logical skip patterns, categorical variables, and complex survey-design elements that require substantial preprocessing before predictive analysis can be performed. This study presents a curated machine-learning-ready benchmark dataset derived from the 2023 NSCH together with a fully reproducible computational pipeline for studying school-age child flourishing. The workflow constructs a binary flourishing outcome from four survey items related to curiosity, task persistence, emotional self-regulation, and interest in doing well in school. After restricting the sample to children aged 6–17 years and retaining only records with valid responses in all four outcome items, the final analytical dataset contained 32,934 observations. Feature selection based on mutual information computed on the training partition, combined with cross-validated subset-size selection, yielded a final benchmark subset of 150 predictors. Baseline experiments using logistic regression and random forest showed stable and reasonably strong predictive performance, with held-out ROC-AUC values around 0.84–0.85 and closely aligned cross-validation results. An exploratory comparison between weighted and unweighted learning further showed that survey weighting did not improve discriminative performance in this benchmark setting, although the magnitude of the effect was modest and model-dependent. By releasing both the curated benchmark dataset and the reproducible pipeline, this study provides a reusable resource for machine-learning research on child well-being and survey-based computational benchmarking. Full article
(This article belongs to the Topic Machine Learning and Data Mining: Theory and Applications)
Show Figures

Figure 1

31 pages, 2038 KB  
Article
Quantifying the Key Performance Indicators of Success: An Exploratory Analysis of Champion Teams in Europe’s Top Football Leagues
by José Gama, Gonçalo Dias, Rodrigo Mendes, Fernando Martins, Rui Sousa Mendes and Vasco Vaz
Data 2026, 11(5), 102; https://doi.org/10.3390/data11050102 - 2 May 2026
Viewed by 2801
Abstract
This study quantified performance indicators associated with match outcomes among champion teams from the five major European football leagues during the 2023–2024 season. Ordinal logistic regression with robust standard errors clustered by team was employed, with analyses stratified by match location (home/away) and [...] Read more.
This study quantified performance indicators associated with match outcomes among champion teams from the five major European football leagues during the 2023–2024 season. Ordinal logistic regression with robust standard errors clustered by team was employed, with analyses stratified by match location (home/away) and opponent quality (high/medium/low). Data from 182 matches were sourced from Wyscout® and included offensive indicators (possession, passes, shots, shots on target, expected goals) and defensive indicators (interceptions, fouls, shots conceded, yellow and red cards). Spearman correlations showed that goals scored (q=0.523) and shots on target (q=0.243) were positively associated with match outcomes, whereas goals conceded (q=0.441) and fouls (q=0.255) were negatively associated. Ordinal regression revealed context-dependent effects. Offensively, shots on target increased the odds of a better outcome at home (OR = 3.76) and against high-quality opponents (OR = 5.24), while expected goals (xG) was the key predictor in away matches (OR = 2.09). Defensively, interceptions were crucial against high-quality opponents (OR = 1.76), while fouls (OR = 0.53) and yellow cards (OR = 0.61) were detrimental against medium-quality opponents. Against low-quality opponents, shots on target conceded (OR = 0.22) and red cards (OR = 66.58) were critical. Volume-based indicators did not retain significant independent effects. For elite champion teams, competitive success is predominantly determined by efficiency-based indicators, shot accuracy, expected goals, and defensive organisation, whose relevance varies systematically with context. These findings provide exploratory insights and a context-sensitive benchmark for performance analysis at the highest level of European football, warranting further validation in future studies. Full article
(This article belongs to the Special Issue Big Data and Data-Driven Research in Sports)
Show Figures

Figure 1

23 pages, 19482 KB  
Data Descriptor
An Open Industrial Energy Dataset with Asset-Level Measurements and High-Coverage 15-Minute Aggregates from a Manufacturing Facility
by Christopher Flynn, Trevor Murphy, Joseph Walsh and Daniel Riordan
Data 2026, 11(5), 101; https://doi.org/10.3390/data11050101 - 1 May 2026
Viewed by 1706
Abstract
Publicly available electricity datasets from operational industrial facilities remain limited due to instrumentation cost, retrofit complexity, and data governance constraints. This paper presents an openly accessible dataset of asset-level electrical energy measurements collected from a medium-scale industrial manufacturing facility over an approximately one-year [...] Read more.
Publicly available electricity datasets from operational industrial facilities remain limited due to instrumentation cost, retrofit complexity, and data governance constraints. This paper presents an openly accessible dataset of asset-level electrical energy measurements collected from a medium-scale industrial manufacturing facility over an approximately one-year observation window, with staged commissioning resulting in heterogeneous temporal coverage. The dataset includes time-series measurements from production machinery, auxiliary systems, and distribution-level assets instrumented using a heterogeneous fleet of Ethernet and RS-485 energy meters integrated via industrial gateways and programmable logic controllers. Measurements were acquired via a SCADA-based logging infrastructure and exported from an operational SQL historian. The publicly released dataset comprises fixed 15 min aggregated energy and power metrics derived from high-frequency SCADA telemetry. In its released ALL-phase representation, the dataset comprises measurements from 43 monitored assets and 1,039,873 15 min windows, corresponding to 2.96 GWh of measured electrical energy. Mean window-level data coverage is 99.99%, and 97.72% of ALL-phase windows satisfy the dataset’s reliability criterion. Interval records include energy consumption, demand, data coverage metrics, and reliability indicators. The dataset reflects real-world industrial monitoring conditions, including mixed communication pathways and irregular sampling behaviour, and is intended to support research in industrial energy analytics, data quality assessment, load profiling, and operational energy modelling. Full article
Show Figures

Figure 1

9 pages, 1210 KB  
Data Descriptor
Preferred Colleague Dataset: A Human-Annotated Dataset of Perceived Colleague Preference
by Deepu Krishnareddy, Bakir Hadžić, Hamid Gazerpour, Michael Danner, Zhuoqi Zeng and Matthias Rätsch
Data 2026, 11(5), 100; https://doi.org/10.3390/data11050100 - 1 May 2026
Viewed by 817
Abstract
Recruitment is a time-consuming process, and AI systems are increasingly being used to support the decision-making process. However, machine learning models used in such systems can inherit bias if the underlying training data reflects biased human preferences. It is essential to analyze and [...] Read more.
Recruitment is a time-consuming process, and AI systems are increasingly being used to support the decision-making process. However, machine learning models used in such systems can inherit bias if the underlying training data reflects biased human preferences. It is essential to analyze and quantify these biases in order to develop fairer AI systems. To address this issue, we collected human judgments of colleague preference for 2200 face images. The face image set includes images of different ethnicities and genders, as well as both real and synthetically generated faces. The images were annotated by humans from diverse backgrounds in terms of age, gender, and ethnicity. Annotators were shown series of pairs of face images and asked to select which individual they would prefer as a colleague. We gathered responses from 451 annotators and aggregated the annotations to compute a preference score for each image. This dataset provides a basis for understanding human bias in colleague preference and can support the development of fair and unbiased AI models for use in recruitment settings. Full article
Show Figures

Figure 1

8 pages, 528 KB  
Data Descriptor
Whole-Genome Sequencing Dataset from Two High-Risk Breast Cancer Families Negative for BRCA1/2 and Other Known Susceptibility Genes
by Silvia González-Martínez, Alejandra Rezqallah Arón, José Manuel Pérez-García, José Palacios, Belén Pérez-Mies, Javier Román, Laia Garrigos, Judith Balmaña, Daniela Camacho, Sandra Íñiguez-Muñoz, Diego M. Marzese and Javier Cortés
Data 2026, 11(5), 99; https://doi.org/10.3390/data11050099 - 30 Apr 2026
Viewed by 927
Abstract
Hereditary breast cancer (BC) remains unexplained in a substantial proportion of families who test negative for BRCA1/2 and other known susceptibility genes. To contribute to the genomic characterization of these unresolved cases, we generated a whole-genome sequencing (WGS) dataset from six women belonging [...] Read more.
Hereditary breast cancer (BC) remains unexplained in a substantial proportion of families who test negative for BRCA1/2 and other known susceptibility genes. To contribute to the genomic characterization of these unresolved cases, we generated a whole-genome sequencing (WGS) dataset from six women belonging to two unrelated high-risk families, each comprising three sisters diagnosed with BC. All participants had previously received negative results in conventional multigene panel testing. WGS was performed on peripheral blood DNA using the Illumina NovaSeq platform, followed by variant calling against GRCh38 and the comprehensive annotation of single-nucleotide variants, indels, and structural variants. For each family, we identified shared ClinVar-annotated variants, rare exonic or splice-site alterations, and intronic variants located within a curated set of 286 cancer-related genes. The dataset includes per-patient VCF files, copy number variation annotations, and family-level variant summaries. Raw and processed data are publicly available through the Sequence Read Archive and Zenodo. This resource supports variant reinterpretation, exploration of regulatory and intronic regions, and methodological benchmarking in the study of familial BC beyond established susceptibility genes. Full article
Show Figures

Figure 1

20 pages, 1275 KB  
Article
Machine Learning Models for Predicting Professional Disqualification in Peruvian Association Members
by Manuel Pretel Pretel, Yeny Chávez Llempén, Abel Angel Sullon Macalupu, Paulo Canas Rodrigues, Javier Linkolk López-Gonzales and Esteban Tocto-Cano
Data 2026, 11(5), 98; https://doi.org/10.3390/data11050098 - 30 Apr 2026
Viewed by 967
Abstract
The disqualification of licensed professionals for non-payment of their monthly fees constitutes a significant operational risk to the financial sustainability of professional associations. This problem highlights the need for predictive tools that can anticipate the risk of disqualification and protect institutional stability. The [...] Read more.
The disqualification of licensed professionals for non-payment of their monthly fees constitutes a significant operational risk to the financial sustainability of professional associations. This problem highlights the need for predictive tools that can anticipate the risk of disqualification and protect institutional stability. The main objective of this study was to develop a supervised machine learning model for estimating the risk of disqualification among registered professionals based on historical and contextual variables. An empirical, applied, and quantitative study was conducted by analyzing more than 5.7 million financial records corresponding to 27,964 registered professionals. Multiple supervised classification algorithms, including ensemble models such as CatBoost and XGBoost, were evaluated using stratified cross-validation and class-balancing techniques to address the substantial imbalance in the data. The results indicated that CatBoost performed best (F1-score = 57.96%; AUC = 0.72), whereas XGBoost showed greater stability across cross-validation folds. In conclusion, the model developed supports the timely identification of members at high-risk of disqualification, enabling the implementation of early warning systems and proactive institutional financial management strategies. Full article
Show Figures

Figure 1

12 pages, 2488 KB  
Article
Bibliometric Analysis of the Literature Regarding MRI-Linac: A Paradigm Shift in Radiation Oncology
by Andrea Emanuele Guerini, Paolo Rondi, Federico Mastroleo, Stefania Volpe, Stefano Riga, Stefania Nici, Marco Luzzara, Giulio Ferrazzi, Marco Krengli, Davide Farina, Luigi Spiazzi, Barbara Alicja Jereczek-Fossa, Marco Ravanelli and Michela Buglione di Monale e Bastia
Data 2026, 11(5), 97; https://doi.org/10.3390/data11050097 - 28 Apr 2026
Viewed by 825
Abstract
Background: By integrating an MRI scanner and a linear accelerator, MR-linac systems provide superior soft tissue imaging and allow to perform adaptive radiotherapy adjusted on daily anatomical changes. The advent of this technology represents a revolution in radiation oncology and could improve treatment [...] Read more.
Background: By integrating an MRI scanner and a linear accelerator, MR-linac systems provide superior soft tissue imaging and allow to perform adaptive radiotherapy adjusted on daily anatomical changes. The advent of this technology represents a revolution in radiation oncology and could improve treatment accuracy and clinical outcomes. We performed a comprehensive bibliometric analysis with the aim of displaying the available scientific literature and trends regarding MR-linac. Methods: Scopus database was investigated, considering documents published up to 6 April 2025. Keywords encompassed terms related to “MR-linac” or “MRI-linac” and possible combinations and acronyms. BibTeX data file was imported into Biblioshiny (Bibliometrix package—v. 4.1.4) and analysis was conducted using R code (R version 4.3.2) and the Bibliometrix package (version 4.1.4). Results: A total of 1624 articles on MR-linac were identified. The number of annual publications gradually increased from 21 in 2008, peaking at 211 in 2022 and then remaining substantially stable in subsequent years. Most of the papers were original articles (79.2%) and the majority was published by the 10 journals with the largest output. Remarkably, of 6385 identified authors, over 85% were from one of the 10 most represented countries (including European, North American and Asian nations). Consistently, the 10 institutions with the larger output were North American, Australian or European and provided over 60% of the articles. International co-authorship was found in only 23.6% of the articles. Keyword and co-occurrence analyses identified MR-guided radiotherapy, SBRT, dosimetry, and adaptive strategies as core themes, with emerging trends in radiomics, diffusion metrics, and deep learning. Conclusions: Bibliometric analysis identified trends and patterns of scientific publications regarding MR-linac, highlighting a growing interest in the topic. Nonetheless, it should be considered that the majority of the papers were published by a few journals and over 85% of authors were from 10 countries, demonstrating an evident disparity across nations. Multicentric international research protocols and common frameworks could foster the transition towards collaborative practice-changing studies. Full article
Show Figures

Figure 1

9 pages, 1748 KB  
Data Descriptor
Draft Genome Sequence Data of Multidrug-Resistant Pseudomonas aeruginosa, Strain ASK-80
by Shilippreet Kour, Shilpa Sharma, Achhada Ujalkaur Avatsingh, Prem Prashant Chaudhary and Nasib Singh
Data 2026, 11(5), 96; https://doi.org/10.3390/data11050096 - 26 Apr 2026
Cited by 1 | Viewed by 951
Abstract
In this study, we report the draft genome sequence of Pseudomonas aeruginosa strain ASK-80, a multidrug-resistant bacterium isolated from municipal wastewater in Baddi, district Solan, Himachal Pradesh, India. The whole genome was sequenced through Illumina MiSeq sequencing (150 bp paired-end). The size of [...] Read more.
In this study, we report the draft genome sequence of Pseudomonas aeruginosa strain ASK-80, a multidrug-resistant bacterium isolated from municipal wastewater in Baddi, district Solan, Himachal Pradesh, India. The whole genome was sequenced through Illumina MiSeq sequencing (150 bp paired-end). The size of the assembled genome was 6,261,345 bp, and the genome annotation revealed 5834 genes, including 5778 CDSs, 5748 protein-coding genes, 56 RNA genes and 30 pseudo genes. Genomic characterization revealed the occurrence of multiple antibiotic resistance genes (blaOXA-396, blaOXA-486, blaOXA-494, blaPAO, blaPDC-8, aph(3)-IIb, catB7, fosA and others), virulence genes (algB, chpA, clpV1, exsA, flgA, pilB, pvcA, toxA, tse1, and waaA), insertion sequences, transposable elements and phage sequences. This genome data may serve as a valuable resource for comparative genomics of P. aeruginosa and research on the antibiotic resistance surveillance of wastewater. Full article
(This article belongs to the Special Issue Benchmarking Datasets in Bioinformatics, 3rd Edition)
Show Figures

Figure 1

Previous Issue
Next Issue
Back to TopTop