Sign in to use this feature.

Years

Between: -

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (1,480)

Search Parameters:
Journal = Data

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
20 pages, 512 KB  
Data Descriptor
Phase 2: Agricultural Life Cycle Inventory Dataset (Inputs, Outputs, and Data Sources)
by Rahmah Alhashim and Aavudai Anandhi
Data 2026, 11(9), 213; https://doi.org/10.3390/data11090213 - 26 Aug 2026
Abstract
Life cycle inventory (LCI) data constitute Phase 2 of life cycle assessment (LCA) and serve as the basis for calculating environmental impacts. In agricultural LCA studies, LCI data are often reported using different scopes, stages, units, and calculation methods, making it difficult to [...] Read more.
Life cycle inventory (LCI) data constitute Phase 2 of life cycle assessment (LCA) and serve as the basis for calculating environmental impacts. In agricultural LCA studies, LCI data are often reported using different scopes, stages, units, and calculation methods, making it difficult to reproduce results and compare studies. The objective of this study is to compile and standardize Phase 2 LCI data reported in agricultural LCA studies into a structured dataset. The dataset is based on 184 peer-reviewed agricultural LCA studies published between 1999 and 2025. Data were collected through a systematic review using Google Scholar, and studies were included if they applied LCA to crop production systems and reported inventory data such as inputs, outputs, emission factors, or calculation equations. Inventory data were manually extracted from each study, including inputs and outputs, emission factors, equations, and data sources. The dataset is provided as an Excel workbook containing linked sheets for study identifiers, inputs, outputs, emission factors, equations, and data sources. Rather than providing newly harmonized inventory values, the dataset organizes extracted and categorized information reported in the reviewed studies using standardized identifiers and categories. The dataset includes more than 2000 inputs, around 1000 outputs, about 200 emission factors, and over 600 data sources. It is intended for researchers, practitioners, and tool developers to support LCI development, cross-study comparison, and integration into databases, knowledge bases, and decision-support tools. Full article
(This article belongs to the Section Information Systems and Data Management)
Show Figures

Figure 1

10 pages, 2139 KB  
Opinion
Digital Barriers Still Hindering the Retrieval and Analysis of Historical Dark Data in Phenology
by Nagai Shin, Taku M. Saitoh and Chifuyu Katsumata
Data 2026, 11(9), 212; https://doi.org/10.3390/data11090212 - 24 Aug 2026
Abstract
To deepen our understanding of human–ecosystem interactions, researchers need to be able to retrieve and analyze historical dark data such as plant and animal phenology, but there are often barriers to doing so. Despite the development of online digitization and other tools, including [...] Read more.
To deepen our understanding of human–ecosystem interactions, researchers need to be able to retrieve and analyze historical dark data such as plant and animal phenology, but there are often barriers to doing so. Despite the development of online digitization and other tools, including library search engines, digital collections, machine translation, OCR (optical character recognition), HTR (handwritten text recognition), and generative AI technologies, and the establishment of standards and frameworks (e.g., FAIR Principles and the International Image Interoperability Framework), barriers to converting analog records to digital records (“digital barriers”) and to translating local languages to an international common language (“language barriers”) still remain. We present a case study example of the use of historical dark data in phenology in Japan and the digital and language barriers encountered. We then briefly summarize factors and challenges hindering use of this data and describe the benefits of further removal of these barriers. Full article
(This article belongs to the Section Featured Reviews of Data Science Research)
Show Figures

Figure 1

18 pages, 382 KB  
Data Descriptor
A Georeferenced Dataset of Electromagnetic Field Exposure Measurements in Colombia
by David L. Ocampo-Rodríguez, Diógenes de Jesus Ramirez-Ramirez and Cristian David Correa-Álvarez
Data 2026, 11(8), 211; https://doi.org/10.3390/data11080211 - 21 Aug 2026
Viewed by 124
Abstract
This Data Descriptor presents a georeferenced dataset of electromagnetic-field exposure measurements collected in Colombia between 2 November 2023 and 30 June 2024. The records originate from the Sistema de Monitoreo de Campos of the Agencia Nacional del Espectro (ANE), which uses isotropic probes [...] Read more.
This Data Descriptor presents a georeferenced dataset of electromagnetic-field exposure measurements collected in Colombia between 2 November 2023 and 30 June 2024. The records originate from the Sistema de Monitoreo de Campos of the Agencia Nacional del Espectro (ANE), which uses isotropic probes to monitor broadband radiofrequency electromagnetic fields from 100 kHz to 8 GHz and reports six-minute averages of incident power density in W/m2. The comma-separated value file contains 1,286,346 timestamped records and seven source variables, including Well-Known Text (WKT) point geometry. The principal 2024 subset comprises 995,238 records from 24 monitoring stations. We provide a reproducible workflow for timestamp and coordinate parsing, structural and numeric validation, station-level aggregation, and spatial sensitivity analysis using inverse distance weighting with the mean, median, and 95th percentile. The dataset does not provide frequency-resolved measurements, calibration certificates, or measurement-uncertainty budgets; these omissions limit the use of the public file for formal regulatory-compliance assessment. The accompanying repository includes validated station-level summaries, descriptive tables, figures, and reproducible R code. The package supports environmental monitoring, geospatial analysis, methodological comparison, and reproducible reuse of electromagnetic-field exposure data. Full article
Show Figures

Figure 1

13 pages, 8160 KB  
Data Descriptor
Dataset on the Biodiversity and Seasonal Dynamics of Horseflies (Tabanidae, Diptera) in Some Regions of European Russia
by Irina A. Budaeva, Sergei V. Pestov, Alexander B. Ruchin, Sergei V. Lukiyanov, Evgeniy A. Lobachev, Mikhail N. Esin and Irina G. Esina
Data 2026, 11(8), 210; https://doi.org/10.3390/data11080210 - 20 Aug 2026
Viewed by 185
Abstract
The regional fauna of many Diptera groups is insufficiently studied. Tabanidae are the largest representatives of blood-sucking Diptera. A description of a dataset on the biodiversity of horseflies (Tabanidae) is presented, which includes information obtained by the authors and their colleagues in 2001, [...] Read more.
The regional fauna of many Diptera groups is insufficiently studied. Tabanidae are the largest representatives of blood-sucking Diptera. A description of a dataset on the biodiversity of horseflies (Tabanidae) is presented, which includes information obtained by the authors and their colleagues in 2001, 2007–2009, and 2012–2025. The information was received from 14 regions of European Russia. A total of 5740 specimens were processed. In total, samples were obtained from 159 localities. Each record includes information about the species, time and place of collection, the collectors and determiners, as well as the administrative affiliation of the locality, making the dataset suitable for faunistic, biogeographical, and ecological studies. In total, 35 species were reliably identified. The most studied was the Republic of Mordovia, from which 33 species of Tabanidae became known. In turn, 12 species are included in the dataset from the Vladimir Region, and 10 species are included in the dataset from the Ryazan Region. Five species (Chrysops caecutiens, Chrysops viduatus, Haematopota pluvialis, Hybomitra bimaculata, Hybomitra muehlfeldi) were represented in the dataset in the largest number, and they accounted for 61.5% of the total specimens. The dataset also includes one specimen each of five species: Atylotus plebeius, Chrysops nigripes, Chrysops sepulcralis, Hybomitra borealis, and Hybomitra kaurii. Tabanidae activity begins in the first ten days of May and ends in the second ten-day period of September. The maximum species richness was observed in the third ten-day period of June and the first ten days of July, with 24 species recorded in each of these periods. Meanwhile, the maximum value of the Shannon index (2.52) and the minimum value of the Berger–Parker index (0.16) are observed in the first ten days of July. The dataset can be used as a basis for further study of the diversity and distribution of Tabanidae in European Russia. Full article
(This article belongs to the Section Spatial Data Science for Environment and Earth)
Show Figures

Figure 1

27 pages, 11115 KB  
Article
Learning Adaptive Cross-Modal Interactions for Multimodal Sentiment Analysis
by Chuhan Cheng, Hangcheng Wu, Junqiao Wang and Yuqi Ouyang
Data 2026, 11(8), 209; https://doi.org/10.3390/data11080209 - 20 Aug 2026
Viewed by 197
Abstract
Multimodal sentiment analysis aims to integrate heterogeneous textual, visual, and acoustic information for effective emotion understanding. However, existing methods often suffer from insufficient cross-modal interaction modeling, limited adaptability in multimodal fusion, and inadequate suppression of modality-specific noise under complex conversational scenarios. To address [...] Read more.
Multimodal sentiment analysis aims to integrate heterogeneous textual, visual, and acoustic information for effective emotion understanding. However, existing methods often suffer from insufficient cross-modal interaction modeling, limited adaptability in multimodal fusion, and inadequate suppression of modality-specific noise under complex conversational scenarios. To address these challenges, this paper proposes a framework for learning adaptive cross-modal interactions for multimodal sentiment analysis. The proposed framework consists of three stages: modality-aware preprocessing, heterogeneous representation learning, and adaptive multimodal fusion. First, a unified preprocessing strategy is designed to improve cross-modal consistency through textual normalization, speaker-aware visual alignment, and utterance-level acoustic representation enhancement. Second, modality-specific encoders are constructed to capture complementary semantic, spatial, and utterance-level acoustic characteristics from textual, visual, and acoustic modalities, respectively. Third, an adaptive fusion framework is introduced to explicitly model cross-modal interactions, dynamically estimate the importance of different modality combinations, and further calibrate discriminative feature channels through channel attention. By jointly performing modality-level interaction learning and channel-wise feature refinement, the proposed framework effectively enhances multimodal representation capability for sentiment classification. Extensive experiments conducted on the CMU-MOSI and MELD benchmark datasets demonstrate that our framework consistently outperforms previous methods. In particular, the proposed model achieves 90.27% accuracy and 90.26% F1-score on CMU-MOSI, together with 66.57% accuracy and 66.21% F1-score on MELD. Additional ablation studies and qualitative analyses further validate the effectiveness of the proposed preprocessing strategy, modality-specific representation learning, and adaptive fusion mechanism. Full article
(This article belongs to the Special Issue Vision-Based AI in the Real World: Data, Robustness and Deployment)
Show Figures

Figure 1

13 pages, 25261 KB  
Data Descriptor
UAV-Hyp: UAV Dataset for Remote Sensing-Based Phenotyping of Hypericum perforatum
by Yujie Zhang, Ahmed EL Menuawy, Katrin Fitza, Frank Marthe and Sven Reichardt
Data 2026, 11(8), 208; https://doi.org/10.3390/data11080208 - 17 Aug 2026
Viewed by 2515
Abstract
Quantitative flowering phenotypes are needed to support breeding and harvest management in Hypericum perforatum L. (St. John’s wort), but manual flower assessment is slow and difficult to standardize under field conditions. UAV-Hyp is a multi-temporal UAV RGB dataset containing 12,653 high-resolution images acquired [...] Read more.
Quantitative flowering phenotypes are needed to support breeding and harvest management in Hypericum perforatum L. (St. John’s wort), but manual flower assessment is slow and difficult to standardize under field conditions. UAV-Hyp is a multi-temporal UAV RGB dataset containing 12,653 high-resolution images acquired at 26 measurement dates across the complete flowering period of 15 H. perforatum accessions. The images represent variable illumination, soil moisture, weed pressure, and developmental stages. The dataset provides 59,163 plant bounding boxes and 107,054 flower bounding boxes. As an application example, cascaded YOLOv8 plant and flower detectors achieved mAP@0.50:0.95 values of 0.977 and 0.950, respectively. UAV-Hyp supports scalable flower quantification and the development of time-series phenotyping methods for genotype comparison and quality-oriented medicinal-plant breeding. Full article
(This article belongs to the Section Spatial Data Science for Environment and Earth)
Show Figures

Figure 1

15 pages, 4123 KB  
Data Descriptor
A Device-Level IoT Network Traffic Dataset with Distributed Capture and Non-IID Characteristics
by Othmane Belarbi, Theodoros Spyridopoulos, Eirini Anthi, Omer Rana, Pietro Carnelli and Aftab Khan
Data 2026, 11(8), 207; https://doi.org/10.3390/data11080207 - 14 Aug 2026
Viewed by 230
Abstract
The development of intrusion detection and network security solutions for securing Internet of Things (IoT) networks is constrained by the limited availability of representative network security datasets. Many existing datasets rely on centralised traffic collection and do not capture the non-Independent and Identically [...] Read more.
The development of intrusion detection and network security solutions for securing Internet of Things (IoT) networks is constrained by the limited availability of representative network security datasets. Many existing datasets rely on centralised traffic collection and do not capture the non-Independent and Identically Distributed (non-IID) characteristics inherent to edge environments. To address this limitation, this work presents a device-level IoT network dataset generated using the open-source Gotham testbed, a virtualised smart city environment. Network traffic is collected in a distributed manner at the interfaces of 78 heterogeneous IoT devices operating across multiple protocols, including MQTT, CoAP, and RTSP. The dataset comprises over 31.8 million packet-level records, each described by 22 features. It includes both benign traffic and multiple attack classes, namely Network Scanning, Brute Force, Infection, Denial of Service (DoS), and Command and Control (C&C) Communication. Ground-truth labels are assigned using a deterministic process based on orchestration logs. The dataset preserves device-level traffic distributions and captures non-IID characteristics without artificial partitioning. It is publicly available and can be used to support reproducible evaluation of intrusion detection approaches and network analysis tasks in both centralised and distributed learning settings. Full article
(This article belongs to the Section Information Systems and Data Management)
Show Figures

Figure 1

9 pages, 674 KB  
Data Descriptor
A Multi-Metric NDVI-Derived Dataset of Vegetation Dynamics and Land Surface Phenology in Southern and Central Europe (1982–2022)
by Caterina Samela, Maria Lanfredi, Rosa Coluzzi and Vito Imbrenda
Data 2026, 11(8), 206; https://doi.org/10.3390/data11080206 - 11 Aug 2026
Viewed by 252
Abstract
This Data Descriptor presents a value-added suite of derived vegetation phenology products for Southern and Central Europe (10° W–28° E; 35° N–50° N), generated from the PKU GIMMS NDVI v1.2 archive (1982–2022) through a standardized processing workflow including quality screening, temporal compositing, phenological [...] Read more.
This Data Descriptor presents a value-added suite of derived vegetation phenology products for Southern and Central Europe (10° W–28° E; 35° N–50° N), generated from the PKU GIMMS NDVI v1.2 archive (1982–2022) through a standardized processing workflow including quality screening, temporal compositing, phenological metric extraction, eco-phenological regionalization, and variability analysis. The final collection includes 82 GeoTIFF raster layers and one CSV file, structured into long-term monthly NDVI climatologies, decadal NDVI composites, eco-phenological regionalization, Land Surface Phenology (LSP) metrics with inter-decadal shift layers, and a Phenology Variability Index (PVI). In addition, cluster-level Mann–Kendall and Theil–Sen trend statistics are provided in tabular form. All products are distributed in WGS84 (EPSG:4326) at 0.0833° spatial resolution. The dataset provides ready-to-use vegetation phenology indicators, variability metrics, and spatially consistent climatological products, supporting applications in ecosystem monitoring, climate impact assessment, biodiversity studies, and large-scale environmental analysis across European bioclimatic regions. Full article
(This article belongs to the Section Spatial Data Science for Environment and Earth)
Show Figures

Figure 1

26 pages, 1860 KB  
Data Descriptor
Topoclimatic Graph Dataset for Frost Prediction in the Tropical High-Mountain Altiplano Cundiboyacense, Colombia
by Evelin Calderón Caro, Dario Antonio Castañeda Sánchez, John R. Ballesteros and John W. Branch-Bedoya
Data 2026, 11(8), 205; https://doi.org/10.3390/data11080205 - 11 Aug 2026
Viewed by 296
Abstract
Frost prediction in tropical high-mountain agricultural regions is difficult because sparse meteorological networks must represent strong terrain-driven microclimatic variability. This article presents a topoclimatic graph dataset for frost prediction in the Altiplano Cundiboyacense, Colombia. The released core dataset contains 23 agricultural weather stations [...] Read more.
Frost prediction in tropical high-mountain agricultural regions is difficult because sparse meteorological networks must represent strong terrain-driven microclimatic variability. This article presents a topoclimatic graph dataset for frost prediction in the Altiplano Cundiboyacense, Colombia. The released core dataset contains 23 agricultural weather stations and is seasonally focused on recurrent November–February frost periods rather than year-round continuous monitoring. It includes four consistently available meteorological variables at 30 min resolution: air temperature, relative humidity, dew point temperature, and solar radiation. The data were consolidated from multiple operational sources, harmonized to a common temporal grid, subjected to physical and consistency-based quality control, and completed through temporal and spatial reconstruction with traceability labels. The final release also provides binary frost_event and frost_warning_6h labels, point-based topographic descriptors, 1 km buffer-based raster summaries, land-cover proportions, station-level static feature vectors, and graph products including edge lists and adjacency matrices. These data products support graph-based deep learning, multimodal spatiotemporal analysis, and frost early warning experiments in a tropical mountain agroecosystem. The dataset offers a reproducible framework for integrating heterogeneous environmental observations into graph-ready representations while preserving sufficient environmental context for benchmarking frost prediction methods in data-sparse regions. Full article
Show Figures

Graphical abstract

17 pages, 3333 KB  
Data Descriptor
Vibration Dataset for Crack Analysis and Detection in a Rotating Bladed System
by Adolfo Salgado-Ancona, José Billerman Robles-Ocampo, Edwin Neptalí Hernández-Estrada, Andrés López-López, Juvenal Rodríguez-Resendíz and Perla Yazmín Sevilla-Camacho
Data 2026, 11(8), 204; https://doi.org/10.3390/data11080204 - 10 Aug 2026
Viewed by 240
Abstract
This work presents conditioned and normalized vibration signal datasets acquired from the spanwise axis of the three blades of a rotating bladed system operating at a constant rotational speed of 240 rpm. The conditioned dataset was obtained using piezoelectric accelerometers mounted at the [...] Read more.
This work presents conditioned and normalized vibration signal datasets acquired from the spanwise axis of the three blades of a rotating bladed system operating at a constant rotational speed of 240 rpm. The conditioned dataset was obtained using piezoelectric accelerometers mounted at the blade roots. The accelerometer output signals were conditioned and recorded by a dedicated data acquisition system. The signals were acquired under both healthy and damaged operating conditions. Baseline vibration signals were first recorded with all three blades in a healthy condition. Subsequently, cracks were deliberately introduced at three different locations along the blade span—the root, middle, and tip zones. Each crack location was independently evaluated on each of the three blades, resulting in a comprehensive dataset that includes healthy operation and all combinations of blade–damage locations. The datasets enable analysis of the system’s vibratory response and of dynamic information propagation toward the blade root, depending on the crack zone. Their main contribution is to provide reliable experimental data for the development, validation, and benchmarking of vibration-based diagnostic and structural health monitoring techniques. Furthermore, the datasets serve as valuable resources for advancing early crack detection strategies and enhancing the reliability of rotating industrial equipment with blades, such as fans, compressors, turbines, and aerogenerators. Full article
Show Figures

Figure 1

14 pages, 8044 KB  
Data Descriptor
Global Metadata of the Influence of Cover Crops on Key Soil Hydraulic Properties
by Sabin Shrestha, Puja Sapkota, Bharat Sharma Acharya, Jason de Koff, Bharat Pokharel and Resham Thapa
Data 2026, 11(8), 203; https://doi.org/10.3390/data11080203 - 7 Aug 2026
Viewed by 302
Abstract
We present a global metadata comprising results from studies investigating the effects of cover crops (CCs) on six key soil hydraulic properties, namely total porosity, infiltration rate, saturated hydraulic conductivity, water retention at field capacity and permanent wilting points, and available water holding [...] Read more.
We present a global metadata comprising results from studies investigating the effects of cover crops (CCs) on six key soil hydraulic properties, namely total porosity, infiltration rate, saturated hydraulic conductivity, water retention at field capacity and permanent wilting points, and available water holding capacity. This data repository is the result of a global meta-analysis entitled “Cover Crop Performance and Functional Groups Regulate Improvements in Soil Hydrology: A Global Meta-analysis”. Globally, numerous studies have investigated the role of CCs on soil hydraulic properties, but the results have varied across sites and years. Hence, the objective of the meta-analysis was to synthesize the existing knowledge base to assess the overall effects of CCs on these soil hydraulic properties and evaluate how environmental and management factors moderate these overall CC responses. We searched for peer-reviewed research articles published through 5 October 2024 in the ISI Web of Science database, with reference checking following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines. A total of 146 relevant articles were identified from which data on CC responses were extracted. The metadata consists of 1007 pairwise observations comparing CC vs. no-CC controls across diverse geographic regions worldwide. Moreover, we collected associated metadata for each pairwise comparison that includes a broad set of bibliographic, geographic, soil, climate, and management variables. Categorical variables were grouped into pre-defined factor levels or classes. Missing soil and climate data were filled using publicly available data products. Our data repository can be a valuable resource for the field and modeling community to identify knowledge gaps and guide future research. Full article
(This article belongs to the Section Spatial Data Science for Environment and Earth)
Show Figures

Figure 1

24 pages, 1085 KB  
Data Descriptor
MUTra-CDMX: Multisource Urban Traffic Dataset for the Insurgentes Sur Corridor in Mexico City
by Arturo Rodríguez-Roman, Alicia Martínez-Rebollar, Hugo Estrada Esquivel, Ernesto de la Cruz-Nicolás and Eddie Clemente
Data 2026, 11(8), 202; https://doi.org/10.3390/data11080202 - 6 Aug 2026
Viewed by 251
Abstract
The growing complexity of urban mobility requires datasets that integrate dynamic traffic observations with meteorological, geometric, and urban-context information. This study presents MUTra-CDMX, a multisource urban traffic dataset covering a 14.72 km section of the Insurgentes Sur corridor in Mexico City. Traffic data [...] Read more.
The growing complexity of urban mobility requires datasets that integrate dynamic traffic observations with meteorological, geometric, and urban-context information. This study presents MUTra-CDMX, a multisource urban traffic dataset covering a 14.72 km section of the Insurgentes Sur corridor in Mexico City. Traffic data were obtained from TomTom at five-minute intervals for 20 consecutive road segments from 1 November 2024 to 28 February 2025. Hourly meteorological data were retrieved from Meteosource, while segment-level geometry, topology, signalized locations, and nearby points of interest were derived from TomTom metadata and OpenStreetMap. The primary analytical file contains 691,200 segment–timestamp records and 12 variables describing traffic and free-flow conditions, meteorological information, derived operational indicators, and reconstruction status. Of these records, 682,264 are original observations and 8936 are reconstructed segment–timestamp combinations, identified by the Boolean variable is_imputed. Technical validation confirmed complete temporal coverage, preservation of original traffic observations, consistent weather alignment, and reconstruction performance through artificial masking. Predictive utility was evaluated through chronological travel-time forecasting under a leakage-controlled protocol. At the 30 min horizon, XGBoost achieved a mean absolute error of 12.84 s, a root mean squared error of 37.91 s, and a coefficient of determination (R2) of 0.771, outperforming a persistence baseline. MUTra-CDMX supports congestion analysis, imputation studies, spatiotemporal modeling, and travel-time forecasting. Full article
(This article belongs to the Section Spatial Data Science for Environment and Earth)
Show Figures

Figure 1

31 pages, 3233 KB  
Article
Mapping Data-Driven Governance in Sharing Economy Platforms: Algorithmic Management, Platform Control, and Value-Creation Mechanisms
by Maria-Francisca Blasco-Lopez, Ramón Alberto Carrasco and Sulaiman Krayem
Data 2026, 11(8), 201; https://doi.org/10.3390/data11080201 - 6 Aug 2026
Viewed by 274
Abstract
Research on sharing economy platforms has expanded rapidly, yet the literature remains fragmented across studies on platform business models, gig work, algorithmic management, trust, reputation systems, artificial intelligence, and data-driven value creation. This article addresses this fragmentation through a bibliometric and systematic review [...] Read more.
Research on sharing economy platforms has expanded rapidly, yet the literature remains fragmented across studies on platform business models, gig work, algorithmic management, trust, reputation systems, artificial intelligence, and data-driven value creation. This article addresses this fragmentation through a bibliometric and systematic review of 660 documents retrieved from Scopus and Web of Science covering the period from 2010 to May 2026. A PRISMA-based protocol guided identification, deduplication, screening, eligibility assessment, and final corpus construction. The analysis combined performance indicators, co-citation analysis, keyword co-occurrence mapping, country collaboration analysis, longitudinal thematic evolution, strategic diagrams, and systematic content coding using Bibliometrix/Biblioshiny 5.4.1, VOSviewer 1.6.21, and SciMAT 1.1.04. The results show a marked acceleration of the field after 2020 and identify major research clusters around algorithmic labour and platform control, algorithmic management, trust and reputation, and dynamic pricing. The systematic coding further indicates that algorithmic management, reputation systems, dynamic pricing, surveillance, matching, and AI-enabled mechanisms recur across governance and value-creation processes. The study develops an integrative framework that interprets these patterns through four connected elements: data inputs, algorithmic mechanisms, governance functions, and value outcomes. This framework provides managers and regulators with a basis for assessing transparency, accountability, participant autonomy, value distribution, and the legitimacy of platform governance. Full article
(This article belongs to the Section Information Systems and Data Management)
Show Figures

Figure 1

25 pages, 3791 KB  
Article
Machine Learning in FinTech for Financial Fraud Data Detection
by Sanjaikanth E. Vadakkethil Somanathan Pillai and Wen-Chen Hu
Data 2026, 11(8), 200; https://doi.org/10.3390/data11080200 - 6 Aug 2026
Viewed by 370
Abstract
Financial fraud keeps rising these days. Organizations attempt to stop this trend by using various methods, such as distributing guides on how to avoid scams and frauds and automatically generating alerts when suspicious activities occur. However, this passive approach does not mitigate the [...] Read more.
Financial fraud keeps rising these days. Organizations attempt to stop this trend by using various methods, such as distributing guides on how to avoid scams and frauds and automatically generating alerts when suspicious activities occur. However, this passive approach does not mitigate the problem, as the trend is worsening, and it is usually too late when victims realize they have been scammed. Therefore, active approaches must be employed before scams reach the victims. A wide variety of preventive methods, such as neural networks and data mining, have been used to detect financial fraud data, but none have proven entirely effective in combating scams. Each method has its pros and cons. This research takes advantage of multiple machine learning techniques, such as k-nearest neighbors (kNN) and decision trees, by utilizing data fusion to detect financial fraud accurately. The data fusion function used here is self-adjusting through learning. During the training phase, the system is repeatedly applied to the dataset until an optimal detection rate is achieved. Experimental results from credit card transactions show that the proposed method outperforms each individual method. Parameter or threshold values for the data fusion are set heuristically. Future research will focus on developing reconfigurable data fusion by automatically adjusting the values. Full article
(This article belongs to the Special Issue Artificial Intelligence and Data Science for Fintech)
Show Figures

Figure 1

31 pages, 2093 KB  
Article
A Data-Centric Network Traffic Dataset for Anomaly Detection: Construction, Reproducible Pipeline, and Technical Validation
by Daniel Quirumbay Yagual, Diego Fernández Iglesias, Francisco J. Nóvoa and Daniel Garabato
Data 2026, 11(8), 199; https://doi.org/10.3390/data11080199 - 6 Aug 2026
Viewed by 309
Abstract
The effectiveness of machine learning and deep learning methods for network anomaly detection depends strongly on the quality and representativeness of the datasets used for training and evaluation. Despite recent advances, many publicly available benchmarks rely on synthetic traffic, outdated attack scenarios, or [...] Read more.
The effectiveness of machine learning and deep learning methods for network anomaly detection depends strongly on the quality and representativeness of the datasets used for training and evaluation. Despite recent advances, many publicly available benchmarks rely on synthetic traffic, outdated attack scenarios, or limited representation of encrypted communications. This work presents a network traffic dataset derived from operational firewall logs collected in a heterogeneous institutional environment dominated by HTTPS/TLS traffic. A structured data-centric pipeline was implemented, including preprocessing, behavioral feature engineering, unsupervised pseudo-labeling through the EFMS–KMeans algorithm, class balancing using SMOTE, and the generation of model-oriented sequential representations for deep learning analysis. The resulting dataset contains large-scale flow-level records describing volumetric, behavioral, and temporal traffic characteristics while preserving privacy through anonymization procedures. Technical validation was conducted using statistical analysis, entropy-based measurements, clustering quality metrics, and dimensionality reduction techniques, confirming data consistency, structural diversity, and class separability. The dataset is publicly available through the Mendeley Data repository together with metadata and documentation supporting anomaly detection research, encrypted traffic analysis, and the evaluation of machine learning and deep learning approaches in realistic cybersecurity environments. Full article
(This article belongs to the Topic Data Stream Mining and Processing)
Show Figures

Graphical abstract

Back to TopTop