electronics-logo

Journal Browser

Journal Browser

Data-Centric Artificial Intelligence: New Methods for Data Processing, 2nd Edition

A Special Issue of Electronics (ISSN 2079-9292) belonging to the section "Artificial Intelligence".

Deadline for manuscript submissions: 15 October 2026 | Viewed by 9569

Editor


E-Mail Website
Guest Editor
Department of Intelligent Systems, Faculty of Telecommunications, Computer Science and Electrical Engineering, Bydgoszcz University of Science and Technology, 85-796 Bydgoszcz, Poland
Interests: bee algorithms; fuzzy logic; artificial neural networks and their applications; language models; generative AI
Special Issues, Collections and Topics in MDPI journals

Special Issue Information

Dear Colleagues,

Data-centric artificial intelligence is developing rapidly thanks to advances in machine learning, natural language processing, and data visualization. These modern AI techniques enable a better understanding and processing of huge datasets. They provide companies and scientists with tools for extracting hidden patterns, discovering new knowledge, and automating complex analytical processes. In this Special Issue, we present examples of applications of these AI methods for solving real business and scientific problems.

We would like to invite you to submit a paper to our Special Issue of Electronics dedicated to data-centric artificial intelligence. This Special Issue will focus on the following topics:

  1. New methods and techniques for processing large datasets;
  2. Topics related to machine learning, natural language processing, and data visualization;
  3. Presenting practical applications of these methods in various fields.

This Special Issue will supplement the existing literature by focusing on the latest trends and solutions in this area.

Dr. Dawid Ewald
Guest Editor

Manuscript Submission Information

Manuscripts should be submitted online at www.mdpi.com by registering and logging in to this website. Once you are registered, click here to go to the submission form. Manuscripts can be submitted until the deadline. All submissions that pass pre-check are peer-reviewed. Accepted papers will be published continuously in the journal (as soon as accepted) and will be listed together on the special issue website. Research articles, review articles as well as short communications are invited. For planned papers, a title and short abstract (about 250 words) can be sent to the Editorial Office for assessment.

Submitted manuscripts should not have been published previously, nor be under consideration for publication elsewhere (except conference proceedings papers). All manuscripts are thoroughly refereed through a single-anonymized peer-review process. A guide for authors and other relevant information for submission of manuscripts is available on the Instructions for Authors page. Electronics is an international peer-reviewed open access semimonthly journal published by MDPI.

Please visit the Instructions for Authors page before submitting a manuscript. The Article Processing Charge (APC) for publication in this open access journal is 2400 CHF (Swiss Francs). Submitted papers should be well formatted and use good English. Authors may use MDPI's English editing service prior to publication or during author revisions.

Keywords

  • artificial intelligence
  • machine learning
  • data processing
  • data visualization
  • natural language processing
  • fuzzy logic

Benefits of Publishing in a Special Issue

  • Ease of navigation: Grouping papers by topic helps scholars navigate broad scope journals more efficiently.
  • Greater discoverability: Special Issues support the reach and impact of scientific research. Articles in Special Issues are more discoverable and cited more frequently.
  • Expansion of research network: Special Issues facilitate connections among authors, fostering scientific collaborations.
  • External promotion: Articles in Special Issues are often promoted through the journal's social media, increasing their visibility.
  • Reprint: MDPI Books provides the opportunity to republish successful Special Issues in book format, both online and in print.

Further information on MDPI's Special Issue policies can be found here.

Related Special Issue

Published Papers (5 papers)

Order results
Result details
Select all
Export citation of selected articles as:

Research

23 pages, 3390 KB  
Article
Quantifying the Incremental Value of XRF Elemental Data for Well-Log-Based TOC Prediction: A Fusion Strategy Evaluated Under Strict Cross-Validation
by Yang-Yang Zhong, Yun-Long Dai, Chao Li and Changjun Zhou
Electronics 2026, 15(18), 4075; https://doi.org/10.3390/electronics15184075 - 9 Sep 2026
Viewed by 197
Abstract
Total organic carbon (TOC) is a key parameter for source-rock evaluation, but laboratory measurements are sparse and do not resolve continuous vertical variation. Conventional well logs provide continuous physical responses, whereas X-ray fluorescence (XRF) data add geochemical information. We evaluated the incremental value [...] Read more.
Total organic carbon (TOC) is a key parameter for source-rock evaluation, but laboratory measurements are sparse and do not resolve continuous vertical variation. Conventional well logs provide continuous physical responses, whereas X-ray fluorescence (XRF) data add geochemical information. We evaluated the incremental value of combining these data sources using 58 depth-matched samples from the Yiwan-1 well in the Wangjiawan area, South China. Four well-log variables and six XRF variables selected within each training fold were evaluated with nine models under repeated 5-fold cross-validation (20 repeats). The fused SVR model performed best (RMSE = 0.789 ± 0.267; R = 0.912 ± 0.081). At the repetition level, the absolute RMSE reduction relative to logs alone was 0.430 (95% CI: 0.386–0.478), equivalent to 35.3% (Holm-adjusted p = 1.34 × 10−5). A 500-run permutation test that shuffled complete XRF sample rows gave an empirical p-value of 0.002. In a separate contiguous depth-block analysis with a 0.16 m exclusion buffer, the SVR RMSE decreased from 2.566 to 1.837, corresponding to an improvement of 28.4%, although the benefit was not retained by every linear model. Mo, V, Cr, Cd and Zn were selected in all Pearson-selection folds and had positive mean held-out permutation importance. The results support a complementary contribution from XRF data within this well while also showing that estimated accuracy depends on the validation design. Full article
Show Figures

Figure 1

23 pages, 4067 KB  
Article
Interpretable Machine Learning Models Using SHAP for Hourly and Daily-Maximum Carbon Monoxide Forecasting at Urban Air Quality IoT Monitoring Stations in Greece
by Yiannis Kiouvrekis, Christos Christakis, Ioannis Tsilikas and Theodor Panagiotakopoulos
Electronics 2026, 15(15), 3371; https://doi.org/10.3390/electronics15153371 - 31 Jul 2026
Viewed by 374
Abstract
Carbon monoxide (CO) remains an important urban air quality and traffic-exposure tracer despite rarely exceeding regulatory limits in modern European cities. This study presents an interpretable machine learning framework for short-term CO forecasting at two operational horizons—next hour and next-day maximum—applied identically to [...] Read more.
Carbon monoxide (CO) remains an important urban air quality and traffic-exposure tracer despite rarely exceeding regulatory limits in modern European cities. This study presents an interpretable machine learning framework for short-term CO forecasting at two operational horizons—next hour and next-day maximum—applied identically to four automated monitoring stations spanning central-urban, urban, urban-background, and suburban typologies in the Greater Athens Area (2021–2024). Using only univariate CO history and calendar-derived features, four learners (support vector regression, random forest, gradient boosting, and a multilayer perceptron) were benchmarked against a naïve baseline under a chronological train/validation/test split. At the next-hour horizon, all learners outperformed naïve at every station, with gradient boosting achieving the best or joint-best skill (test R2=0.810.90); Wilcoxon signed-rank tests confirmed that these small but consistent margins were statistically significant. The next-day-maximum task proved substantially harder (R2=0.440.59), with the neural and random forest models overtaking gradient boosting and SVR losing competitiveness. A SHAP (Shapley Additive Explanations) analysis of the next-hour gradient-boosting model showed that the most recent hourly lag dominates the forecast, with an effect nearly an order of magnitude over any other predictor, with hour-of-day encoding and short-lag rolling statistics contributing secondary, sign-consistent effects—providing a transparent, mechanistic account of model behavior rather than a black-box skill score. Unlike ozone, a secondary pollutant whose predictability degrades toward the trafficked urban core, CO concentration and forecastability increase together at the traffic-dominated site, indicating that primary-pollutant forecasts are most reliable precisely where exposure is greatest. These findings support a horizon-specific, interpretable forecasting strategy for operational deployment on real-time monitoring networks. Full article
Show Figures

Figure 1

18 pages, 12250 KB  
Article
A Vision Transformer Model with Hyperparameter Optimization for Oral Cancer Image Classification
by Chun-Tai Huang, Ying-Lei Lin, Chung-Hui Lin and Ping-Feng Pai
Electronics 2026, 15(10), 2230; https://doi.org/10.3390/electronics15102230 - 21 May 2026
Viewed by 605
Abstract
Oral cancer is a significant public health concern and is among the most common malignant tumors of the head and neck. Its incidence and mortality rates remain persistently high, especially in regions where smoking and betel nut chewing are prevalent. Due to its [...] Read more.
Oral cancer is a significant public health concern and is among the most common malignant tumors of the head and neck. Its incidence and mortality rates remain persistently high, especially in regions where smoking and betel nut chewing are prevalent. Due to its high mortality rate, early detection is crucial for improving patient outcomes. However, early symptoms of oral cancer often resemble benign oral lesions, leading to delayed diagnosis. In this study, a vision transformer (ViT) model with Optuna (ViTOPT) is employed to perform classification tasks of identifying oral cancer images. The Optuna is used to determine hyperparameters in ViT. Histological images are obtained from a publicly available dataset. Three classification tasks with histological images namely classifying oral squamous cell carcinoma (OSCC) and leukoplakia (LEUK), classifying the presence of dysplasia, and classifying OSCC and leukoplakia with or without dysplasia are performed in this study. Numerical results reveal that the proposed ViTOPT framework is able to provide satisfactory performance in oral cancer recognition. Thus, the proposed ViTOPT model is a feasible and effective alternative in identifying oral cancer. Full article
Show Figures

Figure 1

37 pages, 883 KB  
Article
Data-Centric AI Manifesto: How Data Quality Drives Modern AI
by Donato Malerba, Antonella Poggi, Mario Alviano, Tommaso Boccali, Maria Teresa Camerlingo, Roberto Maria Delfino, Domenico Diacono, Domenico Elia, Vincenzo Pasquadibisceglie, Mara Sangiovanni, Vincenzo Spinoso and Gioacchino Vino
Electronics 2026, 15(9), 1913; https://doi.org/10.3390/electronics15091913 - 1 May 2026
Viewed by 3096
Abstract
Artificial Intelligence (AI) has traditionally been developed according to a model-centric paradigm, in which progress is driven by increasingly sophisticated learning architectures applied to largely fixed datasets. However, this paradigm exhibits well-known limitations, including sensitivity to label noise, distribution shifts, adversarial perturbations, and [...] Read more.
Artificial Intelligence (AI) has traditionally been developed according to a model-centric paradigm, in which progress is driven by increasingly sophisticated learning architectures applied to largely fixed datasets. However, this paradigm exhibits well-known limitations, including sensitivity to label noise, distribution shifts, adversarial perturbations, and limited transparency and reproducibility. These issues indicate that many of the current bottlenecks of AI systems arise from deficiencies in data rather than from model design. In this paper, we adopt and formalize the Data-Centric Artificial Intelligence (DCAI) paradigm, which places data quality, semantic consistency, and representativeness at the core of the AI lifecycle. From this perspective, performance, robustness, interpretability, and regulatory compliance are primarily achieved through systematic data engineering, including data curation, enrichment, validation, and continuous monitoring, rather than through repeated model re-engineering. The contributions of this work are threefold. First, a conceptual framework is provided to clarify the epistemic and methodological foundations of DCAI and distinguish it from traditional model-centric approaches. Second, a data-centric lifecycle is presented, covering training data development, inference data design, and data maintenance and integrating techniques such as semantic data representation, active learning, synthetic data generation, and drift-aware quality control. Third, the role of DCAI in the context of Generative AI is analyzed, showing how data-centric practices are essential to ensure robustness, accountability, and responsible deployment of large-scale generative models. Overall, this work positions DCAI as a coherent methodological and technological framework for the development of trustworthy, resilient, and sustainable AI systems, making a research contribution and providing a reference model for industrial and regulatory contexts. Full article
Show Figures

Figure 1

28 pages, 4737 KB  
Article
Comparative Evaluation of Perceptual Hashing and Deep Embedding Methods for Robust and Efficient Image Deduplication
by Md Firoz Mahmud, Zerin Nusrat and W. David Pan
Electronics 2026, 15(7), 1493; https://doi.org/10.3390/electronics15071493 - 2 Apr 2026
Viewed by 4729
Abstract
The rapid growth in large-scale image repositories over the past few years has made exact and near-duplicate images increasingly common, creating substantial redundancy that wastes storage resources and reduces retrieval efficiency in practical systems. Even though perceptual hashing and deep learning are promising [...] Read more.
The rapid growth in large-scale image repositories over the past few years has made exact and near-duplicate images increasingly common, creating substantial redundancy that wastes storage resources and reduces retrieval efficiency in practical systems. Even though perceptual hashing and deep learning are promising deduplication strategies, the lack of standardized benchmarks complicates direct comparison. In this study, we conduct a unified, controlled evaluation of five commonly used methods, including four classical perceptual hashes (AHash, DHash, PHash, and WHash) and a CNN-based embedding model. We evaluate all methods on the UKBench and Amazon Berkeley Objects datasets using identical preprocessing, thresholds, and metrics, which include exact duplicates, near-duplicates, and geometrically transformed duplicates. Our experiments highlight a clear trade-off between speed and robustness. Hashing methods are computationally efficient and effective for exact matches, but perform poorly on near-duplicates and under geometric transformations, whereas the CNN model is significantly more robust across all duplicate types, but comes at a high computational cost. Based on these results, we outline practical recommendations for selecting deduplication strategies in large-scale applications. In addition, our evaluation setup serves as a reproducible baseline for future research in image similarity and large-scale deduplication. Full article
Show Figures

Figure 1

Back to TopTop