Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (627)

Search Parameters:
Keywords = FAIR datasets

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
29 pages, 3461 KB  
Article
Benchmarking Class Imbalance Mitigation Strategies Across Deep CNN Architectures for Skin Cancer Classification
by Irshad Ahmad, Muhammad Khubaib and Saleh M. Altowaijri
Diagnostics 2026, 16(16), 2571; https://doi.org/10.3390/diagnostics16162571 - 14 Aug 2026
Abstract
Background/Objectives: Class imbalance is one of the major challenges in automated skin lesion classification since the number of categories of malignant and clinically significant skin lesions is normally less than the benign ones. However, due to this imbalance, deep convolutional neural networks [...] Read more.
Background/Objectives: Class imbalance is one of the major challenges in automated skin lesion classification since the number of categories of malignant and clinically significant skin lesions is normally less than the benign ones. However, due to this imbalance, deep convolutional neural networks (CNNs) tend to overlook minority classes and fail to recognize them with an acceptable accuracy, which leads to a decrease in diagnostic reliability. A wide range of imbalance mitigation techniques has been suggested, but their effectiveness is found to differ significantly depending on CNN architecture, and detailed comparative studies of these techniques for a consistent experimental setup are still limited. Methods: This study proposes a comprehensive benchmarking framework that tests sixteen class imbalance mitigation methods by applying them to six pretrained CNN architectures—EfficientNet-B0, EfficientNet-B3, ResNet50, DenseNet121, InceptionV3 and MobileNetV2—on the official ISIC 2019 skin lesion dataset. The tested techniques are conventional resampling techniques, synthetic sample generation techniques, algorithm-level learning techniques, data augmentation techniques, and hybrid techniques. The dataset was partition into a separate training set and testing set, and stratified cross-validation was only conducted on the training set to ensure the study was fair and reproducible. Both models have been optimized with the same optimizer, learning rate, batch size, epochs and preprocessing pipeline. The performance of the models was evaluated by computing the accuracy, precision, recall and F1-score. Results: The experimental results show that the effect of class imbalance mitigation is very specific to the underlying CNN architecture. The traditional undersampling and oversampling methods yielded only moderate improvements, while feature space and hybrid methods yielded more consistent results. When coupled with EfficientNet-B3, Balanced MixUp improved the overall performance of the model by achieving an accuracy of 92.39%, an increase in precision of 93.3%, a recall of 91.36%, and an F1-score of 92.33%. However, some architectures such as ResNet50 performed better with iterative learning techniques, such as Cumulative Learning and Yielding Multi-Fold Training, which suggests that there is a diversity in how different network architectures react to imbalance mitigation methods. Conclusions: This paper highlights the importance of selecting appropriate technique–architecture combinations for addressing long-tailed data distributions in medical imaging. The proposed benchmarking framework provides valuable insights for developing robust and reliable deep learning systems for skin lesion classification and other medical imaging tasks affected by severe class imbalance. Full article
Show Figures

Figure 1

18 pages, 4050 KB  
Review
Algorithmic Prognostication in Female Oncofertility Counseling: Ethical Challenges of Bias, Autonomy, and Predictive Uncertainty
by Huei-Ying Chiu, Ya-Ting Chuang, Simona Zaami and Tao-An Chen
Healthcare 2026, 14(16), 2538; https://doi.org/10.3390/healthcare14162538 - 13 Aug 2026
Abstract
Advances in machine learning, predictive analytics, and clinical prediction modeling have accelerated the development of algorithmic tools for estimating reproductive outcomes after cancer treatment. In female oncofertility counseling, these models may support individualized assessment of treatment-related amenorrhea, premature ovarian insufficiency, and fertility risk, [...] Read more.
Advances in machine learning, predictive analytics, and clinical prediction modeling have accelerated the development of algorithmic tools for estimating reproductive outcomes after cancer treatment. In female oncofertility counseling, these models may support individualized assessment of treatment-related amenorrhea, premature ovarian insufficiency, and fertility risk, thereby improving risk communication and timely fertility-preservation referral. However, their use raises ethical concerns beyond predictive accuracy. This narrative review examines algorithmic prognostication in female oncofertility counseling, focusing on predictive uncertainty, surrogate reproductive endpoints, missing data, heterogeneous datasets, limited external validation, algorithmic bias, reproductive inequity, and the influence of algorithmic authority on patient autonomy and shared decision-making. We argue that predictive algorithms should be understood as decision-support tools rather than determinants of reproductive futures. Responsible implementation requires transparency, explainability, fairness assessment, ongoing validation, and meaningful human oversight. Algorithmic risk estimates should be communicated as conditional and contextual probabilities within patient-centered counseling, ensuring that predictive tools support informed, transparent, and value-concordant fertility-preservation decisions for women facing cancer treatment. Full article
Show Figures

Figure 1

39 pages, 804 KB  
Review
The Current Generation of Tabular Foundation Models: A Critical Review
by Sergei O. Kurashkin, Vadim S. Tynchenko, Aleksei S. Borodulin, Vladimir A. Nelyub, Nikolay O. Kalutsky and Tee Connie
Mach. Learn. Knowl. Extr. 2026, 8(8), 244; https://doi.org/10.3390/make8080244 - 13 Aug 2026
Abstract
Tabular foundation models (TFMs) have moved tabular machine learning from per-dataset training towards amortised in-context inference, fitting a small-to-medium table in a single forward pass without a training run. The 2024–2026 release train, the TabPFN and TabICL lines and challengers such as Mitra, [...] Read more.
Tabular foundation models (TFMs) have moved tabular machine learning from per-dataset training towards amortised in-context inference, fitting a small-to-medium table in a single forward pass without a training run. The 2024–2026 release train, the TabPFN and TabICL lines and challengers such as Mitra, LimiX and Orion, has produced a generation whose architectures, capabilities and limits are documented mainly in preprints, while existing surveys treat these models as a subsection of tabular deep learning or of language-model table understanding. This review is, to our knowledge, the first organised around the current generation. From a corpus of 961 screened records and 98 retained studies, it taxonomises the architectures by pretraining regime, maps the capability space across five axes, isolates the language-model-on-tabular strand for prediction, feature engineering and generation, and summarises openness and deployment. A dedicated critical synthesis then reads the reported capabilities against independent evidence: on the studies reviewed here, tree-based and deep models retain the lead across 142 curated datasets that go beyond the standard independent and identically distributed setting; on 112 datasets, the models attain the highest accuracy but weaker conditional coverage than gradient-boosted trees; and robustness under feature shift, fairness and generation quality remain open. Amortised in-context prediction is thus a working paradigm whose independent evidence has yet to match its benchmark claims. Full article
(This article belongs to the Section Thematic Reviews)
Show Figures

Figure 1

72 pages, 12680 KB  
Article
iCert-Fair: A Human-Preference-Guided Two-Layer Framework for Multi-Objective Fairness Assessment and Harm Recovery in Credit Scoring
by Rashed Bahlool and Nabil Hewahi
AI 2026, 7(8), 312; https://doi.org/10.3390/ai7080312 - 13 Aug 2026
Abstract
As regulatory requirements increasingly shape automated lending decisions, fairness remains a critical challenge in high-stakes domains, particularly credit scoring. Although artificial intelligence models can achieve strong predictive performance, they may also reproduce biased outcomes that reduce financial inclusion or transfer harm to overlooked [...] Read more.
As regulatory requirements increasingly shape automated lending decisions, fairness remains a critical challenge in high-stakes domains, particularly credit scoring. Although artificial intelligence models can achieve strong predictive performance, they may also reproduce biased outcomes that reduce financial inclusion or transfer harm to overlooked protected groups. Existing fairness interventions commonly operate at a single stage of the decision-making pipeline, despite bias often propagating across representational and decision layers. This study proposes iCert-Fair, a two-layer framework for technical fairness assessment and harm recovery in credit scoring. The first layer adopts a fairness-through-explainability paradigm, using SHAP-based explanations to identify direct and proxy dependence on protected attributes and guide structural dataset repair, while the second layer applies targeted threshold-policy adjustments to recover residual harm while preserving decision utility. Experiments on the German and Taiwanese credit datasets show that fairness gains are model- and dataset-specific and may be collective, concentrated, transferred, or recovered unevenly across protected attributes. The direct comparison with representative pre-processing, in-processing, and post-processing methods revealed that baseline methods targeting one protected attribute at a time frequently transferred residual harm to other monitored attributes. In contrast, the fairness-focused recommendations generated by iCert-Fair achieved larger collective fairness improvements across all considered protected attributes while avoiding residual harm. These gains were obtained while preserving predictive utility on the German dataset and with utility degradation remaining below 5% across the evaluated performance metrics on the Taiwanese dataset, alongside consistently lower false-negative risk. The empirical findings support the use of complementary structural and policy-level interventions and demonstrate the importance of jointly evaluating aggregate disparity, worst-case attribute-level harm, cross-attribute transfer, and predictive utility. Full article
(This article belongs to the Special Issue Human-Computer Interaction and Human-Centered AI)
32 pages, 10723 KB  
Article
Hybrid Recommendation System for the Marketing of Artisanal Agricultural Products in Dispersed Rural Markets
by Mary Elsy Arzuaga-Ochoa, Jorge Gómez, Melisa Acosta-Coll and Mauricio Barrios-Barrios
Appl. Sci. 2026, 16(16), 8074; https://doi.org/10.3390/app16168074 - 13 Aug 2026
Abstract
Artisanal farmers in developing regions capture limited commercial value due to information asymmetries, market fragmentation, and informal intermediation, where commercialization, rather than production, is the main bottleneck for rural income. This study formalizes artisanal agricultural commercialization as a buyer-recommendation problem in a sparse [...] Read more.
Artisanal farmers in developing regions capture limited commercial value due to information asymmetries, market fragmentation, and informal intermediation, where commercialization, rather than production, is the main bottleneck for rural income. This study formalizes artisanal agricultural commercialization as a buyer-recommendation problem in a sparse bilateral market. We propose a context-aware hybrid system that combines content-based and collaborative filtering with economic and geographic variables. Because no transactional dataset exists for the target region (Cesar, Colombia), a proof-of-concept evaluation (Phase A) uses the Brazilian Olist e-commerce dataset as a structural proxy, benchmarking 13 models under a temporal leave-one-out protocol. A Reciprocal Rank Fusion ensemble achieved the best ranking performance (NDCG@10 = 0.4031), an item-based neighborhood model had the lowest rating-prediction error, and content-based filtering proved to be the most informative component under extreme sparsity. Differences were confirmed using Friedman, Nemenyi, bootstrap, and effect size tests. These offline indicators show technical feasibility but do not demonstrate gains in producer income or field adoption rates. Real-world validation, contextual adaptation, and fairness assessment in rural Cesar are proposed for future work (Phase B). Full article
Show Figures

Figure 1

26 pages, 5467 KB  
Article
BC-XAIA: A Blockchain-Based Recruitment Framework with Explainable AI and Smart Contract Integration
by Hebat Allah Adel, Sayed AbdelGaber and Wessam H. El-Behaidy
Appl. Sci. 2026, 16(16), 8064; https://doi.org/10.3390/app16168064 - 13 Aug 2026
Viewed by 51
Abstract
Ensuring transparency and security in digital recruitment systems remains a critical challenge. This study proposes BC-XAIA, a unified framework that integrates blockchain, smart contracts, explainable artificial intelligence (XAI), and agile methodology to enable consistent, secure, and traceable recruitment decision-making. Smart contracts, implemented in [...] Read more.
Ensuring transparency and security in digital recruitment systems remains a critical challenge. This study proposes BC-XAIA, a unified framework that integrates blockchain, smart contracts, explainable artificial intelligence (XAI), and agile methodology to enable consistent, secure, and traceable recruitment decision-making. Smart contracts, implemented in Solidity and deployed using the Remix Ethereum IDE, automate key processes such as identity verification, data access control, and behavior monitoring, reducing reliance on centralized intermediaries. To support intelligent decision-making, multiple machine learning models, including Random Forest, Logistic Regression, and Support Vector Machine (SVM), were trained and evaluated on a recruitment dataset, with Random Forest achieving the highest performance, reaching an accuracy of 93%. To enhance transparency, SHAP and LIME were employed to provide both global and local interpretability of model predictions. Furthermore, agile methodology is embedded to drive continuous adaptation, iterative development, and stakeholder feedback throughout the recruitment lifecycle. Unlike existing recruitment systems that treat blockchain, AI, and explainability separately, BC-XAIA unifies these technologies within an agile and decentralized architecture. Overall, BC-XAIA establishes a secure, transparent, and explainable decentralized recruitment ecosystem that enhances trust, fairness, and intelligent decision-making in next-generation HR systems. Full article
Show Figures

Figure 1

28 pages, 2416 KB  
Article
Fairness Evaluation Paradox: How Biased Test Data Masks True Group Fairness Assessment
by Sašo Karakatič, Ivona Colakovic and Tjaša Heričko
Mathematics 2026, 14(16), 2894; https://doi.org/10.3390/math14162894 - 11 Aug 2026
Viewed by 159
Abstract
The EU AI Act makes fairness metrics for high-risk AI systems’ compliance evidence, turning their trustworthiness into a safety and accountability concern. Fairness audits assume that test data reflects the properties of real-world conditions, whereas standard evaluation protocols use test data drawn from [...] Read more.
The EU AI Act makes fairness metrics for high-risk AI systems’ compliance evidence, turning their trustworthiness into a safety and accountability concern. Fairness audits assume that test data reflects the properties of real-world conditions, whereas standard evaluation protocols use test data drawn from the same biased records as the training data. Studies measuring how strongly this bias in test data distorts fairness metrics between the validation phase and real-world deployment are still very rare. We conduct an experiment on synthetic and real data, measuring this discrepancy across five fairness interventions on four unfairness types (1000 repetitions per combination, 20,000 total runs). Synthetic data lets us encode human bias in labels and compare fairness metrics on biased test labels (data available during development) against clean labels (conditions models face in deployment). We find that a systematic evaluation bias is present across all metrics, so the same models on the same test data can support opposite fairness conclusions and mask the mistreatment of the most disadvantaged groups. This pattern of fairness misevaluation is confirmed by a real-world validation on the Adult Census Income dataset. We conclude that trustworthy fairness auditing and regulatory standards should require bias-aware evaluation protocols, in which observed labels are not treated as ground truth. Full article
Show Figures

Figure 1

46 pages, 19062 KB  
Article
A Physics-Informed Benchmarking Framework for Machine Learning and Tree-Based Ensembles in IIoT-Enabled Predictive Maintenance
by Yi-Kai Su and Chun-Jan Tseng
Sensors 2026, 26(16), 5026; https://doi.org/10.3390/s26165026 - 7 Aug 2026
Viewed by 365
Abstract
Reliable Predictive Maintenance (PdM) in Industrial Internet of Things (IIoT) environments is challenged by severe class imbalance, heterogeneous sensor variables, inconsistent experimental protocols, and deployment constraints. This study proposes a Physics-Informed Benchmarking Framework that integrates engineering-guided feature construction, Mutual Information (MI)-based feature relevance [...] Read more.
Reliable Predictive Maintenance (PdM) in Industrial Internet of Things (IIoT) environments is challenged by severe class imbalance, heterogeneous sensor variables, inconsistent experimental protocols, and deployment constraints. This study proposes a Physics-Informed Benchmarking Framework that integrates engineering-guided feature construction, Mutual Information (MI)-based feature relevance analysis, standardized model development, and deployment-oriented evaluation within a unified and reproducible workflow. Using the AI4I 2020 Predictive Maintenance Dataset, Logistic Regression, Isolation Forest, Random Forest, and Extreme Gradient Boosting (XGBoost) were evaluated using identical feature representations, train–test partitions, preprocessing procedures, and imbalance-handling strategies. The engineered feature space incorporates thermal, mechanical, interaction, and degradation-related information derived from the original sensor measurements. The results show that tree-based ensembles provide the strongest overall performance under severe class imbalance. Random Forest achieved an accuracy of 0.986, an F1-score of 0.722, and a ROC-AUC of 0.983, providing the best balance between failure detection and false-alarm control. XGBoost achieved an accuracy of 0.978, a recall of 0.853, and the lowest inference latency of 0.35 ms, indicating its suitability for latency-sensitive IIoT deployment. These findings demonstrate that combining engineering-guided feature representation with a standardized evaluation protocol enables fair comparison of representative learning paradigms while preserving engineering interpretability and deployment relevance. Full article
Show Figures

Figure 1

24 pages, 9984 KB  
Article
Shape-Prior Dynamic Refinement Network with Diffusion Wavelets for Fine-Grained Ship Detection in Remote Sensing Images
by Zhanchao Huang, Weiwang Guan, Jiajun Zhou, Wenjun Hong and Hua Su
Remote Sens. 2026, 18(16), 2649; https://doi.org/10.3390/rs18162649 - 7 Aug 2026
Viewed by 248
Abstract
Fine-grained ship detection in remote sensing images is important for maritime surveillance and port management. Although deep learning methods have achieved promising results, fine-grained ship detection remains difficult because of complex background interference and difficulty in distinguishing similar objects. To address these limitations, [...] Read more.
Fine-grained ship detection in remote sensing images is important for maritime surveillance and port management. Although deep learning methods have achieved promising results, fine-grained ship detection remains difficult because of complex background interference and difficulty in distinguishing similar objects. To address these limitations, we propose a Shape-Prior Dynamic Refinement Network (SPDR-Net) for fine-grained ship detection. First, we design a Gaussian Mixture Shape Representation (GMSR) mechanism to guide the learning of shapes and texture details by learnable Gaussian mixture models. Second, we develop Re-Parameterized Selective Dynamic Kernels (RSDKs) based on the GMSR shape prior to dynamically adjust the receptive field of each scale to more accurately capture morphological features. Furthermore, a Diffusion-Supervised Wavelet Refinement (DSWR) strategy is developed, which brings diffusion-based noising–denoising adversarial learning mechanism into wavelet reconstruction to recover high-frequency details, strengthening boundaries and fine-grained differential information while suppressing noise. Extensive experiments on public datasets demonstrate that the proposed method achieves state-of-the-art performance with mAP50s of 79.3% on DOSR, 47.97% on FAIR1M, 98.1% on HRSC2016, and 97.7% on SSDD, respectively. It exhibits that SPDR-Net effectively handles both offshore and complex nearshore scenarios, offering a robust solution for fine-grained ship detection. Full article
Show Figures

Figure 1

15 pages, 698 KB  
Article
Operationalization of Equity Through a Quantitative Analysis of Stakeholder Salience
by Isabel M. del Águila and José del Sagrado
Software 2026, 5(3), 34; https://doi.org/10.3390/software5030034 - 7 Aug 2026
Viewed by 124
Abstract
The incorporation of human values into software engineering processes is increasingly recognised as crucial, yet there is still a lack of practical methods to ensure their inclusion. A key challenge in this area is to identify a representative set of stakeholders while maintaining [...] Read more.
The incorporation of human values into software engineering processes is increasingly recognised as crucial, yet there is still a lack of practical methods to ensure their inclusion. A key challenge in this area is to identify a representative set of stakeholders while maintaining equality and diversity throughout the development process. This study aims to leverage stakeholder salience—defined by the attributes of power, legitimacy, and urgency—to partition stakeholders using a quantile-based statistical approach, establishing stakeholder groups that prevent discrimination and marginalization in requirements engineering processes. We propose a quantile-based statistical method for systematic stakeholder partitioning to ensure that diverse interests are considered and formalize equality and diversity through quantitative metrics in order to operationalize fairness. We applied the proposed method to the RALIC dataset to evaluate its applicability, considering two and three levels of granularity (binary medians and ternary terciles) for each salience component to assess its effectiveness in maintaining equality and diversity. Integrating quantile-based stakeholder salience partitioning into the identification process provides a structured approach to incorporating human values into software engineering, demonstrating that tercile partitioning significantly optimizes group balance and spatial coverage, ultimately contributing to more inclusive and equitable decision-making in requirements gathering. Full article
Show Figures

Figure 1

31 pages, 11789 KB  
Article
AFCANet: An Axis-Factorized Convolution–Attention Network for Portfolio-Level Customer Baseline Load Estimation
by Faraj H. Alyami, Sheeraz Iqbal, Md Shafiullah and Saleh Al Dawsari
Mathematics 2026, 14(15), 2860; https://doi.org/10.3390/math14152860 - 6 Aug 2026
Viewed by 295
Abstract
In incentive-based demand response, an aggregator is paid for the gap between a customer’s metered load and the baseline load that would have occurred without a curtailment signal. This baseline is never recorded during the event, yet the settlement depends on it, so [...] Read more.
In incentive-based demand response, an aggregator is paid for the gap between a customer’s metered load and the baseline load that would have occurred without a curtailment signal. This baseline is never recorded during the event, yet the settlement depends on it, so it must be reconstructed from the load observed before and after the curtailment window. At the portfolio level on which settlement is cleared, this amounts to filling a single contiguous gap, aligned with the daily peak, in an otherwise complete record. To estimate the portfolio-level customer baseline load (CBL), we propose AFCANet, which folds the one-dimensional CBL time series into a period-aligned two-dimensional tensor whose two axes describe different things. The intra-period axis traces the shape of a single daily cycle, which is locally smooth and strongly autocorrelated, while the inter-period axis links the same clock time across successive days, a longer-range and less locally smooth dependency. At the core of AFCANet is the Axis-Factorized Convolution–Attention (AFCA) block, which assigns a convolution to the intra-period axis, where its locality and weight-sharing suit the smooth daily shape, and self-attention to the inter-period axis, where its ability to link distant positions suits the cross-day dependency. Experiments use metered load from the Low Carbon London trial dataset with half-hourly resolution, evaluated under a control-group protocol in which the masked baseline is directly verifiable; AFCANet attains a MAE of 20.88 kWh, a MAPE of 1.50%, and a near-zero bias of 4.56 kWh, improving on averaging, regression, and learning-based imputation baselines. A controlled ablation that swaps the two operators confirms that the matched axis assignment is the source of the gain. Since the evaluation relies on synthetic curtailment windows in which no behavioral response is present, the reported accuracy should be read as an upper bound on the performance attainable in live demand-response events. The near-zero bias is of direct practical value to load aggregators, as a baseline free of systematic over- or under-estimation supports accurate curtailment measurement and fair financial settlement in incentive-based demand response. Full article
Show Figures

Figure 1

17 pages, 599 KB  
Article
Comparative Performance Evaluation of Six Federated Learning Frameworks Under Locked FedAvg: Native SDKs and a Shared Reference Harness for Edge-Oriented 6G Applications
by Vasileios D. Batsios and Constantinos T. Angelis
Future Internet 2026, 18(8), 416; https://doi.org/10.3390/fi18080416 - 6 Aug 2026
Viewed by 167
Abstract
Federated learning (FL) enables privacy-preserving collaborative training at the network edge, a core capability envisioned for sixth-generation (6G) wireless systems. While surveys and scale-oriented benchmarks advance FL methodology, documented, head-to-head comparisons of mainstream Python frameworks under identical FedAvg settings remain scarce. We benchmark [...] Read more.
Federated learning (FL) enables privacy-preserving collaborative training at the network edge, a core capability envisioned for sixth-generation (6G) wireless systems. While surveys and scale-oriented benchmarks advance FL methodology, documented, head-to-head comparisons of mainstream Python frameworks under identical FedAvg settings remain scarce. We benchmark six frameworks—Flower, TensorFlow Federated (TFF), FedML, NVIDIA FLARE, OpenFL, and PySyft—distinguishing two native SDK integrations (Flower, TFF) from four runs of a shared PyTorch FedAvg reference harness (FedML, NVIDIA FLARE, OpenFL, PySyft) in a controlled two-phase study on a Proxmox virtualized testbed with containerized runners, formalize the FedAvg objective and communication-cost model, and position our contribution against prior surveys, scale benchmarks, and single-framework documentation. Each framework–dataset pair is repeated over five IID partitions (random seeds 42–46); we report round-10 mean ± standard deviation for accuracy, wall time, and simulated communication volume. Phase 1 (MNIST) confirms protocol fairness (99.22±0.0799.29±0.06% accuracy) with moderate wall-time spread; Phase 2 (CIFAR-10) exposes stack-dependent accuracy gaps (TFF 71.16±0.23% vs. ≈68% for PyTorch runners). We report per-round accuracy and loss curves with variability bands, wall-time comparisons, and simulated parameter traffic for all six frameworks across nine figures. The experimental protocol, model topology, and hyperparameters are specified in full; per-round JSON metrics and global model checkpoints are published. The study provides a documented baseline for 6G edge framework selection and for follow-on network-constrained and security experiments. Full article
(This article belongs to the Special Issue 5G/6G and Beyond: The Future of Wireless Communications Systems)
Show Figures

Figure 1

22 pages, 1009 KB  
Article
The Explanation Cost of Fairness: How Bias Mitigation Affects Explanation Faithfulness in Machine Learning Intrusion Detection
by Khalid Alalawi
Electronics 2026, 15(15), 3457; https://doi.org/10.3390/electronics15153457 - 5 Aug 2026
Viewed by 227
Abstract
Machine learning intrusion detectors are accurate but opaque, so explainable AI justifies their alerts and bias mitigation evens out detection across imbalanced categories. These are usually studied independently, and whether making a detector fairer changes how faithfully it can be explained is unknown. [...] Read more.
Machine learning intrusion detectors are accurate but opaque, so explainable AI justifies their alerts and bias mitigation evens out detection across imbalanced categories. These are usually studied independently, and whether making a detector fairer changes how faithfully it can be explained is unknown. In an interventional design, we train an unmitigated baseline and three imbalance mitigations on a one-dimensional CNN, an FT-Transformer, and XGBoost, applying class weighting and oversampling to all three and class-weighted focal loss to the neural models. Faithfulness is measured with deletion-based comprehensiveness and sufficiency, before and after each mitigation, per category, on CICIoT2023 and CICIDS2017. Bias mitigation affects faithfulness unevenly. On the primary dataset, the rare categories most helped gain recall and more faithful explanations, while the largest penalty fell on a category whose detection scarcely changed. XGBoost, explained by exact TreeSHAP, showed no comprehensiveness cost of its own. Among the neural mitigations, focal loss was the most costly, lowering comprehensiveness by up to 0.13 while reducing the recall parity gap by 0.42–0.45, and oversampling was nearly free; this pattern held on both datasets, more weakly on the second, where class weighting and oversampling were not reliably separated. Fairness and interpretability need not be in tension when the mitigation is chosen well, and faithfulness should be evaluated per category whenever a fairness intervention is applied. Full article
Show Figures

Figure 1

33 pages, 850 KB  
Article
Workflow-Level Data Valuation with Stable Compensation via Least-Core: A Cooperative Game-Theoretic Framework
by Shiqian Liu, Peizheng Wang and Chao Wu
Mathematics 2026, 14(15), 2807; https://doi.org/10.3390/math14152807 - 5 Aug 2026
Viewed by 161
Abstract
Existing data valuation methods assign value to individual data points, ignoring the workflow structure through which value is generated in machine learning (ML) pipelines. In practice, value emerges from a chain of interdependent actions—collection, preprocessing, feature engineering, and model training—performed by different agents [...] Read more.
Existing data valuation methods assign value to individual data points, ignoring the workflow structure through which value is generated in machine learning (ML) pipelines. In practice, value emerges from a chain of interdependent actions—collection, preprocessing, feature engineering, and model training—performed by different agents with intertwined incentives. We propose a workflow-level data valuation framework that jointly addresses ownership attribution and compensation stability. We formalize the Data Chain (DC) and derive an ownership assignment rule from incomplete contract theory: The agents with the highest marginal contributions are retained in the compensation coalition. Under supermodular characteristic functions and a stated contextual-dominance condition, this choice does not increase the Least-core deficit relative to any competing retained set reachable by single-action swaps. Recognizing that workflow structures demand coalition stability over distributional fairness, we adopt the Least-core with contribution-aware coefficients to allocate compensation. We prove that the surplus game transformation preserves the core structure (core translation invariance), that the Data Chain Shapley value satisfies the efficiency axiom, and that our LP formulation with Wi=|ψi| guarantees feasibility. Experiments on multiple datasets show that, with a linear base classifier, competitive integration achieves full or near-full coalition-constraint satisfaction while preserving per-agent feasibility. We also assess supermodularity empirically and identify a narrower stability regime when high-capacity models are substitutable. These results define the framework’s present scope and scalability limits. Full article
Show Figures

Figure 1

29 pages, 1533 KB  
Article
A Clinician-in-the-Loop Framework for Validating and Selecting Synthetic Paediatric Dermatology Images
by Ali Tariq Nagi, Chiara Bellatreccia, Andrea Borghesi, Arianna Dondi, Luca Pierantoni, Daniele Zama, Iria Neri, Marcello Lanari and Roberta Calegari
Information 2026, 17(8), 749; https://doi.org/10.3390/info17080749 - 1 Aug 2026
Viewed by 179
Abstract
Synthetic data are increasingly proposed as a strategy for addressing data scarcity and representation imbalance in medical AI, particularly for paediatric populations and darker skin tones. However, visually plausible synthetic images may still contain clinically implausible features or fairness-relevant inconsistencies that are not [...] Read more.
Synthetic data are increasingly proposed as a strategy for addressing data scarcity and representation imbalance in medical AI, particularly for paediatric populations and darker skin tones. However, visually plausible synthetic images may still contain clinically implausible features or fairness-relevant inconsistencies that are not adequately captured by automatic image-quality metrics. In this study, we present and empirically evaluate a clinician-guided framework for validating and selecting synthetic paediatric dermatology images. The framework combines a clinician-facing evaluation platform with structured assessments of visual realism, mask quality, diagnostic plausibility, confidence, and skin-tone relevance. Four clinicians with complementary expertise in paediatrics and dermatology completed 282 assessments of 93 real and synthetic images. Synthetic images were often rated as visually realistic but showed lower inter-rater agreement and weaker mask-quality assessments than real images. Clinician realism and confidence ratings were then used to divide 30 synthetic images into 18 approved and 12 non-approved images. To assess downstream utility, we compared a real-only ResNet50 classifier with classifiers augmented using all synthetic images, clinician-approved synthetic images, or non-approved synthetic images. Across three patient-level experimental splits, the clinician-approved condition achieved the strongest overall classification performance and the largest gains for the under-represented Dark-Skin subgroup. Because the Dark-Skin subgroup contained only seven patients and the synthetic subsets differed in size and disease composition, these fairness results should be interpreted as exploratory. The present study therefore provides evidence for clinician-guided validation and data curation rather than for a completed iterative generator-retraining process. Future work will evaluate whether clinician feedback can also support repeated generative-model refinement in larger, multi-centre datasets. Full article
(This article belongs to the Special Issue Information Technology for Smart Healthcare)
Show Figures

Figure 1

Back to TopTop