Next Article in Journal
A Study on the Identification of the Water Army to Improve the Helpfulness of Online Product Reviews
Next Article in Special Issue
Enhancing Mirror and Glass Detection in Multimodal Images Based on Mathematical and Physical Methods
Previous Article in Journal
Pricing of a Binary Option Under a Mixed Exponential Jump Diffusion Model
Previous Article in Special Issue
Implicit Stance Detection with Hashtag Semantic Enrichment
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Data Quality-Aware Client Selection in Heterogeneous Federated Learning

1
School of Computer Science and Engineering, Changchun University of Technology, Changchun 130012, China
2
College of Bigdata and Internet, Shenzhen Technology University, Shenzhen 518118, China
3
School of Electronic Information, Sichuan University, Chengdu 610017, China
*
Author to whom correspondence should be addressed.
Mathematics 2024, 12(20), 3229; https://doi.org/10.3390/math12203229
Submission received: 10 September 2024 / Revised: 3 October 2024 / Accepted: 8 October 2024 / Published: 15 October 2024

Abstract

Federated Learning (FL) enables decentralized data utilization while maintaining edge user privacy, but it faces challenges due to statistical heterogeneity. Existing approaches address client drift and data heterogeneity issues. However, real-world settings often involve low-quality data with noisy features, such as covariate drift or adversarial samples, which are usually ignored. Noisy samples significantly impact the global model’s accuracy and convergence rate. Assessing data quality and selectively aggregating updates from high-quality clients is crucial, but dynamically perceiving data quality without additional computations or data exchanges is challenging. In this paper, we introduce the FedDQA (Federated learning via Data Quality-Aware) (FedDQA) framework. We discover increased data noise leads to slower loss reduction during local model training. We propose a loss sharpness-based Data-Quality-Awareness (DQA) metric to differentiate between high-quality and low-quality data. Based on the DQA, we design a client selection algorithm that strategically selects participant clients to reduce the negative impact of noisy clients. Experiment results indicate that FedDQA significantly outperforms the baselines. Notably, it achieves up to a 4% increase in global model accuracy and demonstrates faster convergence rates.
Keywords: heterogeneous federated learning; data quality; loss sharpness; noisy data space heterogeneous federated learning; data quality; loss sharpness; noisy data space

Share and Cite

MDPI and ACS Style

Song, S.; Li, Y.; Wan, J.; Fu, X.; Jiang, J. Data Quality-Aware Client Selection in Heterogeneous Federated Learning. Mathematics 2024, 12, 3229. https://doi.org/10.3390/math12203229

AMA Style

Song S, Li Y, Wan J, Fu X, Jiang J. Data Quality-Aware Client Selection in Heterogeneous Federated Learning. Mathematics. 2024; 12(20):3229. https://doi.org/10.3390/math12203229

Chicago/Turabian Style

Song, Shinan, Yaxin Li, Jin Wan, Xianghua Fu, and Jingyan Jiang. 2024. "Data Quality-Aware Client Selection in Heterogeneous Federated Learning" Mathematics 12, no. 20: 3229. https://doi.org/10.3390/math12203229

APA Style

Song, S., Li, Y., Wan, J., Fu, X., & Jiang, J. (2024). Data Quality-Aware Client Selection in Heterogeneous Federated Learning. Mathematics, 12(20), 3229. https://doi.org/10.3390/math12203229

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop