DPCK: An Adaptive Differential Privacy-Based CK-Means Clustering Scheme for Smart Meter Data Analysis
Abstract
1. Introduction
- We propose a CK-means clustering method by improving the K-means algorithm. This method only computes data similarity between data points and the adjacent cluster center set, effectively reducing the computational workload. In addition, during iterations, data that do not require repeated computation are placed into a stability area, avoiding repeated computation. This method significantly decreases the computational overhead of data analysis.
- We design an adaptive differential privacy mechanism to protect smart meter data. During the CK-means analysis of electricity consumption data, this mechanism calculates an appropriate privacy budget for each cluster based on its distribution and adds Laplace noise. It protects data privacy while enhancing data availability in the clustering process.
- Theoretical analysis demonstrates that DPCK provides differential privacy protection and effectively protects user privacy. Experimental results show that, compared to baseline methods, DPCK effectively reduces the computational overhead of data analysis and improves data availability by 11.3% while preserving data privacy.
2. Related Work
2.1. K-Means Clustering
2.2. Differential Privacy-Based K-Means
3. System Model and Related Definitions
3.1. System Model
- Smart Meters: Responsible for collecting users’ electricity consumption data and transmitting them to the data clusterer.
- Data Clusterer: Responsible for performing clustering analysis on the collected electricity consumption data, applying adaptive differential privacy to protect the data, and finally sending the noise-added clustering results to the control centers.
- Control Centers: Responsible for analyzing the received clustering results statistically and classifying users based on these data, facilitating operations such as service provision by the electric company.
3.2. Definitions
3.2.1. K-Means Clustering
3.2.2. Differential Privacy
4. Our Proposed DPCK Scheme
4.1. Initialization
4.2. CK-Means Clustering
4.2.1. Calculating the Adjacent Cluster Centers Set
| Algorithm 1 Calculating the adjacent cluster center set. |
Input: Dataset X, k, initial centers c. Output: , . 1: for to k do 2: if then 3: Calculate the movement of the center of cluster ; 4: Calculate the radius ; 5: end if 6: end for 7: if
then 8: Calculate the distance ; 9: else 10: for to k do 11: for to k do 12: if then 13: ; 14: end if 15: end for 16: end for 17: end if 18: for
do 19: if then 20: Append to ; 21: end if 22: end for 23: return , . |
4.2.2. Data Assignment
| Algorithm 2 Data assignment. |
Input: Dataset X, k, , . Output: The set of centers c, the set of clusters C. 1: for to k do 2: if and and then 3: continue 4: else 5: Sort by distance in ascending order; 6: for each x in do 7: if then 8: continue 9: else 10: if x in the i-th circular area then 11: Compute the distance from x to its first i closest centers; 12: Assign x to the nearest cluster; 13: end if 14: end if 15: end for 16: end if 17: end for 18: for to k do 19: if is stable then 20: = TRUE 21: else 22: = FALSE 23: end if 24: end for 25: ; 26: return c, C. |
4.3. Adaptive Differential Privacy
5. Theoretical Analysis
5.1. Complexity Analysis
5.2. Privacy Analysis
6. Experiments
6.1. Exprimental Settings
- Iris: This dataset includes three classes, with 50 instances per class, totaling 150 instances, and four attributes. Each class represents a type of iris plant. Notably, the 35th sample needs to be manually modified to 4.9, 3.1, 1.5, 0.2, ‘Iris-setosa’; the 38th sample needs to be manually modified to 4.9, 3.6, 1.4, 0.1, ‘Iris-setosa’.
- Wine: This dataset includes 178 instances, divided into three classes, each containing 13 attribute, with no missing values.
- Electrical Grid data: This dataset includes 10,000 instances, 12 attributes, and no missing values. It is a simulated dataset designed for studying the stability of electrical grid systems.
- Gamma: This dataset includes 19,020 instances, 10 attributes, and no missing values.
- Eco dataset: The Eco dataset provides total electricity consumption data at 1 Hz, collected as part of the smart meter service project at ETH Zurich [42]. Each file contains 86,400 rows (i.e., one row per second), with rows with missing measurements represented by “−1”.
6.2. Evaluation Criteria Metrics
6.3. Discussion of Experiments
6.3.1. Computational Overhead
6.3.2. Convergence
6.3.3. Data Availability
6.3.4. Clustering Performance
7. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Mari, A.; Remlinger, C.; Castello, R.; Obozinski, G.; Quarteroni, S.; Heymann, F.; Galus, M. Real-time estimates of Swiss electricity savings using streamed smart meter data. Appl. Energy 2025, 377, 124537. [Google Scholar] [CrossRef] [Scilit]
- Athanasiadis, C.L.; Papadopoulos, T.A.; Kryonidis, G.C.; Doukas, D.I. A review of distribution network applications based on smart meter data analytics. Renew. Sustain. Energy Rev. 2024, 191, 114151. [Google Scholar] [CrossRef] [Scilit]
- Gumz, J.; Fettermann, D.C. User’s perspective in smart meter research: State-of-the-art and future trends. Energy Build. 2024, 308, 114025. [Google Scholar] [CrossRef] [Scilit]
- Zhang, S.; Mao, X.; Choo, K.-K.R.; Peng, T.; Wang, G. A trajectory privacy-preserving scheme based on a dual-K mechanism for continuous location-based services. Inf. Sci. 2020, 527, 406–419. [Google Scholar] [CrossRef] [Scilit]
- Xiong, A.; Zhou, H.; Song, Y.; Wang, D.; Wei, X.; Li, D.; Gao, B. A multi-task based clustering personalized federated learning method. Big Data Min. Anal. 2024, 7, 1017–1030. [Google Scholar] [CrossRef] [Scilit]
- Wang, S.; Song, A.; Qian, Y. Predicting smart cities’ electricity demands using k-means clustering algorithm in smart grid. Comput. Sci. Inf. Syst. 2023, 20, 657–678. [Google Scholar] [CrossRef] [Scilit]
- Yuan, L.; Zhang, S.; Zhu, G.; Alinani, K. Privacy-preserving mechanism for mixed data clustering with local differential privacy. Concurr. Comput. Pract. Exp. 2023, 35, e6503. [Google Scholar] [CrossRef] [Scilit]
- Choksi, K.A.; Jain, S.; Pindoriya, N.M. Feature based clustering technique for investigation of domestic load profiles and probabilistic variation assessment: Smart meter dataset. Sustain. Energy Grids Netw. 2020, 22, 100346. [Google Scholar] [CrossRef] [Scilit]
- Rafiq, H.; Manandhar, P.; Rodriguez-Ubinas, E.; Barbosa, J.D.; Qureshi, O.A. Analysis of residential electricity consumption patterns utilizing smart-meter data: Dubai as a case study. Energy Build. 2023, 291, 113103. [Google Scholar] [CrossRef] [Scilit]
- Xu, H.; Yao, S.; Li, Q.; Ye, Z. An improved k-means clustering algorithm. In Proceedings of the 2020 IEEE 5th International Symposium on Smart and Wireless Systems within the Conferences on Intelligent Data Acquisition and Advanced Computing Systems (IDAACS-SWS), Dortmund, Germany, 17–18 September 2020; pp. 1–5. [Google Scholar]
- Alguliyev, R.M.; Aliguliyev, R.M.; Sukhostat, L.V. Parallel batch k-means for Big data clustering. Comput. Ind. Eng. 2021, 152, 107023. [Google Scholar] [CrossRef] [Scilit]
- Nie, F.; Li, Z.; Wang, R.; Li, X. An effective and efficient algorithm for K-means clustering with new formulation. IEEE Trans. Knowl. Data Eng. 2022, 35, 3433–3443. [Google Scholar] [CrossRef] [Scilit]
- Zhang, S.; Pan, Y.; Liu, Q.; Yan, Z.; Choo, K.-K.R.; Wang, G. Backdoor attacks and defenses targeting multi-domain AI models: A comprehensive review. ACM Comput. Surv. 2024, 57, 1–35. [Google Scholar] [CrossRef] [Scilit]
- Parker, K.; Hale, M.; Barooah, P. Spectral differential privacy: Application to smart meter data. IEEE Internet Things J. 2021, 9, 4987–4996. [Google Scholar] [CrossRef] [Scilit]
- Zhu, P.; Hu, J.; Li, X.; Zhu, Q. Using blockchain technology to enhance the traceability of original achievements. IEEE Trans. Eng. Manag. 2023, 70, 1693–1707. [Google Scholar] [CrossRef] [Scilit]
- Wang, J.; Wu, L.; Zeadally, S.; Khan, M.K.; He, D. Privacy-preserving data aggregation against malicious data mining attack for IoT-enabled smart grid. ACM Trans. Sens. Netw. (TOSN) 2021, 17, 1–25. [Google Scholar] [CrossRef] [Scilit]
- Zhang, S.; Chen, W.; Li, X.; Liu, Q.; Wang, G. APBAM: Adversarial perturbation-driven backdoor attack in multimodal learning. Inf. Sci. 2025, 700, 121847. [Google Scholar] [CrossRef] [Scilit]
- Hu, J.; Zhu, P.; Li, J.; Qi, Y.; Xia, Y.; Wang, F.-Y. A secure medical information storage and sharing method based on multiblockchain architecture. IEEE Trans. Comput. Soc. Syst. 2024, 11, 6392–6406. [Google Scholar] [CrossRef] [Scilit]
- Zhang, S.; Li, X.; Tan, Z.; Peng, T.; Wang, G. A caching and spatial K-anonymity driven privacy enhancement scheme in continuous location-based services. Future Gener. Comput. Syst. 2019, 94, 40–50. [Google Scholar] [CrossRef] [Scilit]
- Zhang, S.; Choo, K.-K.R.; Liu, Q.; Wang, G. Enhancing privacy through uniform grid and caching in location-based services. Future Gener. Comput. Syst. 2018, 86, 881–892. [Google Scholar] [CrossRef] [Scilit]
- He, J.; Wang, N.; Xiang, T.; Wei, Y.; Zhang, Z.; Li, M.; Zhu, L. ABDP: Accurate Billing on Differentially Private Data Reporting for Smart Grids. IEEE Trans. Serv. Comput. 2024, 17, 1938–1954. [Google Scholar] [CrossRef] [Scilit]
- Gough, M.B.; Santos, S.F.; AlSkaif, T.; Javadi, M.S.; Castro, R.; Catalão, J.P.S. Preserving privacy of smart meter data in a smart grid environment. IEEE Trans. Ind. Inform. 2021, 18, 707–718. [Google Scholar] [CrossRef] [Scilit]
- Zheng, Z.; Wang, T.; Bashir, A.K.; Alazab, M.; Mumtaz, S.; Wang, X. A decentralized mechanism based on differential privacy for privacy-preserving computation in smart grid. IEEE Trans. Comput. 2021, 71, 2915–2926. [Google Scholar] [CrossRef] [Scilit]
- MacQueen, J. Some methods for classification and analysis of multivariate observations. In Proceedings of the 5th Berkeley Symposium on Mathematical Statistics and Probability, Berkeley, CA, USA, 21 June–18 July 1965 and 27 December 1965–7 January 1966; University of California Press/University of California: Berkeley, CA, USA, 1967; pp. 281–298. [Google Scholar]
- Xiang, Y.; Hong, J.; Yang, Z.; Wang, Y.; Huang, Y.; Zhang, X.; Chai, Y.; Yao, H. Slope-Based Shape Cluster Method for Smart Metering Load Profiles. IEEE Trans. Smart Grid 2020, 11, 1809–1811. [Google Scholar] [CrossRef]
- Khan, I.; Luo, Z.; Shaikh, A.K.; Hedjam, R. Ensemble clustering using extended fuzzy k-means for cancer data analysis. Expert Syst. Appl. 2021, 172, 114622. [Google Scholar] [CrossRef] [Scilit]
- Hu, H.; Liu, J.; Zhang, X.; Fang, M. An effective and adaptable K-means algorithm for big data cluster analysis. Pattern Recognit. 2023, 139, 109404. [Google Scholar] [CrossRef] [Scilit]
- Chen, Y.; Tan, P.; Li, M.; Yin, H.; Tang, R. K-means clustering method based on nearest-neighbor density matrix for customer electricity behavior analysis. Int. J. Electr. Power Energy Syst. 2024, 161, 110165. [Google Scholar] [CrossRef] [Scilit]
- Yang, M.; Huang, L.; Tang, C. K-means clustering with local distance privacy. Big Data Min. Anal. 2023, 6, 433–442. [Google Scholar] [CrossRef] [Scilit]
- Zhang, S.; Liu, Q.; Wang, T.; Liang, W.; Li, K.-C.; Wang, G. FSAIR: Fine-grained secure approximate image retrieval for mobile cloud computing. IEEE Internet Things J. 2024, 11, 23297–23308. [Google Scholar] [CrossRef] [Scilit]
- Gupta, A.; Ligett, K.; McSherry, F.; Roth, A.; Talwar, K. Differentially private combinatorial optimization. In Proceedings of the Twenty-First Annual ACM-SIAM Symposium on Discrete Algorithms, Austin, TX, USA, 17–19 January 2010; pp. 1106–1125. [Google Scholar]
- Yu, Q.; Luo, Y.; Chen, C.; Ding, X. Outlier-eliminated k-means clustering algorithm based on differential privacy preservation. Appl. Intell. 2016, 45, 1179–1191. [Google Scholar] [CrossRef] [Scilit]
- Wu, F.; Du, M.; Zhi, Q. Density-based clustering with differential privacy. Inf. Sci. 2024, 681, 121211. [Google Scholar] [CrossRef] [Scilit]
- Xiong, J.; Ren, J.; Chen, L.; Yao, Z.; Lin, M.; Wu, D.; Niu, B. Enhancing privacy and availability for data clustering in intelligent electrical service of IoT. IEEE Internet Things J. 2018, 6, 1530–1540. [Google Scholar] [CrossRef] [Scilit]
- Zhang, M.; Zhou, J.; Zhang, G.; Cui, L.; Gao, T.; Yu, S. APDP: Attribute-based personalized differential privacy data publishing scheme for social networks. IEEE Trans. Netw. Sci. Eng. 2022, 10, 922–933. [Google Scholar] [CrossRef] [Scilit]
- Zhang, S.; Zhang, L.; Peng, T.; Liu, Q.; Li, X. VADP: Visitor-attribute-based adaptive differential privacy for IoMT data sharing. Comput. Secur. 2025, 156, 104513. [Google Scholar] [CrossRef] [Scilit]
- Balcan, M.F.; Dick, T.; Liang, Y.; Mou, W.; Zhang, H. Differentially private clustering in high-dimensional Euclidean spaces. In Proceedings of the International Conference on Machine Learning, Sydney, Australia, 6–11 August 2017; pp. 322–331. [Google Scholar]
- He, Z.; Wang, L.; Cai, Z. Clustered federated learning with adaptive local differential privacy on heterogeneous IoT data. IEEE Internet Things J. 2023, 11, 137–146. [Google Scholar] [CrossRef] [Scilit]
- Al Shalabi, L.; Shaaban, Z. Normalization as a preprocessing engine for data mining and the approach of preference matrix. In Proceedings of the 2006 International Conference on Dependability of Computer Systems, Szklarska Poreba, Poland, 25–27 May 2006; pp. 207–214. [Google Scholar]
- Zhou, H.B.; Gao, J.T. Automatic method for determining cluster number based on silhouette coefficient. Adv. Mater. Res. 2014, 951, 227–230. [Google Scholar] [CrossRef] [Scilit]
- Dua, D.; Graff, C. UCI Machine Learning Repository; School of Information and Computer Sciences, University of California: Irvine, CA, USA, 2007; Available online: https://archive.ics.uci.edu (accessed on 22 April 2025).
- Kleiminger, W.; Beckel, C.; Santini, S. Household Occupancy Monitoring Using Electricity Meters. In Proceedings of the 2015 ACM International Joint Conference on Pervasive and Ubiquitous Computing (UbiComp 2015), Osaka, Japan, 9–11 September 2015. [Google Scholar]
- Arthur, D.; Vassilvitskii, S. k-Means++: The Advantages of Careful Seeding; Technical Report 2006-13; Stanford InfoLab: Stanford, CA, USA, 2006; Available online: http://ilpubs.stanford.edu:8090/778/ (accessed on 13 May 2025).








| Method | Computational Overhead | Privacy Mechanism | Data Utility |
|---|---|---|---|
| [24] | ✕ | None | ✓ |
| [25] | ✕ | None | ✓ |
| [26] | ✕ | None | ✓ |
| [27] | ✓ | None | ✓ |
| [28] | ✓ | None | ✓ |
| [31] | ✕ | DP | ✕ |
| [32] | ✕ | DP | ✕ |
| [33] | ✕ | DP | ✓ |
| [34] | ✕ | DP | ✓ |
| [37] | ✕ | DP | ✓ |
| [38] | ✕ | Localized DP | ✓ |
| DPCK (Ours) | ✓✓ | Adaptive DP | ✓✓ |
| Symbol | Description |
|---|---|
| X | The dataset |
| N | The size of the dataset |
| k | The number of clusters |
| The i-th cluster | |
| The center of | |
| The radius of | |
| The adjacent cluster center set of | |
| t | The number of iterations |
| The privacy budget | |
| The sum of all data points in | |
| The number of all data points in |
| Dataset | DP K-Means | PADC [34] | DPCK | |||
|---|---|---|---|---|---|---|
| P | R | P | R | P | R | |
| Iris | 0.7675 | 0.8804 | 0.7856 | 0.8855 | 0.8042 | 0.9029 |
| Wine | 0.7292 | 0.6240 | 0.7554 | 0.6283 | 0.7679 | 0.6555 |
| Electrical Grid Data | 0.5253 | 0.4172 | 0.5652 | 0.4414 | 0.5837 | 0.4583 |
| Gamma | 0.6654 | 0.3122 | 0.7860 | 0.3424 | 0.7957 | 0.3514 |
| Dataset | DP K-Means | PADC [34] | DPCK | |||
|---|---|---|---|---|---|---|
| P | R | P | R | P | R | |
| Iris | 0.8214 | 0.9320 | 0.8299 | 0.9416 | 0.8311 | 0.9441 |
| Wine | 0.7920 | 0.6747 | 0.8045 | 0.6750 | 0.8112 | 0.6756 |
| Electrical Grid Data | 0.6828 | 0.5097 | 0.6972 | 0.5325 | 0.6973 | 0.5333 |
| Gamma | 0.9972 | 0.5165 | 0.9995 | 0.5241 | 0.9997 | 0.5510 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
Share and Cite
Zhang, S.; Zhu, J.; Luo, E.; Zhu, X.; Yang, Q. DPCK: An Adaptive Differential Privacy-Based CK-Means Clustering Scheme for Smart Meter Data Analysis. Electronics 2025, 14, 2074. https://doi.org/10.3390/electronics14102074
Zhang S, Zhu J, Luo E, Zhu X, Yang Q. DPCK: An Adaptive Differential Privacy-Based CK-Means Clustering Scheme for Smart Meter Data Analysis. Electronics. 2025; 14(10):2074. https://doi.org/10.3390/electronics14102074
Chicago/Turabian StyleZhang, Shaobo, Jielu Zhu, Entao Luo, Xiaoyu Zhu, and Qing Yang. 2025. "DPCK: An Adaptive Differential Privacy-Based CK-Means Clustering Scheme for Smart Meter Data Analysis" Electronics 14, no. 10: 2074. https://doi.org/10.3390/electronics14102074
APA StyleZhang, S., Zhu, J., Luo, E., Zhu, X., & Yang, Q. (2025). DPCK: An Adaptive Differential Privacy-Based CK-Means Clustering Scheme for Smart Meter Data Analysis. Electronics, 14(10), 2074. https://doi.org/10.3390/electronics14102074

