Unsupervised Real-Time Anomaly Detection in Hydropower Systems via Time Series Clustering and Autoencoders
Abstract
1. Introduction
- Static or point anomalies refer to isolated data points that differ significantly from expected behavior, such as a sudden spike in energy consumption, an abrupt rise in pipeline pressure, a malfunctioning sensor, or an unexpected drop in water level.
- Contextual refers to data points that are only considered anomalous within a specific context, such as time, location, or operational conditions—for example, a high temperature that is normal in summer but abnormal in winter.
- Dynamic anomalies, also called temporal or collective anomalies, involve a sequence of data points that together form an anomalous pattern over time. An example is the gradual deterioration of a turbine, where multiple variables slowly drift from normal ranges until failure occurs.
1.1. Related Works
1.1.1. Anomaly Detection in Hydropower Systems
1.1.2. Unsupervised Anomaly Detection Techniques
1.1.3. Unsupervised Anomaly Detection in Multivariate Systems
1.1.4. Comparative Analysis
1.2. Research Significance
2. Materials and Methods
2.1. Problem Statement
- Scalability: Designing and maintaining separate anomaly detection models is inefficient.
- Temporal behavior: Anomalies can be either point anomalies (sharp outliers) or temporal anomalies (gradual drift), requiring distinct detection strategies.
- Similarity among variables: Some variables may exhibit similar patterns, suggesting potential for grouping and shared modeling.
2.2. Proposed System Architecture
2.2.1. Anomaly Detection Models Creation
- Anomaly-Free Data Preparation: To develop reliable anomaly detection models, it is essential to start with a dataset that is free of anomalies. This phase included data preprocessing steps such as anomalies cleaning, handling of missing values, and variable standardization. This ensures that the models learn the normal behavior patterns without being biased by abnormal conditions.
- Time Series Clustering: Detecting anomalies for each variable independently would require developing a large number of individual models. To mitigate this, we propose a time series clustering approach that groups variables exhibiting similar temporal behavior. The dataset was transposed so that each row represents a variable and each column a timestamp. The K-Means algorithm was applied to identify clusters of variables with similar time series patterns, enabling the development of a single detection model per cluster. This step allowed the identification of common behavioral patterns and the development of specialized models for each cluster. K-means was chosen for its simplicity and efficiency in partitioning large datasets.
- Cluster Centroids Definition: For each cluster obtained, a centroid was calculated to represent the average time series behavior of the variables within the group. The centroids were then transposed, so that each cluster was represented as a column and each row corresponds to a time point, forming the basis for model training.
- Unsupervised Anomaly Detection Modeling: Each cluster centroid was used to train unsupervised anomaly detection models, including ARIMA, Autoencoder, VAE, LSTM, Sliding-window Autoencoders and Sliding-window VAE. These models were trained using the historical power data and the cluster representative, capturing the typical behavior of the associated variables. Power must be considered because it directly reflects the system’s energy consumption, and anomalous changes in its behavior may indicate operational failures, inefficiencies, or atypical process conditions. Furthermore, variable behavior may vary significantly during system startup or shutdown; in such scenarios, fluctuations should not be flagged as anomalies. However, if fluctuations arise while power remains stable, they may indeed signal anomalous behavior.
- Adaptive Anomaly Thresholding: Anomalies were identified by comparing the predicted value of a variable with its actual value. If the discrepancy exceeds a defined threshold, it is flagged as an anomaly. However, determining an appropriate threshold is nontrivial. A fixed threshold may result in false positives or missed anomalies. To address this, we defined the anomaly threshold based on the predictive reliability of the model. When the model exhibits low prediction error, even small differences may indicate anomalies. In contrast, models with higher error require more tolerant thresholds to avoid misclassification.
2.2.2. Real-Time Anomaly Detection
- Real-Time Data Ingestion: During the real-time operation of the hydroelectric power plant, sensor data is continuously collected and stored in the cloud. A scheduled process reads data in fixed time windows (e.g., every 15 min) from the database. Once a time window is retrieved, the anomaly detection routine is executed, with each task placed into a processing queue. This queuing mechanism ensures that results are written back to the database in a sequential and controlled manner.
- Variable-Based Anomaly Detection: Anomalies are detected for each variable individually, using the reconstruction error produced by the autoencoder model corresponding to the assigned cluster. For example, if Variable 1 is assigned to Cluster 0, it will be evaluated using the detection model specifically trained for that cluster. Importantly, each variable uses a specific anomaly threshold, calculated as twice the standard deviation of its own prediction error. This ensures that the detection process is tailored to the behavior and variability of each variable, rather than applying a common threshold across all variables. By combining cluster-specific models with variable-specific thresholds, the method enhances the precision of anomaly detection while reducing false positives.
- Anomaly Score per Variable: For each variable, a new column is added containing the anomaly score. This score quantifies the degree of deviation from normal behavior, as learned by the corresponding model.
- Anomaly Logging and Response: Anomaly scores are stored in the database for future analysis and visualization. Depending on the severity of the detected anomalies, automated responses could be triggered, such as pausing or adjusting operational processes, notifying plant operators, or activating contingency systems.
2.3. Methodology
- Business Understanding: A detailed understanding of the operational context of the power generation systems was developed in collaboration with domain experts. The main objective was defined as the detection of anomalies in high-dimensional sensor data from power plants.
- Data Acquisition and Understanding: Historical datasets from a power generation unit were explored to assess data quality, completeness, and relevance. Particular attention was given to selecting the main operational variables and understanding their typical behavior under normal operating conditions.
- Data Preparation: The time series data were preprocessed to address missing values, outliers, and synchronization issues. Feature engineering was also performed to derive meaningful attributes for the subsequent modeling phase.
- Modeling: The proposed approach involved three main components such as (1) Time series clustering and centroid definition, used to group variables with similar temporal behavior. This step enabled contextual analysis while reducing the total number of models required; (2) Unsupervised anomaly detection models, developed to identify unusual patterns without relying on labeled data; and (3) Adaptive anomaly thresholding, designed to adjust decision boundaries dynamically according to the statistical characteristics of each cluster. This modular structure allowed for flexible experimentation and model calibration, ensuring the proposed solution could generalize across multiple operating scenario.
- Evaluation: Model outputs were evaluated using historical anomalies data and expert judgment. Iterative feedback loops enabled the refinement of model parameters.
- Deployment and Monitoring: A real-time anomaly detection system was implemented using a four-phase architecture: (1) Real-Time Data Ingestion, where streaming data from monitored assets was collected and pre-processed; (2) Variable-Based Anomaly Detection, in which each variable was evaluated independently using the trained models to compute anomaly scores; (3) Anomaly Score per Variable, allowing fine-grained detection and interpretation of abnormal behavior; and (4) An additional module, Anomaly Logging and Response, was used to store alerts and support timely operational response. A real-time dashboard was developed to support continuous monitoring, displaying key operational metrics and detected anomalies.
2.4. Technical Implementation Specifications
2.5. Use of Generative AI Tools
3. Results
3.1. Anomaly Detection Models Results
3.1.1. Anomaly-Free Data Preparation
- Training Set: This subset was used to train machine learning models for anomaly detection based on two-months historical operational behavior. The training set undergoes a cleaning process to ensure that the models were trained on data free from anomalies.
- Testing Set: This one-month subset was reserved for evaluating the performance of the models and validating the error thresholds used to differentiate between normal and anomalous behavior. The test set is divided into two subsets: one clean and another containing known anomalies.
3.1.2. Time Series Clustering
- The inertia continues to decrease meaningfully up to k = 16, suggesting that additional clusters do not contribute valuable segmentation without overfitting the data.
- The distribution of variables across clusters shows a balanced and interpretable structure, with some clusters capturing dominant patterns and others revealing more specialized behaviors.
3.1.3. Cluster Centroids Definition
3.1.4. Unsupervised Anomaly Detection Modeling
3.1.5. Adaptive Anomaly Thresholding
- For models with low prediction error, the threshold becomes narrower, enabling the detection of even minor anomalies.
- For models with higher variability, the threshold becomes wider, reducing the likelihood of false positives.
3.2. Evaluation
3.3. Real-Time Implementation and Monitoring
4. Discussion
4.1. Discussion on Anomaly Detection Models Creation
- The ARIMA model showed the weakest performance, with a wide range of R2 values, including negative scores, indicating poor fit and limited ability to capture the non-linear dynamics of the data. Its high MAPE values also reveal large reconstruction errors, confirming its inadequacy for modeling temporal dependencies in complex signals.
- Autoencoder and LSTM models achieved consistently high R2 scores, with the Autoencoder exhibiting the lowest variability, indicating strong and stable reconstruction capabilities. Correspondingly, the Autoencoder obtained the lowest MAPE values across all clusters, confirming its superior reconstruction accuracy and robustness against dynamic variations in the data.
- VAE demonstrated high variability in both R2 and MAPE, with some clusters showing poor reconstruction performance. This suggests sensitivity to the complexity of local patterns.
- The sliding-window variants exhibited mixed results, while some samples achieving high R2 and low MAPE values, others showed the opposite behavior. This indicates that their performance is sensitive to the selected window length and local fluctuations in the data.
4.2. Discussion on Evaluation
- First column: Displays the prediction error over time. Anomalies are flagged when the error exceeds two standard deviations from the mean. This statistical threshold highlights instances where the model’s predictions deviate notably from the observed data, suggesting possible abnormal system behavior.
- Second column: The blue points represent the actual measured values of the variable, while the yellow points correspond to the model’s predicted values. When there is a strong alignment, the model is considered accurate. In several cases, the yellow and blue curves diverge clearly around detected anomaly points, reinforcing the validity of the flags.
- Third column: Overlays the actual values (blue), anomalies (red), and power output (purple). The inclusion of power output provides critical operational context. For several variables, anomalies coincide with fluctuations or drops in power, which suggests a possible causal relationship. For example, rows 1, 3, and 6 exhibit clear anomaly peaks in red that align temporally with power dips, hinting at a connection between operational instability and variable behavior.
4.3. Anomaly Detection Models Limitations
4.4. Operational Validation and Expert Assessment
4.5. Inference and Retraining Time
5. Conclusions and Future Work
- Anomaly detection with unlabeled data: the concept of using process variable data to identify changes in condition and equipment deterioration without the need to label data with specific faults opens up a huge range of applications for effective machine condition monitoring for hundreds of industries that do not have sufficient fault-related data with which to train supervised models. Anomalies detected with unsupervised monitoring systems are a source of information for labeling data and complementing monitoring with supervised systems.
- Effective modeling using cluster centroids: The anomaly detection models were trained on cluster centroids derived from historical operational data, rather than the full dataset. This strategy significantly reduced the training complexity while preserving the essential patterns of system behavior. Despite this abstraction, the models achieved high predictive performance, showing that the centroids captured the main behaviors of the variables without losing relevant operational details.
- Autoencoder as the best-performing model: Among the evaluated techniques, the Autoencoder achieved the highest predictive performance and was selected for the anomaly detection stage. The model learns the typical behavior of each cluster under normal operating conditions by minimizing reconstruction error during training; when applied in operation, deviations between the actual and reconstructed values are interpreted as anomalies. This approach allows the autoencoder to be effectively used for anomaly detection by anticipating deviations from expected behavior before they escalate into failures.
- High performance of unsupervised models: The robustness of the proposed framework was confirmed by the R2 and MAE metrics, which showed that the models accurately reconstructed the signals.
- Real-time anomaly detection: The implementation of the real-time system allows for the identification of significant deviations between predicted and actual values, facilitating the timely generation of alerts that can be viewed through a Grafana control panel.
- Intuitive visualization for monitoring: the way the results are presented is crucial to achieving the operator’s level of awareness and focus when working with a large number of variables. Grafana is a platform that achieves this goal by presenting the status of systems that associate groups of variables, allowing the operator to identify the status of an entire machine in a single display in a matter of seconds. Additionally, the operator has the ability to drill down to view the variable that affects the status of a system and the machine.
- The proposed system is scalable and can be extended to other power plants, systems, or industries with similar characteristics, allowing for adaptation to various industrial or energy monitoring contexts.
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Alagöz, İ.; Bulut, M.; Geylani, V.; Yildirim, A. Importance of Real-Time Hydro Power Plant Condition Monitoring Systems and Contribution to Electricity Production. Turk. J. Electr. Power Energy Syst. 2021, 1, 1–11. [Google Scholar] [CrossRef] [Scilit]
- Yu, X.; Zhang, S.; Zhang, X.; Yin, X.; Qiu, W.; Wang, Z.; Lin, Q. A Review of the Recent Research on Hydraulic Excitation-Induced Vibration in Hydropower Plants. J. Phys. Conf. Ser. 2024, 2752, 012093. [Google Scholar] [CrossRef] [Scilit]
- Fanan, M.; Baron, C.; Carli, R. Anomaly Detection for Hydroelectric Power Plants: A Machine Learning-Based Approach. In Proceedings of the 2023 IEEE 21st International Conference on Industrial Informatics, Lemgo, Germany, 18–20 July 2023; pp. 1–6. [Google Scholar]
- Wang, F.; Jiang, Y.; Zhang, R.; Wei, A. A Survey of Deep Anomaly Detection in Multivariate Time Series: Taxonomy, Applications, and Directions. Sensors 2025, 25, 190. [Google Scholar] [CrossRef] [Scilit]
- Aung, K.H.H.; Kok, C.L.; Koh, Y.Y.; Teo, T.H. An Embedded Machine Learning Fault Detection System for Electric Fan Drive. Electronics 2024, 13, 493. [Google Scholar] [CrossRef] [Scilit]
- Foorthuis, R. On the Nature and Types of Anomalies: A Review of Deviations in Data. Int. J. Data Sci. Anal. 2021, 12, 297–331. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Fährmann, D.; Martín, L.; Sánchez, L.; Damer, N. Anomaly Detection in Smart Environments: A Comprehensive Survey. IEEE Access 2024, 12, 64006–64049. [Google Scholar] [CrossRef] [Scilit]
- Vargas, J.; Oviedo, A.; Ortega, N.; Orozco, E.; Gómez, A. Machine-Learning-Based Predictive Models for Compressive Strength, Flexural Strength, and Slump of Concrete. Appl. Sci. 2024, 14, 4426. [Google Scholar] [CrossRef] [Scilit]
- De Santis, R.B.; Costa, M.A. Extended Isolation Forests for Fault Detection in Small Hydroelectric Plants. Sustainability 2020, 12, 6421. [Google Scholar] [CrossRef] [Scilit]
- Rai, A. Unsupervised Learning Algorithms for Hydropower’s Sensor Data. In Cybernetics, Cognition and Machine Learning Applications; Springer: Singapore, 2021; pp. 89–94. [Google Scholar] [CrossRef] [Scilit]
- Guan, S.; He, Z.; Ma, S.; Gao, M. Multivariate Time Series Anomaly Detection with Variational Autoencoder and Spatial–Temporal Graph Network. Comput. Secur. 2024, 142, 103877. [Google Scholar] [CrossRef] [Scilit]
- Yang, B.; Lyu, Z.; Wei, H. A Study of an Anomaly Detection System for Small Hydropower Data Considering Multivariate Time Series. Int. Trans. Electr. Energy Syst. 2024, 1, 8108861. [Google Scholar] [CrossRef] [Scilit]
- Himeur, Y.; Alsalemi, A.; Bensaali, F.; Amira, A. Smart Power Consumption Abnormality Detection in Buildings Using Micromoments and Improved K-nearest Neighbors. Int. J. Intell. Syst. 2021, 36, 2865–2894. [Google Scholar] [CrossRef] [Scilit]
- Carletti, M.; Terzi, M.; Susto, G.A. Interpretable Anomaly Detection with Diffi: Depth-Based Feature Importance of Isolation Forest. Eng. Appl. Artif. Intell. 2023, 119, 105730. [Google Scholar] [CrossRef] [Scilit]
- Ghiasi, R.; Khan, M.A.; Sorrentino, D.; Diaine, C. An Unsupervised Anomaly Detection Framework for Onboard Monitoring of Railway Track Geometrical Defects Using One-Class Support Vector Machine. Eng. Appl. Artif. Intell. 2024, 133, 108167. [Google Scholar] [CrossRef] [Scilit]
- Oti, E.U.; Olusola, M.O.; Eze, F.C.; Enogwe, S.U. Comprehensive Review of K-Means Clustering Algorithms. Int. J. Adv. Sci. Res. Eng. 2021, 7, 64–68. [Google Scholar] [CrossRef] [Scilit]
- Jin, F.; Wu, H.; Liu, Y.; Zhao, J.; Wang, W. Varying-Scale HCA-DBSCAN-Based Anomaly Detection Method for Multi-Dimensional Energy Data in Steel Industry. Inf. Sci. 2023, 647, 119479. [Google Scholar] [CrossRef] [Scilit]
- Li, Z.; Van Leeuwen, M. Explainable Contextual Anomaly Detection Using Quantile Regression Forests. Data Min. Knowl. Discov. 2023, 37, 2517–2563. [Google Scholar] [CrossRef] [Scilit]
- Zamanzadeh Darban, Z.; Webb, G.I.; Pan, S.; Aggarwal, C.; Salehi, M. Deep Learning for Time Series Anomaly Detection: A Survey. ACM Comput. Surv. 2024, 57, 42. [Google Scholar] [CrossRef] [Scilit]
- Kozitsin, V.; Katser, I.; Lakontsev, D. Online Forecasting and Anomaly Detection Based on the ARIMA Model. Appl. Sci. 2021, 11, 3194. [Google Scholar] [CrossRef] [Scilit]
- Mavikumbure, H.S.; Wickramasinghe, C.S.; Marino, D.L.; Cobilean, V.; Manic, M. Anomaly Detection in Critical-Infrastructures Using Autoencoders: A Survey. In Proceedings of the IECON 2022–48th Annual Conference of the IEEE Industrial Electronics Society, Brussels, Belgium, 17–20 October 2022. [Google Scholar]
- Sun, C.; He, Z.; Lin, H.; Cai, L.; Cai, H.; Gao, M. Anomaly Detection of Power Battery Pack Using Gated Recurrent Units Based Variational Autoencoder. Appl. Soft Comput. 2023, 132, 109903. [Google Scholar] [CrossRef] [Scilit]
- Hussein, L.; Mohan, M.; Rajeswari, P.; Gurupandi, D.; Naga, N. Anomaly Detection in IoT Data Streams Based on Long Short-Term Memory. In Proceedings of the International Conference on Intelligent Algorithms for Computational Intelligence Systems, Hassan, India, 23–24 August 2024. [Google Scholar]
- Zhong, Z.; Zhao, Y.; Yang, A.; Zhang, H.; Qiao, D.; Zhang, Z. Industrial Robot Vibration Anomaly Detection Based on Sliding Window One-Dimensional Convolution Autoencoder. Shock. Vib. 2022, 2022, 1179192. [Google Scholar] [CrossRef] [Scilit]
- Moschini, G.; Houssou, R.; Bovay, J.; Robert-Nicoud, S.; Rojas, F.; Herrera, L.J.; Pomare, H. Anomaly and Fraud Detection in Credit Card Transactions Using the ARIMA Model. Eng. Proc. 2021, 5, 56. [Google Scholar] [CrossRef] [Scilit]
- Sun, G.; Yin, C.; Xia, T.; Lu, Y.; Mao, J. An Improved ARIMA Based Anomaly Detection Method for Time Series Data. In Proceedings of the 2024 IEEE 8th Conference on Energy Internet and Energy System Integration (EI2), Shenyang, China, 29 November–2 December 2024; pp. 5132–5138. [Google Scholar] [CrossRef] [Scilit]
- Wang, H.; Liu, X.; Ma, L.; Zhang, Y. Anomaly Detection for Hydropower Turbine Unit Based on Variational Modal Decomposition and Deep Autoencoder. Energy Rep. 2021, 7, 938–946. [Google Scholar] [CrossRef] [Scilit]
- Finke, T.; Krämer, M.; Morandini, A.; Mück, A.; Oleksiyuk, I. Autoencoders for Unsupervised Anomaly Detection in High Energy Physics. J. High Energy Phys. 2021, 2021, 161. [Google Scholar] [CrossRef] [Scilit]
- Schneider, S.; Antensteiner, D.; Soukup, D.; Scheutz, M. Autoencoders-a Comparative Analysis in the Realm of Anomaly Detection. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), New Orleans, LA, USA, 19–20 June 2022. [Google Scholar]
- Hajimohammadali, F.; Fontana, N.; Tucci, M.; Crisostomi, E. Autoencoder-Based Fault Diagnosis for Hydropower Plants. In Proceedings of the 2023 IEEE Belgrade PowerTech, PowerTech 2023, Belgrade, Serbia, 25–29 June 2023. [Google Scholar] [CrossRef] [Scilit]
- Zhang, C.; Wang, X.; Zhang, J.; Li, S.; Zhang, H.; Liu, C.; Han, P. VESC: A New Variational Autoencoder Based Model for Anomaly Detection. Int. J. Mach. Learn. Cybern. 2023, 14, 683–696. [Google Scholar] [CrossRef] [Scilit]
- Kim, H.; Kim, H. Contextual Anomaly Detection for Multivariate Time Series Data. Qual. Eng. 2023, 35, 686–695. [Google Scholar] [CrossRef] [Scilit]
- Zhang, A.; Zhao, X.; Wang, L. CNN and LSTM Based Encoder-Decoder for Anomaly Detection in Multivariate Time Series. In Proceedings of the IEEE Information Technology, Networking, Electronic and Automation Control Conference, ITNEC 2021, Xi’an, China, 15–17 October 2021; pp. 571–575. [Google Scholar] [CrossRef] [Scilit]
- Lindemann, B.; Maschler, B.; Sahlab, N.; Weyrich, M. A Survey on Anomaly Detection for Technical Systems Using LSTM Networks. Comput. Ind. 2021, 131, 103498. [Google Scholar] [CrossRef] [Scilit]
- Van Leeuwen, R.; Koole, G. Anomaly Detection in Univariate Time Series Incorporating Active Learning. J. Comput. Math. Data Sci. 2023, 6, 100072. [Google Scholar] [CrossRef] [Scilit]
- Satyanarayana, K.; Venkatesh, K. IoT Univariate and Multivariate Time-Series Data Anomaly Detection: A Literature Review. In Proceedings of the Applications of Computational Intelligence in Management and Mathematics I, Arunachal Pradesh, India, 4–5 August 2025; Springer Proceedings in Mathematics and Statistics. Volume 492, pp. 249–264. [Google Scholar] [CrossRef] [Scilit]
- Belay, M.A.; Blakseth, S.S.; Rasheed, A.; Salvo Rossi, P. Unsupervised Anomaly Detection for IoT-Based Multivariate Time Series: Existing Solutions, Performance Analysis and Future Directions. Sensors 2023, 23, 2844. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chen, X.; Deng, L.; Huang, F.; Zhang, C.; Zhang, Z.; Zhao, Y.; Zheng, K. DAEMON: Unsupervised Anomaly Detection and Interpretation for Multivariate Time Series. In Proceedings of the 2021 IEEE 37th International Conference on Data Engineering (ICDE), Chania, Greece, 19–22 August 2021; pp. 2225–2230. [Google Scholar] [CrossRef] [Scilit]
- Wickramasinghe, A.; Muthukumarana, S.; Loewen, D.; Schaubroeck, M. Temperature Clusters in Commercial Buildings Using K-Means and Time Series Clustering. Energy Inform. 2022, 5, 1–14. [Google Scholar] [CrossRef] [Scilit]
- Yu, D.; Liu, G.; Guo, M.; Liu, X. An Improved K-Medoids Algorithm Based on Step Increasing and Optimizing Medoids. Expert Syst. Appl. 2018, 92, 464–473. [Google Scholar] [CrossRef] [Scilit]
- Cai, B.; Huang, G.; Samadiani, N.; Li, G.; Chi, C.H. Efficient Time Series Clustering by Minimizing Dynamic Time Warping Utilization. IEEE Access 2021, 9, 46589–46599. [Google Scholar] [CrossRef] [Scilit]
- Ran, X.; Xi, Y.; Lu, Y.; Wang, X.; Lu, Z. Comprehensive Survey on Hierarchical Clustering Algorithms and the Recent Developments. Artif. Intell. Rev. 2022, 56, 8219–8264. [Google Scholar] [CrossRef] [Scilit]
- Kulkarni, O.; Burhanpurwala, A. A Survey of Advancements in DBSCAN Clustering Algorithms for Big Data. In Proceedings of the 2024 3rd International Conference on Power Electronics and IoT Applications in Renewable Energy and its Control, PARC 2024, Mathura, India, 23–24 February 2024; pp. 106–111. [Google Scholar] [CrossRef] [Scilit]
- Raj, R.S.; Hema, L.K. Dynamic Clustering Optimization for Energy Efficient IoT Network: A Simple Constrastive Graph Approach. Expert Syst. Appl. 2025, 264, 125875. [Google Scholar] [CrossRef] [Scilit]
- Zavaleta-Sanchez, M.; Benitez-Guerrero, E.; Molero-Castillo, G.; Mezura-Godoy, C.; Gerardo Montane-Jimenez, L. A Framework for Dynamic User Modeling Integrating Data Stream Mining and Process Mining in Educational Contexts. IEEE Access 2025, 13, 166078–166103. [Google Scholar] [CrossRef] [Scilit]
- Novo, R.; Marocco, P.; Giorgi, G.; Lanzini, A.; Santarelli, M.; Mattiazzo, G. Planning the Decarbonisation of Energy Systems: The Importance of Applying Time Series Clustering to Long-Term Models. Energy Convers. Manag. X 2022, 15, 100274. [Google Scholar] [CrossRef] [Scilit]
- Fahrudin, N.F.; Rindiyani, R. Comparison of K-Medoids and K-Means Algorithms in Segmenting Customers Based on RFM Criteria. E3S Web Conf. 2024, 484, 02008. [Google Scholar] [CrossRef] [Scilit]
- Holder, C.; Middlehurst, M.; Bagnall, A. A Review and Evaluation of Elastic Distance Functions for Time Series Clustering. Knowl. Inf. Syst. 2024, 66, 765–809. [Google Scholar] [CrossRef] [Scilit]
- Microsoft What Is TDSP? Available online: https://www.datascience-pm.com/tdsp/ (accessed on 29 July 2023).
- Miraftabzadeh, S.M.; Colombo, C.G.; Longo, M.; Foiadelli, F. K-Means and Alternative Clustering Methods in Modern Power Systems. IEEE Access 2023, 11, 119596–119633. [Google Scholar] [CrossRef] [Scilit]
- Hamka, M.; Ramdhoni, N. K-Means Cluster Optimization for Potentiality Student Grouping Using Elbow Method. AIP Conf. Proc. 2022, 2578, 060011. [Google Scholar] [CrossRef] [Scilit]
- Lehmann, R. 3σ-Rule for Outlier Detection from the Viewpoint of Geodetic Adjustment. J. Surv. Eng. 2013, 139, 157–165. [Google Scholar] [CrossRef] [Scilit]










| Reference | ML Technique | Description |
|---|---|---|
| [25,26] | ARIMA | Captures temporal patterns using autoregressive and moving average components |
| [27,28,29,30] | Autoencoders | Neural network trained to reconstruct input data; anomalies are inferred from reconstruction error. |
| [22,31] | Variational Autoencoder (VAE) | Learns latent representations and reconstructs input; high reconstruction error suggests anomalies |
| [32,33,34] | LSTM | Recurrent neural network that learns long-term dependencies in time series. |
| Process | Activity | Methods | Quality Measure | Success Criteria |
|---|---|---|---|---|
| Anomaly Detection Model Creation | Anomaly-Free Data Preparation | Direct query from the database | Number of records retrieved | Records retrieved = Records in the database |
| Time series clustering |
|
|
| |
| Cluster Centroids Definition |
|
| Cluster centroid should be similar to the variables within its cluster | |
| Unsupervised Anomaly Detection Modeling |
|
| Prediction error < 10% of true values R2 > 0.8 | |
| Adaptive Anomaly Thresholding | Based on statistical analysis of model error distribution | Threshold per variable or cluster | Thresholds defined and validated for all monitored variables | |
| Real-Time Anomaly Detection | Real-time data acquisition | Direct reading from database | Number of records retrieved | Records retrieved = Records in the database |
| Anomaly detection per variable | Apply pre-trained models per variable | Not applicable | One anomaly detection model applied to each variable | |
| Anomaly score per variable | Anomaly identification based on thresholding | Not applicable | Anomaly score generated for each variable | |
| Anomaly logging in a database | Direct write to database; update if score already exists | Number of correctly inserted records | One score per variable successfully stored in the database |
| Date | Cluster 1 | Power |
|---|---|---|
| 13 September 2023 19:30 | 0.471516 | 0.836539 |
| 13 September 2023 19:31 | 0.471516 | 0.836539 |
| 13 September 2023 19:32 | 0.471516 | 0.836539 |
| 13 September 2023 19:33 | 0.471516 | 0.836539 |
| 13 September 2023 19:34 | 0.471516 | 0.836539 |
| Variable | Mean Error (μ) | Std. Deviation (σ) | Anomaly Threshold (2σ) |
|---|---|---|---|
| 1 | 0.006 | 0.014 | 0.029 |
| 2 | 0.005 | 0.011 | 0.022 |
| 3 | 0.030 | 0.002 | 0.004 |
| 4 | 0.004 | 0.006 | 0.012 |
| 5 | 0.023 | 0.032 | 0.064 |
| R2 Value | Number of Variables | Variables |
|---|---|---|
| R2 between 0.8–1.0 | 145 | - |
| R2 between 0.4–0.7 | 1 | 37 |
| R2 < 0.4 | 3 | 7, 32 and 44 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
Share and Cite
Oviedo, A.I.; Vargas, J.F.; Hincapie, R.C.; Molina, A.F.; Vergara, E.A.; Tello, D.M. Unsupervised Real-Time Anomaly Detection in Hydropower Systems via Time Series Clustering and Autoencoders. Technologies 2025, 13, 534. https://doi.org/10.3390/technologies13110534
Oviedo AI, Vargas JF, Hincapie RC, Molina AF, Vergara EA, Tello DM. Unsupervised Real-Time Anomaly Detection in Hydropower Systems via Time Series Clustering and Autoencoders. Technologies. 2025; 13(11):534. https://doi.org/10.3390/technologies13110534
Chicago/Turabian StyleOviedo, Ana I., John F. Vargas, Roberto C. Hincapie, Andres F. Molina, Edimerk A. Vergara, and Diana M. Tello. 2025. "Unsupervised Real-Time Anomaly Detection in Hydropower Systems via Time Series Clustering and Autoencoders" Technologies 13, no. 11: 534. https://doi.org/10.3390/technologies13110534
APA StyleOviedo, A. I., Vargas, J. F., Hincapie, R. C., Molina, A. F., Vergara, E. A., & Tello, D. M. (2025). Unsupervised Real-Time Anomaly Detection in Hydropower Systems via Time Series Clustering and Autoencoders. Technologies, 13(11), 534. https://doi.org/10.3390/technologies13110534

