1. Introduction
Power transformers play a key role in electrical power systems. They keep the electricity supply stable and manage voltage changes at different stages [
1,
2]. When these transformers fail, they can cause major economic and operational problems. Common failure modes include short-circuit failures, which are identified as the most critical, along with electrical, mechanical, chemical, and environmental stresses affecting various transformer components, including the core, windings, bushings, tank, and cooling systems [
3,
4]. Transformer failures result in power supply interruptions, high repair costs, revenue losses, and environmental damages [
5]. They are also expensive and vital parts of power systems [
6,
7], so they need strong protection, better monitoring, advanced diagnostics, and solid maintenance plans. These steps help manage risks, lower costs [
8], prevent damage, and keep systems reliable. There are many ways to check transformer health; DGA is one of the most common methods for spotting problems early and assessing their condition. The primary focus of this study is DGA, which falls under chemical diagnostics, and will remain the primary focus from now on. DGA measurements can be obtained through various techniques, such as gas chromatography, hydrogen online monitoring, and photo-acoustic spectroscopy [
9]. Several DGA interpretation methods have been used, including Key Gas, Rogers Ratio, IEC Ratio, Doernenburg, and Duval Triangle and Pentagon [
10].
Traditional DGA methods do not always get it right. Their diagnostic accuracy can be pretty low [
11], and they often give mixed results or even outright conflicting diagnoses when trying to figure out what is wrong with a power transformer [
12]. Methods like Doernenburg, Rogers, and the IEC ratios have accuracy rates ranging from just under 42% to about 77%, but even then, they tend to produce many confusing “out-of-code” results [
13,
14]. The problem is that these methods use fixed thresholds and just cannot keep up when fault scenarios get complicated or when gas formation behaves in nonlinear ways [
15]. Lately, though, researchers have started tackling these issues by adding Artificial Intelligence (AI) to the mix. Building on the old-school DGA techniques, more advanced statistical analysis is now helping to sharpen fault detection and set clearer, more reliable boundaries for diagnostics.
Lately, statistical methods for diagnosing power transformers with DGA have taken a big leap forward. Researchers have developed new visual tools like the Circle Method that do a much better job than older techniques at distinguishing overlapping fault types [
16]. There are also smarter statistical tests now that look for things like two peaks in gas data, giving early warning signs when something is wrong [
17]. Of course, it is not all easy; problems like small or uneven datasets still make things difficult, but adaptive over-sampling methods such as ASMOTE help fix that [
18]. Even with these improvements, many experts still suggest using DGA along with other diagnostic tools to get a fuller, more accurate picture of transformer health [
19]. Statistical methods are great for spotting patterns and understanding how gases behave, but ML goes further, making it possible to identify faults automatically and even predict what might go wrong next.
Machine Learning (ML) approaches have been proposed to integrate multiple interpretation techniques to enhance accuracy and consistency [
20]. Various approaches have been explored, including traditional supervised methods like Random Forest and Support Vector Machines (SVM) [
21], as well as advanced deep learning models such as Long Short-Term Memory (LSTM) [
22]. Researchers are getting creative with new model designs like N-Trans, which blends N-gram and Transformer techniques [
23], and they have also experimented with adversarial training to make detection systems tougher and more reliable [
24]. Recent studies have put various ML and Deep Learning (DL) models head-to-head, and interestingly, they have found that with carefully crafted features, top-performing ML methods can go toe-to-toe with the best DL models [
25].
Some researchers have focused on inline detection using deep networks [
26], while others have investigated the use of unsupervised character n-gram embeddings to enhance classifier performance, especially with limited training data [
27]. ML-based transformer diagnostics face several challenges, including the limited availability of abnormal-condition data, which impacts diagnostic performance [
28]. When it comes to fault detection and diagnosis in industry, issues like data quality, model interpretability, and integration remain major headaches [
29].
Among the machine learning techniques used for transformer diagnostics, decision trees stand out; they are popular because they are easy to understand and useful in practice. Time and again, decision tree models do a better job than the classic DGA interpretation methods like the Duval Triangle, Rogers Ratio, and IEC Ratio [
30], and in several studies also surpass other ML approaches including SVM and naive Bayes [
31]. Ensemble extensions further improve performance, with CatBoost achieving 96.78% accuracy [
32] and the KosaNet ensemble reaching 99.9% [
33].
Results vary a lot from one study to another, mostly because of factors like the size of the datasets (from just 286 samples to nearly 3000) [
34], how many fault types are included (sometimes up to nine different groups) [
35], which features are chosen [
36], and how carefully the data is cleaned and prepared. Still, one thing is clear: decision trees are often praised for being easy to understand; they create simple, helpful rules that make work much easier for engineers [
37,
38].
Decision trees are great for clear and effective fault classification, but you can get even better results by combining them with other diagnostic and statistical methods. Hybrid techniques for power transformer diagnostics using DGA have shown improved accuracy over conventional methods. These approaches combine traditional DGA methods with artificial intelligence techniques like fuzzy logic, neural networks, and SVM [
39]. Evolutionary algorithms and clustering methods have been used to enhance feature selection and classification [
40]. Some hybrid models integrate multiple DGA interpretation methods and employ soft computing for better fault diagnosis. Even though these hybrid techniques usually perform better than the old-school methods, they are not perfect; large datasets are still a must, and there is always a risk of overfitting [
41].
Although transformer fault diagnosis has improved over the years, common DGA methods like the Key Gas Method, Rogers Ratio, and Duval Triangle still often fail in real situations. These methods usually do not work well when faults overlap or when gas values differ from expected patterns. This happens mainly because they use fixed limits and simple ratios, which do not capture the complex, changing behaviour of gases in transformers. Hybrid methods that combine AI and DGA have shown promise, but they often miss important statistical details like how data changes, links between data points, confidence levels, data patterns, percentiles, and data imbalance. Using ML has made diagnosis more accurate, but these methods are often difficult to understand and need a lot of data. They also still have problems like overfitting and being hard to explain. Even with progress in hybrid diagnostic systems, big challenges remain in making results clear, including the statistical details, and ensuring reliable performance in real-world situations. To address these issues, a hybrid transformer fault classification system is introduced. This system creates rule-based fault categories by analysing real DGA data using statistical methods and tests these rules with a Decision Tree model. The framework was developed and validated using 33,778 DGA records collected over five years from actual operating conditions. The objectives of the study are to:
Determine statistically significant DGA patterns to establish data-driven diagnostic thresholds.
Establish rule-based categories for transformer faults by applying statistically determined gas concentration thresholds.
Assess the capacity of the Decision Tree model to learn and generalise statistically derived fault classification rules.
Evaluate the relative importance of DGA variables in transformer fault classification.
The main contribution of this study is a combined fault-classification system that uses statistical DGA diagnostic rules together with Decision Tree learning and tests how well they work using 33,778 real-world operational records.
Unlike traditional methods that depend only on set diagnostic rules or ML, this study introduced a Statistically Guided Decision Tree (SGDT) framework that uses statistical DGA limits together with decision tree learning to better classify transformer faults using real operational data.
The rest of the study is structured as follows.
Section 1 provides a detailed overview of the relevant work;
Section 2 outlines the materials and methods used; and
Section 3 presents the proposed hybrid fault classification method combining rule-based logic and decision tree modelling.
Section 4 summarises key statistical and model results.
Section 5 discusses the key findings and future directions.
Section 6 concludes with the study’s main contributions.
3. Proposed SGDT Framework for Transformer Fault Classification
Figure 1 shows the schematic diagram detailing the process utilised in this study to build the proposed SGDT framework, which integrates advanced statistical analysis with decision tree learning using DGA data from a single power transformer. Using one transformer made it possible to closely study how faults develop and how gas behaves under steady working conditions, reducing differences caused by transformer design, load, maintenance, and environment. The results show that the method works well on a real monitored transformer. However, testing with data from many transformers is valuable to consider in the future.
The dataset was collected directly from an online DGA monitoring system over five years and kept in its original form to keep the real-world measurement details. Small negative values showed up in a few gas measurements, mainly for acetylene (C2H2), with rare cases in H2, CH4, CO, and CO2. These values come from measurement uncertainty, sensor baseline shifts, calibration changes, and limits in analysis when gas levels are close to the detection limit. Although these values are not physically negative gas amounts, they are actual readings from the monitoring equipment during operation. So, the values were kept to ensure the fault classification model was built and tested using realistic field data that reflects real use conditions.
The work began with identifying the transformer and collecting its dissolved gas records. Following this was data validation; the relevant gases were extracted, and TDCG values and generation rates were calculated. Descriptive statistical analysis was then performed to understand gas behaviour patterns and establish data-driven thresholds for transformer fault identification. Based on these statistical findings, fault classification rules were developed and used to assign fault labels to the dataset. The labelled dataset was then used to train and test the decision tree model. Traditional DGA diagnostic methods, like the Duval Triangle, provided background and comparison points.
3.1. Selection of Statistical Estimation Methods
To make sure the description of the DGA dataset is accurate, several ways to estimate average values and spread were tested.
Table 1 shows the methods looked at for each statistic, highlights the chosen method, and explains why based on features of the data like unusual values and uneven distributions. The bootstrap method was selected as the preferred approach in all cases due to its robustness and independence from parametric assumptions.
3.2. Classification Rule Application
This study created guidelines to identify transformer problems by looking at gas data from transformer samples. These rules set clear limits to identify specific fault types using summary statistics and gas patterns found in the data. Summary statistics were calculated for all variables, including main combustible gases and other gases like CO
2, O
2, H
2O, TDCG, and generation rates. Threshold values were set for each variable to describe their statistical patterns. However, only six gases (H
2, CH
4, C
2H
2, C
2H
4, C
2H
6, and CO) were selected to establish fault classification rules.
Table 2 describes each transformer fault type based on specific gas concentration thresholds. These thresholds were chosen using descriptive statistics such as the median, Interquartile Range (IQR), 75th percentile, 95th percentile, and maximum, depending on how each gas behaves in the dataset.
For example, C2H2 levels are usually very low, so values above the 95th percentile are seen as abnormal. H2 is found more often, so using limits based on the middle value or spread is better. Gases linked to heat damage, like C2H4 and C2H6, usually rise slowly, so limits set between the middle spread and the 95th percentile help show how serious the fault is. CO, often linked to insulation ageing, is checked using values between the 75th and 95th percentiles. These limits were chosen based on how each gas is spread out and its known connection to certain fault types, helping to make reliable rule-based decisions. Fault rules were made only for the six main gases to keep the diagnosis relevant. In contrast, when training the ML model, all available variables, including those excluded from rule development, were incorporated as features to enable the model to identify additional patterns and enhance fault prediction robustness.
Although eight fault categories were defined through the statistical rule-development process, only six fault classes (D1, PD, T1, T2, Mixed, and Normal) were represented in the classified dataset. The D2 (High Discharge) and T3 (High Thermal) categories were defined based on extreme threshold combinations from the upper tails of gas concentration distributions. Since no observations in the operational dataset met these criteria, only the six represented fault classes were included in the decision tree modelling and performance evaluation.
Rule 1 flags a high discharge fault when C2H2 goes above 1.0 ppm, H2 is over 200 ppm, and CO is higher than 90 ppm. These values are well beyond what is usually seen and point to strong arcing inside the transformer. Rule 2 suggests a low discharge if C2H2 is above 0.5 ppm, CH4 is over 70 ppm, and H2 is more than 150 ppm. This combination indicates a less intense but still abnormal discharge. In Rule 3, partial discharge is likely when H2 is greater than 300 ppm while both CH4 and C2H2 stay low, which is a typical pattern for ionisation or corona activity.
Rule 4 indicates high thermal stress when C2H4 and C2H6 both exceed their usual limits, and CO also rises, all pointing to overheating. Rule 5 describes a medium thermal fault if C2H4, C2H6, and CO fall between common and extreme values, meaning there is heating but not at critical levels. Rule 6 is used when C2H4, C2H6, and C2H2 remain low, which matches what is expected from minor heating. Rule 7 covers cases where two or more gases from the group C2H2, H2, CO, CH4, C2H4, and C2H6 are elevated at the same time, showing signs of both thermal and electrical faults happening together. Rule 8 is used when all gas levels are normal, showing that the transformer is working without any problems or damage. Stray gassing was not listed as a separate category in the dataset, so it was not studied. Future work will add labelled stray-gassing cases to see whether the SGDT system can distinguish stray gassing from actual transformer faults.
3.3. Decision Tree Model Configuration
The decision tree model was created to check if the gas concentration rules for fault labelling matched real data patterns. The goal was not to diagnose transformer faults on its own but to see if the fault classification rules based on statistics could be reliably learned and applied to new data. A dataset of 33,778 records was split into training (64%), validation (16%), and test (20%) parts. The model was set with a maximum depth of 4, needed at least 20 records to make a split, and required at least 15 records in each end node. Feature scaling and a fixed random seed were used to keep results consistent. A complexity penalty of 0.1 was added to reduce overfitting. The full model reached 98% accuracy on the test data with 119 splits. A simpler decision tree with a maximum depth of 3 was also made to help with visualisation. This simpler model made it easier to understand while keeping important splits based on gases like C2H2, H2, CH4, and C2H4, matching the original rule-based system.
Table 3 presents decision tree model performance across multiple random-seed iterations. To assess the robustness of the proposed Decision Tree model against variations in data partitioning, the model was evaluated over ten iterations using different random seeds. The average Accuracy, Precision, Recall, F1-score, and AUC were 0.9935 ± 0.0005, 0.9820 ± 0.0019, 0.9810 ± 0.0019, 0.9812 ± 0.0017, and 0.9865 ± 0.0021, respectively. The low standard deviations indicate that the model exhibits stable performance and good generalisation capability across different dataset partitions.
6. Conclusions
The proposed SGDT framework, which combined advanced statistics applications together with decision tree learning, was accurate and reliable at identifying faults with real-world DGA transformer data. The results reached an accuracy of 99.3%, precision of 0.981, recall of 0.980, F1-score of 0.980, and AUC of 0.987, showing excellent prediction, balanced classification, and strong separation between fault types. Testing with 33,778 operational DGA records showed that this method provided a simple, statistically based, and widely useful tool for checking transformer condition. The results showed that the method improved fault detection, helped spot problems early, and supported better transformer health management and maintenance decisions.
Although the hybrid model worked well, a few issues remain and should be explored in future work: first, negative gas values were found in the data. This shows a need for better data checks before analysis. Secondly, only linear relationships were studied in this work. Therefore, more advanced models may help detect complex fault signs. Thirdly, the method was tested on one dataset; therefore, the proposal is that the same method should be applied to more transformers under different conditions, such as age, load history, geographical location, and maintenance practices. Furthermore, future work will consider stray gassing cases, and other parameters that may be effective for fault detection.