Next Article in Journal
Do High-Possession Teams Have Greater Reported Injury Burden? A Machine Learning Cross-League Analysis of Europe’s Big Five Football Competitions
Previous Article in Journal
A Transformation-Based AI Framework for Equitable RTI Tier Classification in Qatar
Previous Article in Special Issue
A Review of Human-AI Complementarities Across Multiple Dimensions of Organisational Complexity
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Systematic Review

A Comprehensive Systematic Meta-Survey of Energy Theft Detection: From Traditional Methods to Generative AI

Department of Computer Engineering, Modeling, Electronics and Systems (DIMES), University of Calabria, 87036 Rende, Italy
*
Author to whom correspondence should be addressed.
Big Data Cogn. Comput. 2026, 10(9), 297; https://doi.org/10.3390/bdcc10090297
Submission received: 23 June 2026 / Revised: 13 August 2026 / Accepted: 27 August 2026 / Published: 2 September 2026
(This article belongs to the Special Issue Big Data and Cognitive Computing in 2026)

Abstract

Electrical energy theft presents a serious and significant challenge for utility companies worldwide. It poses substantial risks to energy infrastructure, reduces efficiency, and destabilizes distribution networks. The introduction of the Advanced Metering Infrastructure (AMI) in the last decade has significantly boosted the development of new methods and techniques for detecting electrical energy theft. This advancement is primarily due to the availability of electric consumption and other data with greater granularity (e.g., every 15 min), which has enabled the development of analytic and data-driven approaches as opposed to mass inspections alone. In this context, this paper aims to analyze the landscape of energy theft detection by providing a systematic analysis of the state-of-the-art manuscripts in the form of a tertiary study (i.e., a review of literature reviews and surveys) in accordance with the PRISMA methodology guidelines. Consequently, a comparative framework is presented, along with well-formulated research questions designed to explore the past, present, and future directions of energy theft detection.

1. Introduction

1.1. Background and Motivation

Electrical energy theft is a major form of non-technical loss (NTL) and remains a persistent challenge for electricity utilities worldwide. It causes direct financial losses and can adversely affect the reliability and efficiency of distribution networks by distorting the consumption information available to utilities. In particular, inaccurate consumption measurements can affect demand estimation, load forecasting, energy balancing, and distribution-network planning [1]. Global estimates have placed the economic impact of electricity theft at tens of billions of dollars annually, highlighting the scale and persistence of the problem [2,3]. The problem is particularly relevant in developing and rapidly digitalizing electricity systems, where high levels of NTL can coexist with heterogeneous infrastructure, limited monitoring capabilities, and diverse patterns of fraudulent consumption [4,5].
NTL encompasses losses that are not attributable to the physical loss of electrical energy in network components. Unlike technical losses, which arise from the physical characteristics of distribution lines and transformers, NTL includes electricity theft, unauthorized connections, meter tampering, billing irregularities, and other forms of unmetered or incorrectly recorded consumption [6]. Although some components of NTL can be reduced through improved metering and administrative procedures, electricity theft remains particularly challenging because fraudulent behavior may be intermittent, partial, or deliberately designed to avoid conventional detection mechanisms [7].
Traditional theft-detection practices have relied heavily on the physical inspection of customer premises. Such inspections require substantial human, financial, and operational resources and become increasingly difficult to scale when distribution system operators (DSOs) serve large numbers of customers. Moreover, physical inspection may not reliably identify intermittent or partial theft. For example, a customer may legitimately consume part of the electricity through the official meter while diverting another part through an alternative electrical connection. In such cases, automated methods can provide valuable risk indicators before an inspection is conducted. In several regulatory environments, including Italy, however, an on-site inspection remains necessary to formally certify the occurrence of fraud. Consequently, the practical role of automated NTL-detection systems is not necessarily to replace physical inspection, but rather to support risk-based selection and prioritization of inspection campaigns [7,8]. The trend of DSOs and electricity operators is towards a digital transition that will bring the gradual worldwide replacement of traditional measuring devices with smart meter systems, as shown in Figure 1. This will lead to an increasing use of data-driven systems for detecting electricity fraud. Evidence of this is the presence in recent years of articles found in search databases increasingly oriented toward data-driven detection systems and solutions.
The increasing digitalization of distribution networks has substantially changed this scenario. The deployment of AMI enables electricity consumption to be recorded at much finer temporal resolutions than traditional manual meter-reading systems, thereby creating opportunities for analytical and data-driven detection approaches [10]. AMI also enables bidirectional communication between smart meters and utility systems and supports functions such as remote meter management, supply activation and deactivation, tariff management, and monitoring of meter status [11]. These capabilities provide utilities with richer information for identifying abnormal consumption patterns and potential electricity theft.
At the same time, the digitalization of metering infrastructure introduces new vulnerabilities. Smart meters, gateways, concentrators, communication channels, and associated software can become targets for cyberattacks or data manipulation [12]. Consequently, future NTL detection cannot be considered exclusively as a problem of identifying physical meter tampering. Increasingly digital distribution systems create a cyber-physical threat landscape in which conventional physical theft may coexist with the manipulation of measurements, communication data, or AMI infrastructure. False-data injection, measurement manipulation, communication disruption, and compromised smart-meter information therefore represent emerging issues that should be considered when defining the future scope of NTL detection.

1.2. Data-Driven NTL Detection

The increasing availability of high-resolution AMI data has stimulated extensive research into data-driven NTL detection [13]. Depending on data availability and the assumed level of grid digitalization, researchers have investigated statistical methods [14,15], supervised and unsupervised machine learning [16,17,18,19,20,21,22], anomaly detection [4,23], deep learning [24,25,26,27], and hybrid [25,28] or ensemble [20,21,29,30] approaches. These approaches generally attempt to distinguish legitimate consumption behavior from anomalous or potentially fraudulent patterns using historical consumption measurements and, in some cases, additional meter or network information.
The effectiveness of a particular detection paradigm, however, depends strongly on its underlying assumptions. Supervised approaches can provide strong predictive performance when sufficiently labeled examples of electricity theft are available, whereas unsupervised or anomaly based methods are attractive when confirmed theft labels are scarce [31,32]. Deep-learning methods can learn complex temporal and behavioral representations from high-resolution consumption profiles, but they may require substantial computational resources and sufficiently representative training data. Hybrid and ensemble approaches attempt to combine complementary detection capabilities but may increase model complexity.
Reported performance across these approaches should therefore be interpreted carefully. The secondary literature uses heterogeneous datasets, geographical contexts, theft-generation mechanisms, class distributions, and evaluation protocols. Consequently, a high accuracy or F1-score reported on one benchmark cannot automatically be interpreted as evidence of general superiority across utilities or regions. Practical NTL detection must additionally consider false-positive inspection costs, scalability, interpretability, computational and communication requirements, privacy, concept drift, and the ability to generalize to previously unseen theft behaviors.
The evolution of AMI also motivates a broader transition from centralized data analysis toward edge, distributed, and potentially cooperative intelligence. Neighboring smart meters or edge devices could potentially exchange selected features, anomaly scores, model information, or state estimates through local communication while retaining raw consumption data locally. Such cooperative architectures may reduce dependence on a centralized control center and support privacy-preserving and scalable detection. Nevertheless, the extent to which multi-agent and distributed-control mechanisms can improve NTL detection remains insufficiently established in the existing secondary evidence.
Generative AI represents another emerging direction [8,33,34,35,36,37]. Generative models may potentially support synthetic consumption-profile generation, minority-class augmentation, representation learning, and privacy-preserving data generation in situations where genuine theft observations are limited. However, the representation of Generative AI in the existing secondary literature remains substantially smaller than that of established ML, DL, anomaly-detection, and hybrid approaches. It should therefore be considered an emerging research opportunity rather than an established NTL-detection paradigm. This distinction is important when interpreting the current evidence and defining future research priorities.

1.3. Existing Reviews and the Secondary Literature

The rapid growth of research on NTL detection has resulted in a substantial body of surveys, systematic reviews, and other secondary studies. These studies have provided valuable syntheses of primary research from different perspectives. For example, Savian et al. conducted a systematic review of 121 articles addressing the worldwide NTL landscape, including its impacts, barriers, mitigation strategies, and regulatory aspects [38]. Ahmed et al. proposed a taxonomy and comparative analysis of energy-theft detection techniques, including data-mining, state/network-based, and game-theoretic approaches [37]. Stracqualursi et al. reviewed energy-theft practices and AI-based autonomous detection, considering theft types, AI methodologies, data-collection issues, and challenges related to generalized detection [39]. Other reviews have specifically examined AI-based NTL detection according to algorithms, features, evaluation metrics, and NTL categories [40]. More recent reviews have considered data-driven electricity-theft detection and emerging Generative-AI-based approaches, including challenges related to datasets and synthetic data generation [33].
These studies demonstrate that the NTL-detection literature has already been extensively surveyed. However, the existing secondary literature is heterogeneous in scope, terminology, taxonomy, methodological emphasis, data assumptions, infrastructure requirements, and evaluation practices. Some reviews primarily classify theft types, whereas others focus on machine-learning or deep-learning algorithms. Some emphasize datasets and feature engineering, while others address AMI security, privacy, anomaly detection, or broader smart-grid perspectives.
This heterogeneity makes direct comparison across reviews difficult. For example, different reviews may classify the same detection method under different methodological categories, emphasize different performance metrics, or reach different conclusions because their underlying primary studies use different datasets and experimental assumptions. Similarly, research gaps identified in one review may receive limited attention in another. Therefore, although individual secondary studies provide valuable methodological summaries, they do not necessarily establish what the secondary literature collectively agrees upon, where it differs, and which research dimensions remain insufficiently addressed.
Table 1 summarizes the distinction between representative existing secondary studies and the present work.
The important distinction is that the previous reviews primarily synthesize primary NTL-detection studies, whereas the present study treats the secondary studies themselves as the units of analysis. Thus, the present work operates at a higher level of evidence synthesis.

1.4. Research Gap

The preceding analysis identifies four principal gaps in the existing secondary literature. First, there is no common analytical framework through which the heterogeneous secondary studies can be directly compared. Existing reviews use different classifications of attack types, detection paradigms, data requirements, and infrastructure assumptions. A common framework is therefore necessary to harmonize their findings. Second, the existing reviews provide limited cross-survey synthesis. Although individual reviews summarize primary studies, they do not systematically determine which findings are consistently reported across reviews, which findings are dependent on particular datasets or methodological assumptions, and where the reviews disagree. Third, the secondary literature has not been sufficiently examined in terms of methodological quality and evidence strength. Reviews cannot automatically be treated as equivalent sources of evidence because they differ in search transparency, eligibility criteria, screening procedures, synthesis methodology, and reporting of limitations. A tertiary synthesis should therefore consider not only the findings reported by secondary studies but also the methodological quality of those studies. Fourth, there is limited quantitative characterization of research coverage across the secondary literature. It is important to determine not merely whether a topic has been discussed, but how consistently it is represented across the available reviews. Such analysis can reveal highly investigated dimensions as well as methodological blind spots and persistent gaps.
To address these issues, the present study conducts a systematic tertiary review of 28 eligible secondary studies. The study does not attempt to reproduce another method-level review of individual NTL-detection algorithms. Instead, it systematically maps and compares the existing secondary evidence using a common framework covering attack types, detection approaches, data and feature requirements, AMI/infrastructure dependence, evaluation practices, and challenges and research gaps.
The study also distinguishes between the 28 core secondary studies, which constitute the evidence base of the tertiary synthesis, and additional primary/contextual studies cited selectively to illustrate recent developments, datasets, emerging technologies, or specific research directions. These supplementary studies are not treated as units of the tertiary analysis and are not included in the quantitative cross-survey statistics.
Based on this scope, the central objective of the study is to answer the five research questions: (1) RQ1: What types of NTL/electricity-theft attacks are addressed in the existing secondary literature? (2) RQ2: What detection approaches and methodological paradigms are reported across the secondary studies? (3) RQ3: What data, features, and AMI/infrastructure requirements are associated with the reported detection approaches? (4) RQ4: How are NTL-detection approaches evaluated, and to what extent are the reported evaluation practices comparable? (5) RQ5: What challenges, limitations, and research gaps are consistently or inconsistently identified across the secondary literature?

1.5. Contributions

The main contributions of this study are as follows:
  • The study provides a PRISMA-based tertiary review of 28 eligible surveys, systematic reviews, and comprehensive reviews on NTL/electricity-theft detection, treating the secondary studies themselves as the units of analysis.
  • The heterogeneous findings of the included secondary studies are mapped to a common framework covering attack types, detection approaches, data and features, AMI/infrastructure requirements, evaluation practices, and challenges/research gaps.
  • The methodological quality of the 28 secondary studies is systematically assessed using predefined criteria addressing review objectives, search transparency, eligibility criteria, screening procedures, synthesis methodology, and reporting of limitations.
  • The study goes beyond a study-by-study summary by systematically identifying areas of agreement, divergence, and context-dependent findings across the secondary literature, particularly with respect to assumptions, methods, datasets, evaluation metrics, limitations, and practical relevance.
  • The study quantifies the coverage of the principal research dimensions across the 28 secondary studies, thereby identifying well-established themes as well as underrepresented and persistent research gaps.
  • The synthesis identifies future research priorities concerning standardized benchmarking, real-world validation, privacy-preserving and distributed detection, cyber-physical NTL and false-data injection threats, cooperative multi-agent detection, and emerging Generative-AI approaches. Emerging technologies are interpreted according to the strength and extent of the available evidence rather than assumed to be established solutions.
  • Based on the cross-survey evidence, the study provides recommendations for selecting and evaluating NTL-detection approaches according to data availability, AMI maturity, computational and communication constraints, privacy requirements, and operational deployment conditions.
Overall, the contribution of this work is not the proposal of another NTL-detection algorithm, but the higher-level synthesis and critical interpretation of the existing secondary evidence. By integrating systematic study selection, quality assessment, common evidence mapping, cross-survey comparison, and quantitative coverage analysis, the study provides a consolidated view of the current NTL-detection research landscape and identifies directions for developing more reliable, scalable, explainable, privacy-preserving, and deployment-oriented solutions.
The remainder of the manuscript is organized as follows. Section 2 presents the background and related secondary literature. Section 3 describes the research questions and methodology. Section 4 presents the tertiary synthesis and results, while Section 5 provides the cross-survey analytical synthesis. Section 6 discusses the major findings and recommendations. Section 7 discusses the threats to validity and limitations of the study, and Section 8 concludes the manuscript.

2. Background and Related Secondary Studies

2.1. Non-Technical Losses and Electricity Theft

In an electrical grid, losses are generally categorized into “technical” and “non-technical” losses (NTLs) [6]. Technical losses arise from the physical characteristics of the power system, including energy loss in distribution lines and transformer-related losses such as winding (copper) and magnetic core losses. Although these losses can be reduced through appropriate network planning, equipment selection, and operational measures, they are inherent to the physical transmission and distribution of electricity [54]. In contrast, non-technical losses represent avoidable losses that are primarily associated with human, administrative, and fraudulent activities. They include electricity theft, billing errors, metering irregularities, and other forms of unauthorized or inaccurate energy accounting [7].
Electricity theft constitutes one of the most significant and operationally challenging components of NTL. It can involve deliberate meter tampering, unauthorized connections, bypassing of meters, or manipulation of consumption measurements. The introduction of smart meters and AMI has reduced some conventional sources of NTL, particularly manual meter-reading and billing errors, by enabling automated collection and transmission of consumption information to utility control systems. Nevertheless, the digitalization of metering infrastructure has also introduced new security and data-integrity concerns, which are discussed in Section 2.4.
The detection of NTL remains challenging because the legal confirmation of electricity theft in many jurisdictions may require an on-site inspection and documented verification in the presence of the customer. Given the large number of customers served by modern distribution utilities, exhaustive inspection campaigns are costly and impractical. Consequently, automated detection and risk-based prioritization can be used to identify a smaller set of customers for targeted inspection [7]. This operational requirement has motivated extensive research into statistical, machine-learning, deep-learning, and other data-driven approaches for NTL detection.

2.2. AMI and Data-Driven Detection

The transition from conventional electricity metering to a smart grid (SG) and AMI has fundamentally changed the availability and granularity of data for NTL detection. Traditional metering systems generally provide periodic or aggregated consumption readings, whereas smart meters can provide fine-grained consumption measurements at substantially higher temporal resolutions. This increased data granularity enables utilities and researchers to analyze temporal consumption patterns and identify deviations that may indicate fraudulent or abnormal electricity use [35,37,39,42].
AMI also establishes bidirectional communication between smart meters and utility systems, enabling remote collection of consumption data, meter monitoring, and other operational functions [8,33,36]. The resulting availability of high-resolution consumption profiles has supported the development of statistical, machine-learning, deep-learning, anomaly detection, and hybrid approaches for NTL detection. Depending on the availability of labeled theft observations, these approaches may be formulated as supervised classification, unsupervised anomaly detection, or hybrid learning problems.
The data characteristics available through AMI strongly influence the design and performance of NTL detection systems. Commonly used information includes historical energy consumption and temporal load profiles, while additional customer or contextual attributes may include geographical information, tariff category, number of phases, payment behavior, and meter-reading characteristics [8,40,46]. Feature engineering is particularly important for conventional machine learning approaches, where statistical and temporal characteristics such as average, maximum and minimum consumption, standard deviation, power-factor information, and changes between current and historical consumption may be explicitly extracted [55,56]. Deep-learning approaches can instead learn representations directly from consumption sequences, reducing the dependence on manually designed features [40,43].
Despite the advantages provided by AMI, the resulting detection problem remains challenging. Real electricity-theft observations are typically scarce compared with legitimate consumption records, resulting in substantial class imbalance. Consequently, many studies consider resampling, cost-sensitive learning [57], synthetic data generation, or generative models to improve the representation of the minority theft class. However, synthetic data may not accurately represent evolving theft behaviors or consumption patterns from different geographical regions, and generative approaches can introduce additional computational and reproducibility challenges [32,33,58]. These issues highlight the importance of evaluating data-driven NTL detection approaches beyond predictive performance and considering data availability, representativeness, generalization, and practical deployment constraints.
The emergence of AMI therefore represents both an opportunity and a challenge for NTL detection. While increasingly granular data enable more sophisticated analytical approaches, the effectiveness of these approaches depends on the availability, quality, representativeness, and security of the underlying data. The specific datasets, features, data-balancing strategies, and evaluation practices reported across the secondary literature are systematically compared in the tertiary synthesis presented in Section 2.6.

2.3. Existing Surveys and Systematic Reviews

The increasing availability of smart-meter data and the rapid development of machine learning techniques have led to a growing body of secondary literature on NTL and electricity-theft detection. These studies differ considerably in their scope, objectives, methodological emphasis, and level of coverage. Based on the secondary literature examined in this study, existing reviews can broadly be grouped into three categories: Early and general NTL reviews, AI- and data-driven reviews, and recent specialized reviews addressing emerging technological and security perspectives.
Early and general NTL reviews. Early secondary studies primarily aimed to establish the foundations of NTL detection by discussing the nature and causes of electricity theft, major attack types, conventional detection strategies, and the broader economic, regulatory, and operational context. Some reviews provide relatively broad coverage of NTL identification, while others concentrate on specific detection techniques or particular application domains [8,37,43]. These studies established important taxonomies of electricity-theft practices and detection strategies and highlighted the continuing importance of physical inspection, hardware-based mechanisms, and network-oriented approaches. However, their methodological classifications and terminology are not uniform, making direct comparison across reviews difficult.
AI- and data-driven reviews. The increasing deployment of AMI and the availability of fine-grained electricity-consumption data have shifted the focus of subsequent reviews toward data-driven detection [13]. These studies examine machine-learning and deep-learning approaches, including classification, anomaly detection, convolutional and recurrent neural networks, hybrid models, and ensemble techniques [33,34,35,36,40]. Several reviews additionally consider the datasets, features, evaluation metrics, and data-imbalance strategies used in NTL detection [59,60,61]. This body of literature demonstrates the growing importance of consumption-based analytics, while also highlighting limitations associated with the availability of labeled theft data, class imbalance, generalization, and deployment conditions.
Recent specialized reviews. More recent secondary studies have expanded the scope beyond conventional consumption-based detection. These reviews address topics including smart meter and AMI security, privacy-preserving detection, adversarial attacks, limited-data learning, and emerging generative models. For example, recent studies consider the vulnerability of NTL detection systems to evasion and poisoning attacks, while other works investigate generative models for addressing class imbalance and incomplete consumption sequences. Generative approaches such as TimeGAN and diffusion-based methods represent an emerging direction for synthetic consumption-data generation and data augmentation [32,58]. However, the available evidence remains substantially smaller than that for established machine-learning and deep-learning approaches. Accordingly, Generative AI is considered in this review as an emerging research direction rather than as an established dominant paradigm.
The reviewed secondary literature therefore demonstrates a clear evolution from general NTL taxonomies and conventional detection mechanisms toward data-driven, AI-enabled, and increasingly cyber-physical perspectives. At the same time, the reviews remain fragmented across different analytical dimensions. Some emphasize attack types and physical theft mechanisms, others focus on machine-learning and deep-learning algorithms, while additional studies concentrate on datasets, AMI infrastructure, privacy, adversarial robustness, or emerging generative approaches. This heterogeneity creates an important limitation in the existing secondary literature: although individual reviews synthesize substantial bodies of primary research, their findings cannot be directly compared without a common analytical framework. Differences in taxonomy, data requirements, infrastructure assumptions, evaluation metrics, and reported research gaps make it difficult to determine which findings are consistent across reviews and which are dependent on particular methodological or application contexts.
Therefore, rather than conducting another method-level survey of primary NTL-detection studies, the present work performs a higher-level synthesis of the existing secondary evidence. The 28 eligible secondary studies constitute the core evidence base of the tertiary review and are systematically mapped according to attack types, detection approaches, data and feature requirements, AMI/infrastructure dependence, evaluation practices, and reported challenges and research gaps. Additional primary or contextual studies cited elsewhere in the manuscript are used only to illustrate specific developments and are not treated as units of the tertiary analysis.

2.4. Emerging Cyber-Physical Threats

The increasing digitalization of electricity distribution networks and the widespread deployment of AMI have expanded the attack surface associated with electricity measurement and NTL detection [12]. While a substantial proportion of electricity theft continues to involve physical manipulation of meters or unauthorized connections, the integration of smart meters, communication networks, data concentrators, and utility information systems introduces additional cyber and cyber-physical attack vectors. These attacks may affect either the measurement process itself or the transmission and processing of consumption information. The reviewed literature therefore suggests that future NTL detection systems should consider both physical and cyber-enabled manipulation rather than treating electricity theft as solely a consumption-classification problem.

2.4.1. False Data Injection Attacks

False Data Injection Attacks (FDIAs) represent an emerging cyber threat in which an adversary deliberately modifies information transmitted to or processed by the utility system. In the context of AMI, an attacker may manipulate reported electricity consumption or other measurement information so that the utility receives values that differ from those actually measured by the smart meter. The reviewed literature describes false-data injection as a mechanism through which an adversary can attempt to bypass existing detection mechanisms and generate erroneous measurements [8,37].
The communication architecture of AMI provides several potential points at which such manipulation may occur. Consumption information may be transmitted through communication technologies such as power-line communication and subsequently through wireless or optical wide-area networks. Vulnerabilities in these communication channels can therefore be exploited to alter information before it reaches the utility’s central systems. An adversary may also masquerade as a legitimate communication or control device with authorized access to smart meters [8]. Such attacks are particularly relevant to NTL detection because the fraudulent information may concern the value used for billing or automated decision-making rather than the physical consumption itself.
An important distinction identified in the reviewed literature is whether the manipulation changes the value stored by the meter or only the value transmitted to the utility. In a communication-level attack, the physical measurement stored in the smart meter may remain correct while an altered value is transmitted to the utility. Consequently, a subsequent physical inspection may reveal the actual consumption value, whereas manipulation of the meter’s measurement circuitry may permanently alter the recorded value [8]. This distinction has important implications for both automated detection and subsequent physical verification.

2.4.2. Measurement Manipulation

Measurement manipulation concerns attacks that directly alter the measurement process or the information stored within the metering device. Such manipulation may involve modifications to signal conditioning circuits, electronic memory, or the firmware of the smart meter [8,37,42]. Physical interference with the measurement circuitry can also reduce the amount of energy recorded by the meter. For example, the reviewed literature describes the use of an external magnetic field to disturb the amperometric measurement circuitry and thereby reduce the measured energy [8,39,46].
Measurement manipulation therefore represents an important bridge between conventional physical theft and emerging cyber-physical threats. Unlike attacks that modify information only during communication, manipulation at the measurement layer can cause the incorrect value to be stored by the meter itself. Consequently, the result may remain even when the meter is subsequently inspected.
The literature also reports more sophisticated forms of physical tampering in which access to the internal components is obtained without necessarily removing the meter from its housing. Such techniques may be used to disable anti-tamper sensors or subsequently modify internal measurement components [39]. These observations indicate that NTL detection cannot rely exclusively on consumption-data analytics and should, where appropriate, incorporate meter-level integrity and tamper information.

2.4.3. AMI Communication Attacks

AMI introduces bidirectional communication between smart meters and utility systems, thereby enabling remote monitoring and management but also creating communication-related security risks. The reviewed literature identifies man-in-the-middle attacks as one example in which an adversary can intercept a communication and transmit an altered consumption value instead of the actual reading [46]. Similarly, an attacker may attempt to impersonate a legitimate device or exploit weaknesses in the communication infrastructure connecting smart meters to utility systems [8].
These attacks are particularly important for NTL detection because the data used by the utility’s billing and analytical systems may no longer correspond to the physical measurement at the customer premises. Therefore, a high-performing consumption-based detector may still be vulnerable if the input data themselves have been manipulated before reaching the detection system. The reviewed literature indicates that robust authentication and encryption mechanisms can substantially mitigate communication-based attacks by increasing the difficulty and cost of exploitation, although they cannot completely eliminate all attack vectors. In addition, AMI-based anti-tamper mechanisms can generate events that are communicated to utility systems through the available bidirectional communication infrastructure, enabling potentially faster detection of physical tampering [37,46].

2.4.4. Smart-Meter Compromise

Smart-meter compromise represents a further cyber-physical threat in which an attacker gains physical or electronic access to the metering device and attempts to modify its operation. The reviewed literature reports several forms of meter tampering, including modifications to signal-conditioning circuits, electronic memories, and firmware [8,37,42]. Physical access can also be used to neutralize or bypass anti-tamper mechanisms before additional modifications are introduced.
To counter such attacks, smart meters may incorporate mechanisms for detecting the opening of the meter enclosure or its removal from the base. Such events can be stored in the meter’s memory and subsequently reported to utility systems in real time or on demand [46]. Physical design measures, including mechanisms that make unauthorized opening of the meter evident during inspection, can provide an additional layer of protection. However, the reviewed literature also notes that some anti-tamper mechanisms may themselves be neutralized when their location and operation are known to the attacker [39].
Overall, the emerging threat landscape demonstrates that electricity theft may occur at multiple layers of the smart-grid infrastructure: the physical meter, the measurement circuitry, the communication channel, or the information-processing system. Nevertheless, the available literature does not support the conclusion that cyberattacks have replaced conventional physical theft. Instead, cyber-physical threats should be regarded as an emerging and increasingly relevant dimension of NTL detection. This distinction is important when interpreting the findings of the tertiary review, since the evidence base continues to contain substantial emphasis on physical attacks, unauthorized connections, and consumption-based detection [8,40,43].

2.4.5. Attacks Against NTL Detection Models

The cyber threat landscape also extends beyond the metering and communication infrastructure to the data-driven NTL detection models themselves. The reviewed literature identifies adversarial evasion and poisoning attacks as potential threats to machine-learning-based NTL detection [36,62,63]. In an evasion scenario, an adversary may deliberately modify a fraudulent consumption profile so that it resembles legitimate consumption and is consequently classified as normal. In a poisoning scenario, the attacker interferes with the training process by corrupting labels or modifying training data, potentially degrading the detector even when the inference pipeline remains protected [62]. These threats indicate that detection accuracy on static benchmark datasets alone is insufficient to establish the robustness of an NTL detection system. Adversarial robustness should therefore be considered alongside conventional predictive-performance measures when evaluating data-driven detection systems.

2.5. Emerging AI Directions

The evolution of NTL detection has closely followed advances in artificial intelligence and the increasing availability of fine-grained smart-meter data. Early data-driven approaches primarily relied on statistical analysis and conventional machine-learning classifiers, while subsequent studies increasingly adopted deep learning models capable of learning complex representations directly from consumption time series. Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), Long Short-Term Memory (LSTM) networks, and hybrid architectures have consequently become important components of data-oriented NTL detection systems [8,33,34,35,36,37].
More recently, the research direction has expanded beyond conventional discriminative models toward generative models, synthetic data generation, self-supervised learning, and representation learning. These developments are particularly relevant because NTL datasets are often highly imbalanced, contain missing observations, and provide only a limited number of confirmed theft cases. The following subsections summarize three emerging directions that are increasingly relevant to the future development of NTL detection: generative AI, synthetic data generation, and representation learning.

2.5.1. Generative AI

Generative Artificial Intelligence (GenAI) represents an emerging direction in NTL detection and is beginning to complement conventional discriminative machine-learning and deep-learning approaches. In the reviewed literature, generative models are primarily used to address data-related limitations rather than directly replace the final NTL classifier. In particular, they can be used to generate additional fraudulent consumption profiles, reconstruct incomplete consumption sequences, learn latent representations, and support the development of more robust detection models.
Variational Autoencoders (VAEs) and Generative Adversarial Networks (GANs) are among the generative architectures reported in the reviewed literature. VAEs learn a latent representation through an encoder and decoder architecture, while GANs learn to generate samples that resemble the distribution of the training data. Compared with simple random oversampling, these approaches can potentially generate more diverse samples while preserving characteristics of fraudulent consumption patterns [33]. VAE- and GAN-based approaches have also been used for feature extraction, followed by clustering or threshold-based anomaly detection [37].
More recent developments have introduced time-series-oriented generative models. For example, TimeGAN-based approaches have been investigated for reconstructing incomplete consumption sequences and generating additional theft samples before training an NTL detector [58]. Diffusion-based approaches represent another emerging direction. Recent work has explored diffusion models combined with self-supervised representation learning and imbalance-aware fine-tuning to improve the treatment of class imbalance and missing data in AMI streams [32]. Outside the specific NTL detection setting, conditional diffusion models have also been applied to the generation of realistic long-horizon energy profiles, demonstrating potential applications in privacy-aware data sharing and synthetic stress testing [64].
Generative AI can therefore contribute to NTL detection at several stages of the analytical pipeline, including data reconstruction, augmentation, representation learning, and model robustness. However, its role should not be overstated. The available evidence remains more limited than that for conventional machine-learning and deep-learning approaches, and generative models introduce additional computational and validation requirements. Their effectiveness depends strongly on whether the generated samples adequately represent real fraudulent behaviors.
Large language models (LLMs) are also being explored in the smart grid domain, though primarily as tools for operator decision support, compliance reporting, and data orchestration rather than as direct substitutes for numerical time-series detectors [65,66,67]. Recent reviews suggest that LLM-based components may be useful in workflows that involve retrieval, interpretation, and human-in-the-loop reasoning, while their use in safety or regulation critical tasks remains limited by reliability concerns, hallucination risk, and data governance constraints [65,66].

2.5.2. Synthetic Data Generation

A major challenge in data-driven NTL detection is the severe imbalance between legitimate and fraudulent consumption records. In practical settings, confirmed theft cases represent only a small proportion of the available consumption data, making it difficult to train supervised learning models effectively. Consequently, synthetic data generation has been widely investigated as a strategy for increasing the representation of the minority class.
Traditional approaches include Random Oversampling (ROS), Random Undersampling (RUS), Synthetic Minority Oversampling Technique (SMOTE), Cluster-Based Oversampling (CBOS), weighting strategies, and cluster-based methods. Comparative studies have reported that techniques such as SMOTE and CBOS can improve AUC and F1-score compared with models trained on the original imbalanced data, while ROS has also shown benefits for some CNN-based classifiers [4,16]. These findings demonstrate that appropriate treatment of class imbalance can be as important as the choice of the underlying classifier.
Generative models provide a more flexible alternative because they can learn characteristics of the observed consumption distribution and generate additional samples. VAEs and GANs have therefore been applied to generate synthetic electricity-theft profiles, with the objective of increasing sample diversity and improving the training of downstream classifiers [33]. More recent approaches have moved toward time-series-aware generation. TimeGAN has been used to reconstruct incomplete consumption sequences and generate theft samples with greater intra-class diversity [58]. Diffusion-based methods further extend this direction by combining synthetic sequence generation with self-supervised learning and imbalance-aware training [32].
Despite these advantages, synthetic data generation introduces several limitations. Generated samples are constrained by the attack patterns and consumption behaviors represented in the training data. Therefore, a model trained with synthetic samples may fail when confronted with previously unseen fraud strategies or consumption patterns from a different geographical or operational context. Furthermore, generative models can increase computational requirements, while the confidential nature of utility datasets can make independent reproduction and validation difficult [58]. Consequently, synthetic data should complement rather than replace real-world validation.
An important implication for future NTL research is that studies should evaluate not only whether synthetic generation improves conventional classification metrics, but also whether generated samples improve cross-region generalization, robustness to unseen attacks, and performance under realistic AMI data distributions. Comparisons with imbalance-aware loss functions and transfer-learning approaches are also necessary to determine when generative augmentation provides a meaningful advantage [68,69,70].

2.5.3. Representation Learning

Representation learning has become an important component of modern NTL detection because electricity consumption data are inherently high-dimensional, temporal, and heterogeneous. Traditional machine learning generally requires manually engineered features, whereas deep learning models can learn task-relevant representations directly from raw or minimally processed consumption sequences.
The evolution of NTL detection has progressed from conventional artificial neural networks toward CNNs, RNNs, LSTMs, GRUs, and hybrid architectures [33,50,71]. CNNs can automatically learn local patterns and non-linear relationships in consumption profiles, reducing the need for manual feature engineering. However, their ability to capture long-term temporal dependencies may be limited. RNN-based models, particularly LSTM and GRU architectures, are more suitable for sequential consumption data because they explicitly model temporal dependencies. Consequently, combinations of CNNs with recurrent architectures have been investigated to jointly capture local consumption characteristics and temporal dynamics [8,33,39].
Recent work has further extended representation learning through attention mechanisms, graph-based architectures, and self-supervised learning. Graph-based models can exploit relationships between customers and periodic consumption structures when such relational information is available [72]. Similarly, ConvLSTM architectures combine spatial or local pattern extraction with temporal modeling, although their performance can be affected by changes in tariffs, seasonal behavior, and consumption patterns [24].
Generative models can also contribute to representation learning. VAE and GAN architectures have been used to obtain latent features from consumption data, with subsequent clustering methods such as k-means being applied to identify abnormal load profiles [37]. More recently, diffusion-based approaches have combined synthetic generation with self-supervised representation learning, indicating a potential transition from purely supervised classification toward representation-driven NTL detection [32].
Despite these developments, learned representations remain dependent on the quality and diversity of the training data. Distribution shifts caused by seasonal changes, tariff modifications, regional differences, or evolving fraud behaviors can reduce the generalization capability of deep models [24,73,74]. Furthermore, increasing model complexity can introduce additional requirements in terms of computational resources, interpretability, and inference latency. Therefore, future NTL detection systems should evaluate learned representations not only according to predictive accuracy but also according to robustness, interpretability, computational efficiency, and transferability across datasets and regions.
Synthesis of Emerging AI Directions
Overall, the evolution of AI for NTL detection indicates a gradual shift from conventional supervised classification toward data-centric and representation-driven approaches. Generative AI contributes primarily through synthetic data generation, missing-data reconstruction, and latent representation learning, while deep representation-learning models increasingly reduce the dependence on manually engineered features. However, the reviewed evidence also indicates that these emerging approaches do not eliminate the fundamental challenges of NTL detection. Data imbalance, limited availability of real theft cases, distribution shifts, privacy constraints, computational requirements, and the absence of standardized evaluation remain significant barriers. Therefore, the future development of AI-based NTL detection is likely to depend not on a single model architecture, but on the integration of robust representation learning, reliable synthetic data generation, privacy-aware learning, and evaluation under realistic operational conditions.

2.6. Positioning of the Present Tertiary Review

The preceding subsections demonstrate that the secondary literature on NTL and electricity-theft detection has evolved from general discussions of theft mechanisms and conventional detection strategies toward data-driven methods, AMI, cyber-physical threats, and emerging artificial intelligence techniques. However, these developments have been examined from different perspectives and with different methodological assumptions.
The present study is positioned at a higher level of evidence synthesis. Rather than treating individual NTL-detection algorithms or datasets as the primary units of analysis, the study considers previously published surveys, systematic reviews, and comprehensive review studies as its units of analysis. Thus, the primary studies underlying those reviews constitute the evidence base of the existing secondary literature, whereas the eligible secondary studies constitute the evidence base of the present tertiary review.
The distinction is important because the objective of this study is not to reproduce the method-level comparisons already performed in individual surveys. Instead, the study systematically examines how the existing secondary literature characterizes NTL detection and whether consistent conclusions can be identified across reviews. To achieve this, the heterogeneous findings reported by the secondary studies are mapped onto a common analytical framework covering five dimensions: attack types, detection approaches, data and feature requirements including AMI, evaluation practices, and challenges and research gaps.
This positioning also allows the study to examine areas where the secondary literature converges as well as areas where conclusions differ. For example, individual reviews may report strong performance for data-driven approaches while simultaneously identifying limitations associated with class imbalance, limited real-world data, generalizability, privacy, or computational requirements. Similarly, AMI can be viewed both as an enabler of fine-grained data-driven detection and as an infrastructure that introduces additional cybersecurity and privacy risks. Such differences cannot be adequately characterized by considering individual reviews in isolation.
Accordingly, the present tertiary review complements rather than duplicates existing secondary studies. The comparison of representative secondary studies with the present review, including the differences in evidence level, analytical scope, cross-review synthesis, and quantitative coverage, is presented in Section 2.3. The methodological procedure used to identify, assess, extract, and synthesize the eligible secondary studies is described in the following section.

3. Research Questions and Methodology

This section describes the methodology adopted for the tertiary review, including the formulation of the research questions, search strategy, eligibility criteria, study-selection process, quality assessment, evidence extraction, screening reliability, and quantitative cross-survey analysis. The review follows the PRISMA methodology [75] to provide a transparent and reproducible procedure for identifying, screening, assessing, and synthesizing the relevant secondary studies.

3.1. Research Questions

The research questions were initially introduced in Section 1.4 to define the scope and objectives of the tertiary review. In this section, they are operationalized for the methodological analysis. The questions were derived from a preliminary examination of the existing secondary literature and were designed to establish a common analytical framework for comparing heterogeneous surveys and systematic reviews. In particular, the RQs cover five complementary dimensions of NTL detection: attack characteristics, detection approaches, data and AMI requirements, evaluation practices, and challenges and research gaps. Each question defines a specific dimension that was subsequently coded in the tertiary evidence matrix and used in the cross-survey synthesis.

3.1.1. RQ1: What Types of NTL/Energy-Theft Attacks Are Addressed?

RQ1 examines how the secondary literature characterizes the different forms of NTL and electricity theft. The analysis considers physical meter manipulation, meter bypassing and unauthorized connections, abnormal or fraudulent consumption patterns, billing-related irregularities, and emerging cyber or data-oriented attacks. The objective is not to establish a new attack taxonomy from primary studies, but to determine how consistently these attack categories are represented across the included secondary studies and to identify underrepresented or emerging attack dimensions.

3.1.2. RQ2: What Detection Approaches Are Reported?

RQ2 investigates the methodological approaches reported by the secondary literature for detecting NTL. The analysis considers conventional and statistical techniques, machine-learning and deep-learning approaches, anomaly detection and clustering, network-oriented and state-estimation approaches, hybrid and ensemble methods, and emerging AI and generative approaches. Particular attention is given to the evolution of detection paradigms and to differences in how individual reviews classify and characterize these methods.

3.1.3. RQ3: What Data, Features, and AMI Requirements Are Considered?

RQ3 examines the data requirements underlying NTL detection as characterized by the secondary literature. This includes electricity consumption profiles, smart-meter and AMI data, customer and contextual attributes, engineered features, and learned representations. The analysis also considers the influence of AMI and smart-grid digitalization on data granularity and availability, together with issues such as limited labeled theft data, class imbalance, synthetic data generation, and data quality.

3.1.4. RQ4: How Are Detection Approaches Evaluated?

RQ4 examines how the reviewed secondary studies assess the effectiveness of NTL detection approaches. The analysis considers the evaluation metrics reported, comparative evaluation practices, validation strategies, datasets and experimental settings where explicitly reported, and practical criteria such as computational cost, response time, and deployment considerations. Particular attention is given to the heterogeneity of evaluation practices and the extent to which this heterogeneity limits direct comparison of reported performance.

3.1.5. RQ5: What Challenges, Limitations, and Research Gaps Are Reported?

RQ5 synthesizes the challenges and limitations identified across the secondary literature. These include data availability and quality, class imbalance, privacy and cybersecurity, scalability, generalizability, computational complexity, interpretability, deployment constraints, and regulatory or policy considerations. The objective is to determine which challenges recur across independent reviews and which research gaps remain insufficiently addressed in the secondary evidence.

3.2. Search Strategy

A systematic search was conducted following the PRISMA methodology (see Supplementary Materials for the completed PRISMA 2020 checklist). The search strategy was designed to identify secondary studies that review, survey, systematically map, or comprehensively synthesize research on NTL and electricity-theft detection. The search procedure comprised database selection, database-specific search strings, and a predefined search period.

3.2.1. Databases

The literature search was conducted in Scopus, IEEE Xplore, and Web of Science. These databases were selected to provide broad coverage of the peer-reviewed literature in electrical engineering, smart grids, artificial intelligence, cybersecurity, and energy systems.

3.2.2. Search Strings

The keyword search string was defined according to several concepts behind the different definitions of NTL in the literature. A manual search of the authors in the database revealed that the concept of “non-technical losses” can also be found as “theft” or “fraud” followed by different concepts of electrical energy such as “energy”, “electricity” or “power”, and this can be found along with the words “loss” or “losses”. Considering such key terms, the following search string was formulated:
( ( energy OR electricity OR power ) AND ( theft OR fraud ) ) OR ( non-technical AND loss * )
We used the wildcard symbol “*” to capture the terms “loss” and “losses”. In order to find relevant results, we applied the search string to the article’s title. Later on, since the object of our work was specifically on NTL secondary studies, we applied the following search string to the manuscripts’ abstracts:
( review OR survey OR literature OR comparative )
The search strings were adapted to fit the individual needs and syntax of the individual databases.

3.2.3. Search Date and Temporal Scope

The search was conducted in July–August 2026. The publication eligibility period covered studies published from January 2012 through March 2026. Thus, March 2026 represents the upper bound of the publication period, whereas July–August 2026 represents the date on which the database search was executed. The two dates therefore refer to different aspects of the review protocol. After the records screening process and selecting reports, we examined a total of 45 works. However, these works relate to specific issues (e.g., zero-day attack [31]) or specific solutions [47], to narrow contributions [45,76] and broader surveys/review papers outlining definitions, goals, and roadmaps for the NTL problem.

3.3. Eligibility Criteria

To ensure a transparent and reproducible selection of secondary evidence, explicit inclusion and exclusion criteria were defined before the final screening of the candidate studies. The criteria were designed to ensure that the selected studies were sufficiently broad and methodologically informative to support the tertiary synthesis.

3.3.1. Inclusion Criteria

A study was included when it satisfied all of the following criteria:
  • It was a secondary study, such as a survey, systematic review, comprehensive review, literature review, or mapping study.
  • It addressed NTL, electricity theft, energy-theft detection, or a directly relevant AMI- or data-driven detection problem.
  • It provided a sufficiently broad synthesis covering multiple methods, studies, datasets, attack types, or detection perspectives, rather than focusing exclusively on a single narrowly defined solution.
  • It contained information relevant to at least one of RQ1–RQ5.
  • It provided sufficient methodological or descriptive information to support tertiary-level data extraction, comparison, and synthesis.
  • It was available as a full-text scholarly publication.

3.3.2. Exclusion Criteria

A study was excluded when it satisfied one or more of the following conditions:
  • It was a primary research article rather than a secondary study.
  • It focused exclusively on a single algorithm, model, technique, dataset, or narrowly defined solution without providing a broader synthesis.
  • It addressed only a narrow application that was not relevant to the broader NTL or electricity-theft detection landscape.
  • It focused solely on non-detection aspects, such as general smart-grid operation or management, without a substantive connection to NTL or electricity-theft detection.
  • It was a preprint, editorial, poster, thesis, dissertation, short paper, or other publication type outside the defined publication requirements.
  • The full text was unavailable.
  • It did not provide sufficient methodological or descriptive information for systematic tertiary-level extraction and comparison.
  • It substantially duplicated another included review without providing sufficiently distinct evidence or scope.
Geographical scope was not used as an exclusion criterion. Secondary studies addressing specific countries or geographical regions were considered eligible when they satisfied the above criteria and provided a sufficiently broad synthesis relevant to the research questions.

3.4. Study Selection

The study-selection procedure followed the PRISMA framework and consisted of identification, duplicate removal, title/abstract screening, and full-text eligibility assessment. The complete selection process is summarized in the PRISMA flow diagram shown in Figure 2.

3.4.1. Screening

During the initial screening stage, titles, abstracts, and keywords were examined against the predefined eligibility criteria. Studies that were clearly primary research, narrowly focused, outside the NTL scope, or otherwise ineligible were removed.

3.4.2. Full-Text Assessment

The remaining studies were assessed in full text for relevance, methodological adequacy, and compliance with the predefined eligibility criteria. Following this assessment, 28 secondary studies were retained as the final evidence base for the tertiary synthesis.

3.5. Core and Supplementary Evidence

A clear distinction was maintained between the core evidence used for the tertiary synthesis and supplementary evidence used to provide context for recent developments.
Core evidence. The core evidence consists of the 28 eligible secondary studies identified through the PRISMA-based selection process. These studies constitute the units of analysis of the tertiary review and are used for the quality assessment, tertiary evidence matrix, RQ1–RQ5 synthesis, and quantitative cross-survey analysis.
Supplementary evidence. Additional primary studies cited in the manuscript are not part of the 28-study tertiary evidence base. They are used only as contextual or supplementary evidence when discussing recent developments, emerging technologies, specific datasets, cyber-physical threats, Generative AI, or other developments that provide useful context for interpreting the secondary literature.
Consequently, findings reported as results of the tertiary synthesis are derived from the 28 included secondary studies. Statements based on individual primary studies are explicitly treated as supplementary or contextual evidence.

3.6. Quality Assessment

To assess the methodological reliability of the secondary evidence included in the tertiary synthesis, a formal quality assessment was conducted for all 28 eligible secondary studies. Since the included literature comprises systematic reviews, surveys, comprehensive reviews, and narrative/state-of-the-art reviews with heterogeneous methodological designs, a common quality-assessment checklist was adopted rather than restricting the assessment to a single systematic-review-specific instrument.
The checklist evaluates six methodological dimensions that are relevant to the reliability and reproducibility of secondary studies.

3.6.1. Quality Assessment Criteria

The six quality-assessment questions (QAs) were applied to each included secondary study: QA1: Are the review objective and scope clearly stated? QA2: Is the literature-search strategy adequately described and sufficiently transparent? QA3: Are the inclusion and exclusion criteria explicitly reported? QA4: Is the study-selection or screening procedure adequately described? QA5: Is the evidence-synthesis procedure systematic and sufficiently reproducible? QA6: Does the study explicitly report limitations, threats to validity, or research gaps?
These criteria were selected to capture the principal methodological features required for assessing the transparency, rigor, and reproducibility of heterogeneous secondary studies. The assessment was performed at the level of the secondary study because these studies constitute the units of analysis in the present tertiary review.

3.6.2. Quality Scoring Procedure

Each quality-assessment criterion was evaluated using a binary scoring scheme. A score of 1 was assigned when the criterion was adequately addressed by the secondary study, whereas a score of 0 was assigned when the criterion was not adequately addressed or the required information was not sufficiently reported. Thus, each secondary study received a total quality score ranging from 0 to 6, calculated as Q i = j = 1 6 Q A i j , where Q i denotes the total quality score of secondary study i and Q A i j { 0 , 1 } represents the score assigned to quality criterion j for study i.
For interpretation, higher scores indicate greater methodological transparency and reporting completeness with respect to the criteria considered. The quality score was not used as an automatic exclusion threshold. Instead, it was used to characterize the methodological strength of the secondary evidence and to inform the interpretation of the tertiary synthesis, particularly where findings were heterogeneous or supported by a limited number of secondary studies.

3.6.3. Quality Assessment Results

The quality assessment was applied to all 28 secondary studies included in the tertiary synthesis. The assessment provides an overview of the methodological transparency of the included reviews and indicates the extent to which the secondary evidence satisfies the predefined quality criteria.
The quality assessment was not used to remove studies after the eligibility stage. Instead, it provides an additional dimension for interpreting the evidence. In particular, findings repeatedly reported by secondary studies with strong methodological transparency can be distinguished from observations derived primarily from studies with limited reporting of search, selection, or synthesis procedures.

3.7. Data Extraction

A structured data-extraction framework was developed to ensure consistent coding of the 28 included secondary studies. For each study, the following information was extracted: bibliographic information, publication year, citation count, principal contribution, and evidence coverage corresponding to RQ1–RQ5. The RQ-specific fields record the attack types, detection approaches, data and feature requirements, AMI/infrastructure considerations, evaluation practices, and reported challenges or research gaps.
The extracted information was organized into the tertiary evidence matrix presented in Table 2. The purpose of the matrix is not merely to summarize individual reviews, but to establish a common representation through which similarities, differences, underrepresented dimensions, and recurring research gaps can be identified across the secondary literature.

3.8. Quantitative Cross-Survey Analysis

To complement the qualitative tertiary synthesis, a descriptive quantitative cross-survey analysis was conducted across the 28 included secondary studies. Because the included studies differ substantially in scope, datasets, primary-study populations, experimental protocols, and evaluation metrics, statistical pooling of predictive performance measures such as accuracy, precision, recall, or F1-score was not considered methodologically appropriate. Instead, the analysis quantifies the extent to which major research dimensions are represented across the secondary literature.
For each research dimension d, coverage was calculated as C o v e r a g e ( d ) = N d N × 100 , where N d denotes the number of included secondary studies that explicitly address dimension d, and N = 28 represents the total number of secondary studies included in the tertiary synthesis.
The analysis was performed for the dimensions associated with RQ1–RQ5. The resulting coverage values are presented and interpreted in Section 6, particularly in Section 4.6. The coverage measure is descriptive and indicates the representation of a research dimension within the secondary literature; it does not measure the methodological quality, predictive effectiveness, or superiority of the corresponding approach.

4. Tertiary Synthesis and Results

Based on the five research questions, the synthesis considers attack types, detection approaches, data and AMI requirements, evaluation practices, and reported challenges and research gaps. In addition to the qualitative synthesis, the coverage of individual research dimensions across the 28 secondary studies is quantified to identify areas of strong, moderate, and limited representation in the existing secondary literature.
The synthesis reveals a clear transition from conventional and inspection-oriented approaches toward data-driven and AI-based detection. However, this transition has not resulted in a single dominant detection paradigm. The effectiveness of individual approaches remains strongly dependent on the availability and quality of consumption data, the type of theft considered, the degree of grid digitalization, and the evaluation setting. This observation is consistent with the heterogeneous conclusions reported across the reviewed literature.

4.1. RQ1: Attack Types

RQ1: What types of NTL/electricity-theft attacks are addressed in the secondary literature? The tertiary synthesis indicates that electricity theft is not a homogeneous phenomenon. The reviewed secondary studies identify several attack categories, ranging from physical manipulation of metering infrastructure to unauthorized connections and increasingly digital manipulation of meter or communication data.

4.1.1. Meter Tampering and Manipulation

Meter tampering is the main way customers commit fraud and involves manipulating metering equipment to reduce or nullify the energy recorded by the device relative to actual consumption [8,33,35,37,39]. This type of attack is perpetrated against both traditional meters and smart meters, although in different ways [8]. The customers themselves can tamper the meter or the modification can be done by a specialized criminal who knows the device and how it works. This type of attack is perpetrated by removing the meter from its base and making some modifications inside (Figure 3, case 1). The rheostat shown in Figure 3, case 1 represents the action of this partial or total reduction in measurement through internal manipulation of the meter’s hardware or software. In some cases, the meter is not removed from the base at all, and the modification is done with sophisticated instruments such as probes through small holes, in the manner of laparoscopic surgery [39]. This method is generally preferred by fraudsters in several cases: when a physical seal is interposed between the base and the meter, and a broken seal could be noticed during a company’s inspection, or i.e., when the meter is provided with an anti-tamper device that does not allow the removal of the meter, as removal could disable power from the instrument’s main switch. The laparoscopic method could be preliminary to Figure 3, case 1 attacks; in fact, if anti-tamper systems are present, a laparoscopic attack could allow any pre-purposed sensors to be neutralized. Once these sensors are neutralized, the meter can be removed from the base in order to launch further attacks. Once access to the interior of the meter is gained, tampering can involve modifications to the signal conditioning circuits (current-to-voltage converter), to the electronic memories where measurement data are stored, or to the firmware of the apparatus [8,37,42]. A particular type of tampering does not involve any physical modifications to the meter. This occurs when a magnetic field is applied near the meter by a physical magnet in order to create a disturbance in the amperometric measuring circuits [8,39,46]. Since in cases of high power utilization the current-voltage transducer of the measuring circuit is implemented through a current transformer, the presence of the magnetic field can partially saturate the ferromagnetic material, resulting in lower current transduction and, consequently, lower energy measurement by up to 50–75% [8]. In the past years, a piece of film was inserted in traditional meters to restrain or reduce the movement of the rotating disk without the need to remove the meter from its housing [37].
The coverage analysis shows that 11 of the 28 secondary studies (39.3%) explicitly address meter tampering or manipulation. This makes it the most frequently represented attack category in the tertiary evidence base.

4.1.2. Unauthorized Connections and Meter Bypass

Unauthorized connections can involve the illegal tapping of electricity directly from the distribution line or transformer without passing through the metering infrastructure [33,35,37,39,46]. There are different types of unauthorized connections. When the perpetrators are regular customers provided with assigned meters, they can create illegal connections to bypass the meter entirely (Figure 3, case 3) or partially (Figure 3, case 2). In the case 2, the meter continues to operate and record consumption but at a reduced amount compared to the customer’s actual consumption. Only a portion of household appliances is legitimately metered (load 1), while the other part is not metered at all (load 2). In case 3, all the household appliances are bypassed through the illegal connection, resulting in zero current being measured by the meter despite energy consumption. Where there is no existing meter and utility contract, users may illegally tap directly into the distribution network to steal energy (Figure 3, case 4). In this case, the fraudster is not present in the electricity company’s customers database. All the cases, except case 5 in Figure 3, can be carried out with both smart and traditional meters.
In 7 of the 28 secondary studies (25.0%), unauthorized connections or meter bypass are explicitly addressed.

4.1.3. Cyber and Data Attacks

With the rise in smart grids and AMI systems, cyberattacks have become a new method of energy theft [8,33,34,37,40,46]. Cyberattacks in the context of smart grids and electricity theft involve manipulating data and systems through digital means (Figure 3, case 5). In an Advanced Metering Infrastructure (AMI), electricity consumption data can be delivered from residential smart meters to utility control centers through a multi-stage communication architecture. A typical configuration may employ power line communication (PLC) for the initial local transmission, followed by a wide-area communication technology such as 5G or optical fiber. Although this architecture facilitates efficient and large-scale data collection, the communication links may expose the system to sophisticated cyber threats. One important threat is false-data injection, in which an attacker deliberately alters transmitted or recorded measurements to evade conventional detection mechanisms and cause incorrect system observations [8,37]. Attackers may also exploit legitimate communication privileges by impersonating authorized devices—for example, posing as a neighborhood-area-network data collector—to gain access to smart meters. Furthermore, physical or interface-level attacks have been demonstrated in which an attacker connects a smart meter to a general-purpose computer through an optical conversion interface, such as an infrared device, and subsequently manipulates the recorded electricity-consumption information [8]. In man-in-the-middle attacks [46], the attacker, for example, following an energy reading request from the system, might send an altered consumption value compared to the actual one, thus deceiving the system. We have noted that it is not mentioned in the literature that although in such cited cyberattacks cases the energy billed to the customer is lower than the real, nevertheless the energy recorded by the meter remains correct. This means that upon the first inspection of the utility operators the actual value recorded on the smart meter is detected and subsequently billed. This differs from other attacks, such as that of altering metering circuits, where the value recorded on the meter is equal to that of the altered value. In the first case, the fraud perpetrated by a cyberattack can be recovered, in the second case it is not.
Machine learning models used for NTL detection are also vulnerable to attacks defined as “adversarial attacks” namely evasion attacks and poisoning attacks [36]. Below we will analyze a typical evasion attack. In a typical evasion attack, a machine learning model trained to detect NTL is the target. This model might use features like consumption patterns and time of day usage. A Generative Adversarial Network (GAN) is trained on historical data of legitimate consumption patterns. The generator learns to create synthetic consumption data that resembles normal behavior and is then used to modify the consumption data of a consumer engaging in theft, mimicking the consumption patterns of a similar legitimate consumer. The modified consumption data is then fed into the utility’s system. The goal is for the NTL detection system to classify this fraudulent activity as normal, allowing the theft to go undetected. Although not easy to implement in the real world, GAN-based evasion attacks pose a significant challenge to NTL detection in the utility sector.
The adversarial threat is not limited to evasion through manipulated consumption profiles. A second relevant attack vector is data poisoning, in which the attacker interferes with the training process itself, for example by corrupting labels or altering reported readings during model retraining, potentially degrading detection performance even when the inference pipeline is otherwise well-protected [62]. It has also been shown that evasion attacks can be deliberately crafted to reduce consumption below detection thresholds while remaining within the expected behavior range, which suggests that performance results obtained on static benchmarks may not reflect actual robustness in deployed systems [63]. These considerations point to the need for adversarial robustness to be treated as a standard evaluation criterion rather than an optional one, alongside design choices such as adversarial training, ensemble detectors, or community-level modeling where applicable [29]. In order to interfere the NTL detection system, adversaries may constantly switch their behaviors constantly between committing electricity theft and honestly consuming electricity (intermittent NTL). Specifically, they may launch cyber/physical attacks to tamper with meter readings for a while and then stop these attacks for another while. These behavior patterns can be repeated many times. Most existing methods are not designed to address this type of attack [8]. In contrast, ref. [46] presents an evaluation of intermittent fraud attacks. A fraudulent attack could be initiated during specific hours of the day and deactivated, for instance, when an inspection is in place, making the inspection process ineffective.
For example, the attack may be activated and deactivated using a radio-electronic device operated by remote control. As described in [39], an attacker can install a radio-electronic device inside the meter, which is then controlled externally to initiate or terminate the fraudulent activity. Detecting intermittent fraud using automated systems is therefore more challenging. Another instance of a tampered meter employing an electronic board as a theft switch is documented in [77]. Future research could focus on developing detection methods specifically targeting this type of attack.
A real case of NTL for a domestic consumer is shown in Figure 4. It can be seen that the sudden drop in consumption suggests that the attack started on May 20.

4.1.4. Synthesis of RQ1

The main finding from RQ1 is therefore a mismatch between the diversity of attack mechanisms and the capabilities of dominant detection approaches. Consumption-based approaches are well suited to detecting changes in observed load profiles, but they are inherently constrained when theft occurs downstream of the meter or when the attacker manipulates the measurement or communication process itself. This finding is particularly important because the reviewed literature suggests that methods considered effective under consumption-based assumptions may not detect meter-less or intermittent forms of theft.
The limited coverage of cyber and data attacks in the secondary literature is notable given the increasing digitalization of distribution networks. Only 3 of the 28 secondary studies (10.7%) explicitly address cyber or data-oriented attacks. Thus, although measurement manipulation, false-data injection, and AMI communication attacks represent important emerging threats, they remain comparatively underrepresented in the existing secondary evidence.
This finding should not be interpreted as evidence that cyberattacks have replaced conventional electricity theft. Rather, it indicates a research gap in integrating cyber-physical threats with established consumption-based NTL detection. This gap motivates the future-oriented recommendations discussed in Section 6.5.

4.2. RQ2: Detection Approaches

RQ2: What detection approaches are reported across the secondary literature?
The tertiary synthesis demonstrates an evolution from conventional and statistical approaches toward machine learning, deep learning, hybrid models, and, more recently, generative approaches. Nevertheless, the available evidence does not support the conclusion that one particular method is universally superior. The context of solutions varies significantly over time, and several approaches have been proposed, which we illustrated in Figure 5.
The quantitative synthesis shows that 16 of the 28 secondary studies (57.1%) explicitly discuss machine-learning or AI-based approaches. Deep learning is reported in 7 studies (25.0%), while statistical/conventional approaches and hybrid/ensemble approaches are represented in 7 studies (25.0%) and 6 studies (21.4%), respectively. Unsupervised, anomaly-detection, and clustering approaches appear in 5 studies (17.9%), while security/intrusion-detection approaches are covered by only 1 study (3.6%).
This distribution demonstrates the increasing importance of data-driven detection. The secondary literature increasingly emphasizes machine learning and deep learning because AMI provides high-resolution consumption data suitable for automated analysis. The reviewed studies also indicate a progression from traditional classification and statistical techniques toward CNNs, RNNs, and more recent generative approaches.

4.2.1. Convergence

A major point of convergence across the secondary literature is that data-driven approaches have become an important component of modern NTL detection. Secondary studies increasingly emphasize machine learning and deep learning because AMI provides high-resolution consumption data that can support automated analysis of customer load profiles and anomalous consumption patterns. This convergence, however, should not be interpreted as evidence that data-driven approaches are universally superior to conventional methods. Rather, their applicability depends on the availability, quality, and representativeness of the underlying consumption data and on the characteristics of the theft scenarios being considered.

4.2.2. Divergence

The secondary literature differs substantially in its assessment of the effectiveness and superiority of deep-learning approaches. Some reviews report improved detection performance when deep architectures are employed, whereas others highlight that such improvements are frequently demonstrated using curated, synthetic, geographically restricted, or otherwise limited datasets. Consequently, higher accuracy under one experimental configuration does not necessarily imply greater robustness or generalizability in operational environments. This divergence indicates that the reported effectiveness of a detection approach should be interpreted jointly with the dataset characteristics, attack assumptions, evaluation protocol, and degree of real-world validation. Therefore, direct ranking of detection approaches across the secondary studies is not methodologically justified.

4.2.3. Generative AI

Generative AI should be interpreted cautiously within the current evidence base. Only 4 of the 28 secondary studies (14.3%) explicitly cover Generative AI or generative models. Therefore, Generative AI represents an emerging research direction rather than an established dominant detection paradigm. Its principal role in the current literature is to address structural data limitations, including class imbalance, missing observations, and the scarcity of confirmed theft examples. Time-series generative models and diffusion-based approaches have been explored for synthetic data generation, consumption-sequence reconstruction, and related data-augmentation tasks. Thus, the tertiary evidence does not support presenting Generative AI as the dominant current solution for NTL detection. Rather, its significance lies in its potential to address persistent limitations of existing data-driven methods, particularly where real-world fraudulent samples are scarce or incomplete.

4.3. RQ3: Data, Features and AMI

RQ3: What data, features, and AMI requirements are considered in the secondary literature?
The synthesis shows that the development of NTL detection has become closely associated with the increasing availability of smart-meter and Advanced Metering Infrastructure (AMI) data. Smart meters provide substantially more granular consumption information than traditional periodic meter readings and enable two-way communication between customers and utilities. This has created opportunities for real-time monitoring, anomaly detection, and data-driven analysis.The quantitative coverage analysis shows that 11 of the 28 studies (39.3%) explicitly discuss smart-meter or AMI data, while 9 studies (32.1%) focus on consumption or load-profile data.

4.3.1. Smart-Meter and AMI Data

Smart-meter and AMI data constitute an important foundation for data-driven NTL detection. The availability of fine-grained consumption measurements enables utilities to monitor customer behavior at a much higher temporal resolution than conventional periodic meter readings. AMI additionally provides bidirectional communication capabilities, which can support remote monitoring, meter management, and automated detection of abnormal consumption behavior. The tertiary synthesis shows that smart-meter or AMI data are explicitly considered in 11 of the 28 secondary studies (39.3%). However, the coverage does not imply that all studies use the same type, granularity, or quality of AMI information. Differences in sampling frequency, available customer attributes, measurement resolution, and data completeness can substantially influence the resulting detection performance.

4.3.2. Consumption and Load-Profile Data

Consumption and load-profile data are another major source of information for NTL detection. Such data can be represented as temporal sequences or customer load profiles and can subsequently be analyzed using statistical, machine-learning, or deep-learning approaches. The quantitative synthesis indicates that 9 of the 28 secondary studies (32.1%) explicitly consider consumption or load-profile data. These approaches are particularly suitable for detecting deviations from expected consumption behavior. However, their effectiveness depends on the assumption that fraudulent activity produces a measurable change in the observed consumption profile. This assumption becomes problematic for meter bypass, partial theft, or attacks in which the reported measurement itself is manipulated. Consequently, consumption-based detection should be interpreted in relation to the specific attack model considered.

4.3.3. Features and Representation

The reviewed secondary literature reports a variety of features derived from electricity consumption and customer information. Typical features include energy consumption, average, minimum and maximum consumption, standard deviation, power-factor-related measures, temporal differences, geographical information, tariff category, number of phases, payment regularity, and reading frequency. The use of such features reflects two broad analytical strategies. Conventional machine-learning approaches generally depend on hand-engineered statistical, temporal, or customer-level features, whereas deep-learning approaches can learn representations directly from consumption sequences. The tertiary evidence therefore indicates a gradual transition from manually designed features toward automatically learned representations. Only 2 of the 28 secondary studies (7.1%) explicitly identify feature representation as a distinct research dimension. This relatively low coverage also suggests that the secondary literature does not consistently report feature construction and representation practices, which can make methodological comparison difficult.

4.3.4. Real-World, Synthetic, and Constrained Data

A substantial limitation of the existing evidence concerns the availability and characterization of evaluation data. Only 1 of the 28 secondary studies (3.6%) explicitly identifies real-world data as an evaluation component, whereas 5 studies (17.9%) explicitly discuss limited, synthetic, or otherwise constrained datasets. This indicates that the availability of AMI does not automatically imply the availability of suitable public research datasets. The reviewed literature continues to rely heavily on private, synthetic, or geographically restricted datasets, limiting reproducibility and cross-study comparison. The limited availability of real-world theft data is particularly important because fraudulent consumption patterns are difficult to obtain, label, and share. Consequently, performance reported on synthetic or restricted datasets may not necessarily generalize to different utility environments or previously unseen theft behaviors.

4.3.5. Synthesis of RQ3

The tertiary synthesis indicates that AMI is simultaneously an enabler and a vulnerability for NTL detection. It provides fine-grained data that facilitate sophisticated data-driven detection, while its digital communication architecture introduces additional privacy, integrity, and cybersecurity risks. Overall, the evidence reveals a gap between the increasing availability of smart-meter infrastructure and the limited availability of standardized, openly accessible, real-world datasets. This gap constrains reproducibility, benchmarking, and the assessment of cross-region generalizability of NTL detection approaches.

4.4. RQ4: Evaluation Practices

RQ4: How are NTL detection approaches evaluated across the secondary literature?
The tertiary synthesis reveals considerable heterogeneity in the evaluation practices reported across the secondary literature. Twelve of the 28 secondary studies (42.9%) explicitly discuss performance or effectiveness metrics, while 9 studies (32.1%) report comparative evaluation. Only 2 studies (7.1%) explicitly consider cost, response time, or execution time.

4.4.1. Performance Metrics

Frequently reported metrics include accuracy, detection rate, precision, recall, F1-score, ROC-related measures, and other classification indicators. These metrics are commonly used to assess the ability of a detection model to distinguish fraudulent from legitimate consumption behavior. However, the secondary literature does not employ a common evaluation protocol. Different studies may use different datasets, class distributions, sampling frequencies, experimental settings, and performance measures. Consequently, reported performance values cannot be interpreted as directly comparable across all secondary studies.

4.4.2. Comparative Evaluation

Comparative evaluation is explicitly reported by 9 of the 28 secondary studies (32.1%). Such evaluations typically compare multiple detection approaches or model families under a common experimental setting. Although comparative studies provide useful evidence about relative performance, the validity of the comparison depends on the consistency of the underlying datasets, preprocessing procedures, attack scenarios, and evaluation metrics. Comparisons conducted on different datasets or under different assumptions should therefore not be interpreted as evidence of universal superiority.

4.4.3. Class Imbalance and Evaluation Bias

A particularly important concern in NTL detection is class imbalance, because fraudulent customers generally constitute a much smaller population than legitimate customers [58,68,78]. Under such conditions, accuracy alone can provide a misleading indication of detector effectiveness. Metrics such as precision, recall, F1-score, detection rate, and related measures can provide more informative assessments when the fraudulent class is relatively rare. Nevertheless, the tertiary evidence indicates that evaluation practices remain heterogeneous, and the treatment of class imbalance is not consistently reported across the secondary literature.

4.4.4. Operational and Computational Evaluation

Only 2 of the 28 secondary studies (7.1%) explicitly consider cost, response time, or execution time. This represents a substantial gap because practical deployment of NTL detection systems requires more than predictive accuracy. Operational considerations include computational requirements, detection latency, scalability, communication overhead, inspection cost, and the economic benefits associated with correctly prioritizing customers for physical inspection. These factors are particularly important for utility-scale systems serving large numbers of customers.

4.4.5. Real-World Validation

The tertiary evidence also indicates limited validation using real-world operational data. Many reported results originate from private, synthetic, geographically restricted, or otherwise constrained datasets. Differences in dataset characteristics and data resolution make it difficult to establish whether reported performance will remain stable when models are deployed in different utility environments. Real-world validation is therefore necessary to assess robustness against changing consumption behavior, seasonal effects, evolving theft strategies, and differences in customer populations.

4.4.6. Synthesis of RQ4

The tertiary synthesis identifies three major weaknesses in the current evaluation landscape: (i) lack of standardized datasets, (ii) lack of standardized evaluation metrics and protocols, and (iii) limited validation using real-world operational data. Consequently, reported numerical performance should not be interpreted as directly comparable across secondary studies unless their datasets, attack assumptions, class distributions, and evaluation protocols are sufficiently aligned. The evidence therefore supports the need for more standardized and deployment-oriented evaluation frameworks for future NTL detection research.

4.5. RQ5: Challenges and Research Gaps

RQ5: What challenges and research gaps are consistently reported across the secondary literature?
The tertiary synthesis identifies several recurring challenges across the 28 secondary studies. These challenges span data availability, privacy and cybersecurity, class imbalance, scalability, generalizability, deployment, computational requirements, interpretability, and regulatory considerations.

4.5.1. Data Availability and Quality

Data availability and quality are among the most persistent challenges. In 8 of the 28 studies (28.6%), data availability, quality, or complexity was identified as a research challenge. The lack of publicly available real-world NTL datasets limits reproducibility and benchmarking. In addition, differences in temporal resolution, data completeness, customer characteristics, and labeling procedures can substantially influence model performance. The scarcity of confirmed theft cases further complicates supervised learning because fraudulent samples are typically much less frequent than legitimate consumption records.

4.5.2. Privacy and Cybersecurity

Privacy and cybersecurity are reported by 9 of the 28 secondary studies (32.1%). This concern becomes increasingly important as AMI systems collect detailed consumption information and communicate continuously with utility infrastructure. Fine-grained electricity consumption data can reveal information about customer behavior and occupancy patterns. At the same time, the communication and digital architecture of AMI introduces potential attack surfaces involving data interception, manipulation, unauthorized access, and smart-meter compromise. Consequently, future NTL detection systems must consider not only detection accuracy but also data confidentiality, integrity, and security.

4.5.3. Class Imbalance

Class imbalance is explicitly identified by 2 of the 28 secondary studies (7.1%). Although the coverage is relatively low, class imbalance represents an important technical problem for supervised NTL detection because legitimate consumption typically dominates the available data. The imbalance can cause classifiers to favor the majority class and therefore achieve apparently high accuracy while failing to detect a substantial proportion of fraudulent cases. Data augmentation, resampling, cost-sensitive learning, and generative approaches have therefore been investigated to improve minority-class detection. The relatively limited explicit coverage of this issue across the secondary literature suggests that its treatment is not consistently reported as a separate research dimension.

4.5.4. Scalability and Generalizability

A total of 4 of the 28 studies (14.3%) explicitly identify scalability or generalizability as challenges. Models trained using data from a particular geographical region, customer population, tariff structure, or utility environment may not transfer reliably to another setting. Similarly, changes in seasonal behavior, tariffs, consumption habits, or theft strategies can cause distribution shifts that reduce detection performance. Scalability is also important because practical utility systems may need to analyze consumption information from very large customer populations. Consequently, future approaches should evaluate both cross-dataset generalization and computational scalability.

4.5.5. Deployment and Practical Implementation

Deployment and practical implementation constitute the most frequently reported practical challenge, with 10 of the 28 secondary studies (35.7%) addressing this dimension. The gap between experimental performance and operational deployment arises from several factors, including data accessibility, integration with existing utility infrastructure, computational requirements, communication constraints, false-positive costs, and the need for subsequent physical inspection. Therefore, a detection model should ultimately be evaluated not only according to its classification performance but also according to its ability to support efficient and economically meaningful inspection campaigns.

4.5.6. Computational Resources and Complexity

Computational resources and model complexity are reported by 6 of the 28 studies (21.4%). This issue becomes increasingly relevant as NTL detection moves toward deep-learning, ensemble, and generative architectures. More complex models may provide improved representation capability but can require greater computational resources, memory, and inference time. These requirements can become important when detection systems must operate at large scale or in resource-constrained environments. Thus, future research should consider the trade-off between detection performance and computational cost rather than optimizing predictive accuracy alone.

4.5.7. Interpretability

Interpretability is explicitly identified by only 1 of the 28 secondary studies (3.6%). This relatively low coverage represents an important research gap. The practical use of NTL detection systems may involve decisions about which customers should be selected for physical inspection. Utility operators may therefore require understandable evidence explaining why a particular customer or consumption profile has been classified as suspicious. The limited coverage of interpretability suggests that future research should give greater attention to explainable and transparent detection models, particularly when automated predictions influence operational or regulatory decisions.

4.5.8. Regulatory and Policy Issues

Regulatory and policy issues are explicitly reported by 2 of the 28 secondary studies (7.1%).
NTL detection operates within legal and regulatory environments that may determine how automated detection results can be used and how fraud must ultimately be verified. In some contexts, automated systems can support the identification and prioritization of suspicious customers, while physical inspection remains necessary to establish the existence of fraud.
This distinction highlights the need to consider the legal and institutional context when designing automated NTL detection systems.

4.5.9. Synthesis of RQ5

The most important tertiary-level finding from RQ5 is that the literature has progressed considerably in algorithmic sophistication, but this progress has not been matched by equivalent progress in data availability, standardized evaluation, deployment validation, interpretability, cost-benefit analysis, and robustness against evolving attacks.
The coverage results further indicate that deployment and practical implementation (35.7%), privacy and cybersecurity (32.1%), and data availability and quality (28.6%) receive substantially greater attention than interpretability (3.6%) and regulatory or policy issues (7.1%). This imbalance suggests that future research should move beyond the development of increasingly complex detection models toward more reproducible, explainable, secure, and deployment-oriented solutions.

4.6. Quantitative Coverage Across Secondary Studies

The descriptive quantitative analysis provides an additional cross-survey perspective on the evidence synthesized across the 28 included secondary studies. The analysis does not compare the predictive performance of individual detection algorithms; rather, it measures how frequently specific research dimensions are addressed across the secondary literature. The resulting coverage statistics are reported in Table 3.

5. Cross-Survey Analytical Synthesis

The preceding section synthesized the evidence according to the five research questions. However, an important objective of a tertiary study is to move beyond individual research questions and identify patterns that emerge when the included secondary studies are compared against a common analytical framework. Therefore, this section provides a cross-survey synthesis of the 28 included secondary studies with respect to their assumptions, methodological approaches, datasets and data requirements, evaluation practices, reported limitations, and practical relevance.
The synthesis focuses on identifying both areas of convergence and divergence across the secondary literature. In contrast to the study-by-study descriptions presented in the tertiary evidence matrix, the analysis here examines the secondary literature as a collective body of evidence. This allows recurring assumptions and methodological patterns to be distinguished from findings that are dependent on particular datasets, attack models, or experimental settings.

5.1. Areas of Agreement

Despite differences in scope and methodological emphasis, several points of agreement emerge across the secondary literature.
First, the reviewed studies consistently recognize electricity theft and non-technical losses as important problems for power utilities and identify automated detection as a necessary complement to conventional inspection-based approaches. Second, there is broad agreement that the increasing availability of smart-meter and AMI data has created opportunities for data-driven NTL detection.
A further area of convergence concerns the increasing role of machine-learning and artificial-intelligence-based approaches. The secondary literature increasingly considers supervised learning, unsupervised anomaly detection, deep learning, and hybrid approaches for identifying abnormal consumption patterns. However, this convergence concerns the growing relevance of data-driven methods rather than agreement on the superiority of a particular algorithm.
The reviewed studies also show substantial agreement regarding persistent challenges, particularly limited availability of representative real-world datasets, class imbalance, privacy and security concerns, computational requirements, and difficulties in transferring experimental results to operational environments.
Thus, the principal consensus across the secondary literature is not that a single detection technique is optimal, but rather that effective NTL detection requires reliable data, appropriate analytical models, robust evaluation, and consideration of the operational context.

5.2. Areas of Divergence

Although several broad conclusions are shared, the secondary studies differ considerably in their characterization of effective detection strategies. The main divergence concerns the relative importance of statistical methods, conventional machine learning, deep learning, hybrid approaches, anomaly detection, and security-oriented methods.
Some reviews emphasize the predictive capability of machine-learning and deep-learning models, whereas others highlight the advantages of simpler statistical, feature-based, or anomaly-detection approaches, particularly when computational resources or labeled data are limited. Similarly, the reported advantages of deep learning are often dependent on the characteristics of the datasets and experimental settings used for evaluation.
Divergence is also evident in the treatment of cyber-physical threats. Some secondary studies primarily consider physical theft and abnormal consumption, whereas others explicitly consider privacy, cybersecurity, false-data injection, data integrity, or attacks against AMI infrastructure.
The secondary literature therefore does not provide a uniform definition of the NTL detection problem. Differences in attack models, data assumptions, infrastructure characteristics, and evaluation objectives contribute to the observed divergence.

5.3. Assumptions

The assumptions underlying the reviewed secondary studies have a direct influence on their conclusions. A major assumption in consumption-based NTL detection is that fraudulent activity produces a detectable deviation in the observed electricity consumption profile. Under this assumption, historical load profiles can be used to identify abnormal behavior.
However, this assumption is not valid for all theft scenarios. Partial meter bypass, downstream unauthorized connections, or manipulation of meter measurements may not produce a sufficiently distinguishable pattern in the observed data.
A second recurring assumption concerns data availability. Many data-driven approaches assume that sufficient historical consumption records and, in supervised settings, reliable labels of fraudulent customers are available. The scarcity and confidentiality of real-world theft data can make this assumption difficult to satisfy.
A further assumption concerns the stability of consumption behavior. Models trained using data from a particular customer population, geographical region, tariff structure, or time period may implicitly assume that the underlying distribution remains sufficiently stable.
The cross-survey synthesis therefore indicates that the applicability of reported detection methods depends strongly on the validity of their underlying assumptions. These assumptions should be made explicit when comparing results across studies.

5.4. Methods

The reviewed secondary studies collectively cover a broad spectrum of detection methodologies. Conventional statistical techniques, classification algorithms, anomaly detection, clustering, deep learning, hybrid models, ensemble learning, and security-oriented approaches are represented across the secondary literature. Machine-learning and AI-based approaches constitute the most frequently represented methodological direction in the tertiary evidence base. However, the cross-survey comparison does not establish a universal ranking among the methods. Rather, different methodological families address different combinations of data availability, attack characteristics, computational requirements, and detection objectives. Deep-learning methods provide the ability to learn complex representations from consumption sequences, while conventional machine-learning approaches may be advantageous when the available datasets are relatively small or when interpretability and computational efficiency are important. Hybrid and ensemble approaches attempt to combine complementary characteristics of multiple techniques.
An important observation is that methodological sophistication does not necessarily correspond to greater practical reliability. The reported effectiveness of a method depends on the data, attack model, class distribution, preprocessing procedure, and evaluation protocol under which it is assessed.
The cross-survey evidence therefore supports a context-dependent selection of detection methods rather than the assumption that increasingly complex models are inherently better.

5.5. Datasets

5.5.1. Dataset Characteristics

Dataset characteristics represent an important source of heterogeneity across the secondary literature. The reviewed datasets differ substantially in terms of customer population, sampling frequency, observation period, geographical coverage, and availability of appliance-level measurements. Table 4 summarizes representative datasets reported in the reviewed literature.
The datasets exhibit substantial variation, with customer populations ranging from 5 to more than 110,000 and sampling intervals ranging from 6 s to one day. They also differ in geographical context and availability of appliance-level information. Such differences influence the features available for detection and the assumptions under which models are evaluated.
This heterogeneity limits direct comparison of reported detection performance and may affect model generalizability across utilities, regions, customer populations, and tariff structures. In particular, models evaluated using high-resolution or appliance-level data may have access to information that is unavailable in aggregate consumption datasets.
The cross-survey synthesis therefore identifies the lack of standardized, representative, and well-documented benchmark datasets as an important barrier to reproducible evaluation and meaningful comparison of NTL detection approaches. Synthetic or augmented datasets can partially address the scarcity of fraudulent samples, but their effectiveness depends on how realistically they represent actual theft behavior.

5.5.2. Dataset Heterogeneity and Generalizability

The datasets reported across the secondary literature exhibit substantial heterogeneity in terms of the number of customers, sampling frequency, observation period, geographical context, and availability of appliance-level information. As illustrated in Table 4, the number of customers ranges from a very small number of households to more than 100,000 customers, while the sampling frequency varies from seconds to daily observations. Such differences reflect substantially different information settings for NTL detection and make direct comparison of reported results difficult.
Sampling frequency is particularly important because high-resolution datasets can capture short-term variations in consumption behavior and enable the extraction of temporal or appliance-level characteristics, whereas coarse-grained measurements provide considerably less information about transient consumption patterns. Similarly, datasets with appliance-level measurements provide information that is not available in aggregate smart-meter datasets. Consequently, detection models evaluated using different datasets may operate under substantially different feature and information assumptions.
Geographical heterogeneity represents another important source of variation. The datasets reported in the reviewed literature originate from different countries and utility environments, where customer consumption patterns, tariff structures, climatic conditions, infrastructure characteristics, and electricity-theft practices may differ. Therefore, a detection model that achieves high performance on one geographical dataset cannot automatically be assumed to generalize to another region.
The tertiary synthesis consequently identifies dataset heterogeneity as a major barrier to cross-study comparability and generalizability. Reported performance should be interpreted in relation to the characteristics of the underlying dataset rather than as an unconditional measure of the effectiveness of a particular detection approach. Future research would benefit from standardized benchmarking protocols that evaluate detection approaches across multiple datasets, geographical settings, sampling resolutions, and customer populations.

5.5.3. Real-World Versus Synthetic/Restricted Data

A further limitation identified across the secondary literature concerns the availability and accessibility of real-world NTL data. Although smart-meter and AMI systems generate increasingly detailed consumption information, electricity-theft data are difficult to obtain because confirmed fraud cases are relatively scarce and utility consumption records are generally subject to privacy, security, and commercial constraints.
The quantitative synthesis shows that only 1 of the 28 secondary studies (3.6%) explicitly identifies real-world data as an evaluation component, whereas 5 studies (17.9%) explicitly discuss limited, synthetic, or constrained data. This limited representation of explicitly identified real-world evidence highlights a gap between the availability of AMI infrastructure and the availability of accessible datasets suitable for independent research and benchmarking.
Synthetic and constrained datasets can nevertheless play an important role in NTL detection research. They can facilitate controlled experimentation, enable the generation of rare or fraudulent consumption patterns, and support the evaluation of algorithms when confirmed theft records are unavailable. However, synthetic data may not fully reproduce the temporal, behavioral, geographical, and operational characteristics of real electricity theft. Consequently, high performance on synthetic or artificially constructed data should not automatically be interpreted as evidence of equivalent performance in operational environments.
Restricted or geographically specific datasets present a related challenge. Although such datasets may contain valuable real-world information, limited accessibility can prevent independent reproduction and make cross-study benchmarking difficult. Furthermore, models trained on a particular utility or geographical population may learn characteristics that are specific to that environment rather than general patterns of electricity theft.
The cross-survey evidence therefore indicates a need for a balanced evaluation strategy in which synthetic or controlled datasets are used for reproducible experimentation, while real-world datasets are used for external validation. Future studies should explicitly report the provenance, sampling characteristics, geographical context, labeling procedure, and limitations of their datasets to enable meaningful comparison and assessment of generalizability.

5.6. Evaluation Metrics

The secondary literature employs a heterogeneous set of evaluation metrics, including accuracy, precision, recall, F1-score, detection rate, and receiver-operating-characteristic-based measures. Some studies additionally consider computational cost, execution time, response time, or economic aspects of detection. The absence of a common evaluation protocol represents a major obstacle to quantitative comparison. In particular, electricity-theft datasets are often highly imbalanced, meaning that accuracy alone may provide an overly optimistic assessment of a detector.
The cross-survey synthesis therefore indicates that evaluation should consider multiple complementary dimensions, including minority-class detection capability, false-positive behavior, computational requirements, scalability, and operational cost. Furthermore, numerical performance values reported by different secondary studies should not be directly ranked unless the underlying datasets, attack assumptions, class distributions, preprocessing procedures, and evaluation protocols are sufficiently comparable.
This observation is important because the objective of the tertiary study is not to identify a universally “best” algorithm, but to identify the conditions under which different methodological choices have been evaluated and the limitations associated with those evaluations.

5.7. Limitations

Several limitations recur across the secondary literature. The most frequently observed limitations concern data availability and quality, privacy and cybersecurity, class imbalance, generalizability, scalability, computational complexity, and limited real-world validation. The lack of representative real-world datasets limits the ability to assess whether reported performance generalizes to operational utility environments. Similarly, geographically restricted datasets may not capture differences in consumption behavior, tariffs, infrastructure, or theft mechanisms across regions. Another recurring limitation is the gap between algorithmic evaluation and operational deployment. High predictive performance on a benchmark dataset does not necessarily imply that a method can operate reliably at utility scale or reduce the cost of physical inspection.
Interpretability is another comparatively underexplored issue. Since automated detection can be used to prioritize customers for inspection, utility operators may require explanations for why a particular customer has been classified as suspicious. Overall, the cross-survey synthesis suggests that the principal limitations are increasingly shifting from the availability of detection algorithms toward the availability of trustworthy data, reproducible evaluation, operational validation, and secure deployment.

5.8. Practical Relevance

The practical relevance of NTL detection methods extends beyond classification performance. In operational environments, automated detection systems are primarily expected to support utilities in prioritizing customers for subsequent inspection and thereby reducing the cost and effort associated with large-scale manual inspection.
This distinction is important because an automated model does not necessarily need to establish fraud conclusively. Instead, it can provide a risk or suspicion score that helps utilities identify a smaller and more targeted group of customers for physical investigation.
The cross-survey evidence indicates that practical deployment depends on several factors beyond predictive accuracy, including computational cost, detection latency, scalability, integration with AMI infrastructure, privacy, cybersecurity, interpretability, and the economic consequences of false positives and false negatives.
Consequently, future NTL detection research should evaluate models according to both technical effectiveness and operational utility. Cost-benefit analysis, real-world validation, and integration with existing inspection workflows should therefore receive greater attention.
An important practical implication is that increasingly sophisticated AI models should not automatically replace simpler approaches. The appropriate solution may instead involve a layered detection architecture in which statistical, machine-learning, anomaly-detection, and security mechanisms complement one another according to the characteristics of the deployment environment.

5.9. Quantitative Coverage

The qualitative synthesis is complemented by the quantitative coverage analysis presented in Section 3.8. For each research dimension, the coverage percentage represents the proportion of the 28 included secondary studies that explicitly address that dimension.
The coverage results reveal an uneven distribution of research attention across the NTL detection landscape. Machine-learning and AI-based approaches are explicitly discussed by 16 of the 28 secondary studies (57.1%), making them the most widely represented detection paradigm. In contrast, security and intrusion-detection approaches are explicitly covered by only 1 study (3.6%).
Similarly, smart-meter/AMI data are addressed by 11 studies (39.3%), while explicit identification of real-world evaluation data occurs in only 1 study (3.6%). This contrast highlights the difference between the importance of AMI as an infrastructure enabler and the limited availability of openly identified real-world datasets for reproducible research.
Within the evaluation dimension, 12 studies (42.9%) explicitly discuss performance or effectiveness, whereas only 2 studies (7.1%) address cost, response time, or execution time. This indicates that predictive performance receives substantially more attention than operational and economic considerations.
For research challenges, deployment and practical implementation are reported by 10 studies (35.7%), privacy and cybersecurity by 9 studies (32.1%), and data availability, quality, or complexity by 8 studies (28.6%). In contrast, interpretability is explicitly identified by only 1 study (3.6%) and regulatory or policy issues by 2 studies (7.1%).
These coverage statistics should not be interpreted as measures of research quality or effectiveness. Rather, they quantify the extent to which particular research dimensions are represented in the existing secondary literature. The results therefore provide an evidence map that complements the qualitative cross-survey synthesis.

5.10. Overall Cross-Survey Interpretation

The cross-survey analysis reveals a clear pattern in the evolution of NTL detection research. The field has progressed from conventional statistical and inspection-oriented approaches toward increasingly data-driven, machine-learning, deep-learning, hybrid, and security-aware approaches. However, methodological advancement has not been accompanied by equivalent standardization in datasets, evaluation protocols, operational validation, or deployment requirements.
The strongest convergence across the secondary literature concerns the importance of data-driven detection and the persistent challenges associated with data availability, privacy, cybersecurity, and deployment. The principal divergence concerns the relative effectiveness of different methodological families and the assumptions under which their reported performance is obtained.
The quantitative coverage analysis further demonstrates an imbalance between research attention to algorithmic development and research attention to operational deployment, interpretability, real-world validation, and regulatory considerations.
Taken together, these findings suggest that the next stage of NTL detection research should move beyond isolated improvements in classification performance toward reproducible, robust, secure, explainable, and deployment-oriented detection frameworks. Such frameworks should be evaluated using representative datasets, standardized protocols, multiple complementary metrics, and realistic utility-scale scenarios.

6. Discussion and Recommendations

The tertiary synthesis of 28 secondary studies reveals that NTL and electricity-theft detection has undergone a substantial methodological transition from conventional inspection and statistical approaches toward machine learning, deep learning, hybrid models, anomaly detection, and emerging generative approaches. However, the cross-survey analysis also reveals that methodological advancement has not been accompanied by equivalent progress in real-world validation, standardized datasets, evaluation protocols, interpretability, cybersecurity, and operational deployment.
The findings therefore suggest that the future development of NTL detection should not be driven solely by increasing algorithmic complexity. Instead, greater emphasis should be placed on reproducibility, generalizability, robustness, security, explainability, economic viability, and integration with existing utility inspection processes. The following subsections discuss the principal implications of the tertiary synthesis and identify concrete research priorities.

6.1. Major Findings

The tertiary synthesis leads to several major findings concerning the current state and future evolution of NTL detection.
First, the secondary literature remains predominantly focused on conventional electricity-theft mechanisms, particularly meter tampering, unauthorized connections, meter bypass, and anomalous consumption patterns. Cyber and data-oriented attacks receive substantially less attention. This indicates that the research community has not yet fully aligned NTL detection with the changing cyber-physical architecture of modern distribution networks.
Second, machine-learning and AI-based approaches represent the most widely covered methodological direction in the secondary literature. However, the evidence does not support the conclusion that one algorithmic family is universally superior. Reported performance is strongly dependent on the characteristics of the dataset, attack scenario, feature representation, class distribution, and evaluation protocol.
Third, AMI has become a major enabler of data-driven NTL detection by providing fine-grained consumption measurements and communication capabilities. At the same time, AMI introduces additional privacy, integrity, and cybersecurity vulnerabilities. AMI should therefore be viewed simultaneously as a source of detection information and as an additional attack surface.
Fourth, the availability of suitable real-world datasets remains a major limitation. The tertiary evidence reveals substantial heterogeneity in customer populations, sampling rates, geographical contexts, observation periods, and feature availability. This heterogeneity limits direct comparison of reported performance and weakens confidence in cross-region generalizability.
Fifth, evaluation practices remain predominantly performance-oriented. Accuracy, precision, recall, F1-score, and related classification metrics receive considerably greater attention than computational cost, inspection workload, economic benefit, detection latency, and deployment scalability. Consequently, high predictive performance cannot by itself establish practical utility.
Finally, emerging directions such as Generative AI, cyber-physical detection, and distributed cooperative detection provide promising opportunities but remain comparatively immature. These directions should therefore be investigated through rigorous validation rather than being presented as established solutions.

6.2. Implications for NTL Detection Research

The findings of the tertiary synthesis have several implications for future NTL detection research. A first implication is that NTL detection should increasingly be formulated as a multi-dimensional problem rather than as a binary classification task. Electricity theft can involve physical meter tampering, unauthorized connections, intermittent theft, manipulation of measurements, and cyberattacks against AMI infrastructure. These scenarios have different observability characteristics and cannot necessarily be detected using the same information source.
A second implication concerns the assumptions underlying consumption-based detection. Many data-driven approaches implicitly assume that fraudulent activity produces a measurable deviation in the observed consumption profile. This assumption may not hold for meter-less theft, partial bypass, downstream unauthorized connections, or manipulation of the measurements themselves. Future research should therefore explicitly characterize the observability of each attack scenario and identify which data sources are sufficient for its detection.
A third implication is that model complexity should be justified by the characteristics of the detection problem. Deep learning and generative models may provide powerful representation and data generation capabilities, but their additional computational and operational requirements should be evaluated against the benefits they provide.
Finally, NTL detection should increasingly be evaluated as a decision-support problem. The practical objective of an automated system is often not to legally establish fraud but to identify and prioritize suspicious customers for subsequent physical inspection. This distinction should be incorporated into future model design and evaluation.

6.3. Recommendations for Dataset Development

The lack of standardized and representative datasets is one of the most significant barriers identified by the tertiary synthesis. Future research should therefore prioritize the development of well-documented benchmark datasets for NTL detection. Such datasets should, where possible, include diverse customer populations, different geographical regions, multiple sampling resolutions, sufficiently long observation periods, and a variety of theft scenarios. Information concerning data provenance, sampling frequency, missing values, preprocessing, customer characteristics, and fraud-label generation should also be reported explicitly.
A particularly important requirement is the inclusion of realistic fraud labels. Confirmed theft cases are difficult to obtain because utilities generally treat customer consumption and inspection records as sensitive information. Privacy-preserving mechanisms should therefore be investigated to facilitate controlled data sharing while protecting customer confidentiality. Synthetic and augmented datasets can complement real-world datasets, particularly for rare theft scenarios and class-imbalance problems. However, synthetic data should not be treated as equivalent to operational data without validation. Their statistical and temporal properties should be compared against real consumption distributions, and models trained using synthetic samples should be evaluated on independent real-world data whenever possible. The development of common benchmark datasets would enable more reproducible comparisons and reduce the current dependence on single-dataset evaluations.

6.4. Recommendations for Standardized Evaluation

The substantial heterogeneity in evaluation practices identified by the tertiary synthesis indicates a need for more standardized evaluation protocols. Future studies should report multiple complementary performance measures rather than relying primarily on accuracy. For imbalanced NTL datasets, precision, recall, F1-score, detection rate, false-positive rate, and related minority-class measures should be reported. Where appropriate, precision-recall curves and receiver-operating-characteristic-based measures can provide additional information regarding detector behavior.
Evaluation should also extend beyond predictive performance. Utility deployment requires consideration of computational requirements, inference latency, communication overhead, scalability, inspection workload, false-positive costs, and economic benefits. A detector that provides high accuracy but generates a large number of unnecessary inspections may have limited operational value. Cross-dataset and cross-geographical validation should also become a standard component of evaluation. Models should ideally be tested on data that differ from the training environment to assess their ability to generalize to unseen customers, regions, and consumption patterns. Accordingly, future benchmarking should evaluate NTL detection using a multi-dimensional framework covering predictive performance, generalizability, computational efficiency, robustness, operational cost, and practical utility.

6.5. Cyber-Physical and AMI-Aware Detection

The increasing digitalization of distribution networks changes the threat landscape for NTL detection. Conventional electricity theft includes physical meter tampering, bypassing, and unauthorized connections, whereas AMI introduces additional attack surfaces involving smart-meter firmware, communication networks, gateways, concentrators, and data-management systems. Consequently, future NTL detection should increasingly be considered not only as a problem of identifying abnormal consumption but also as a problem of ensuring the integrity of the information used for detection. The existing literature already recognizes AMI-related privacy and cybersecurity concerns, but their integration with NTL detection remains comparatively limited.

6.5.1. False Data Injection Attacks

False Data Injection Attacks (FDIAs) represent an important emerging threat in which an adversary manipulates measurements or data streams used by grid monitoring and state-estimation processes. In the context of NTL detection, such manipulation could alter the information on which a detection model relies and potentially conceal abnormal consumption or create misleading consumption patterns. This creates an important distinction between physical electricity theft, in which electricity is bypassed or the meter is physically manipulated, and data-oriented theft, in which reported measurements, communication messages, or associated digital information are manipulated. Future research should therefore investigate the joint detection of consumption anomalies and measurement-integrity violations. Detection mechanisms could compare customer-level consumption patterns with temporal consistency, neighboring measurements, network-level constraints, and, where available, distribution-system state estimates.

6.5.2. AMI Communication and Smart-Meter Security

AMI communication channels and smart meters should be treated as security-critical components of NTL detection infrastructure. Attacks involving unauthorized access, message manipulation, compromised devices, or communication disruption may affect both the measurement process and the detection process itself. Future systems should therefore integrate security indicators with consumption-based detection. Authentication, encryption, secure key management, integrity verification, and appropriate data-governance mechanisms should be considered alongside detection performance. The objective should be to develop NTL detection systems that remain reliable even when parts of the measurement or communication infrastructure are compromised.

6.6. Explainable and Trustworthy AI

Interpretability is comparatively underrepresented in the secondary literature despite its practical importance. Automated NTL detection systems may be used to prioritize customers for physical inspection, and utility operators may therefore require understandable evidence supporting a suspicious classification. Future research should investigate explainable AI mechanisms capable of identifying the consumption patterns, temporal characteristics, or other evidence that contributed to a detection decision. Explanation mechanisms should be evaluated for both fidelity and usefulness rather than being treated merely as visualization tools.
Trustworthiness should also encompass robustness, uncertainty, fairness, privacy, and security. A highly accurate model that is vulnerable to adversarial manipulation or produces unreliable predictions under distribution shifts may not be suitable for operational deployment. Uncertainty estimation is particularly relevant because NTL detection systems often operate under incomplete knowledge of customer behavior and limited confirmed fraud labels. Models should therefore be able to distinguish between confident detections and cases requiring additional investigation.
Explainable and trustworthy AI can consequently support a human-in-the-loop inspection process in which automated predictions assist utility personnel without replacing the final inspection and verification process.

6.7. Role of Generative AI

The tertiary evidence indicates that Generative AI should currently be regarded as an emerging research direction rather than a dominant paradigm for NTL detection. Its relatively limited representation across the secondary literature does not justify presenting it as a mature replacement for established statistical or machine-learning methods. The most immediate opportunity for generative models lies in addressing data-related limitations. Time-series generative models can potentially generate representative consumption sequences, augment minority fraudulent classes, reconstruct missing observations, and support stress-testing of detection systems under rare theft scenarios. However, generated data introduce additional methodological concerns. Synthetic samples may reproduce biases present in the training data, fail to capture realistic theft behavior, or create distributions that are insufficiently representative of operational consumption. The risk of data leakage and the computational cost of training generative models must also be considered.
Consequently, future research should evaluate generative models using independent real-world datasets and explicitly assess distributional fidelity, privacy, data leakage, computational requirements, and downstream detection performance.
Large language models (LLMs) should be considered separately from consumption-generation models. Their near-term role in NTL management may be more realistic as an auxiliary tool for anomaly summarization, inspection-report drafting, knowledge retrieval, documentation, or regulatory support rather than as a direct replacement for consumption-based detection models.
Accordingly, Generative AI should be positioned as a complementary technology that addresses specific limitations of existing NTL detection workflows rather than as a universally superior detection paradigm.

6.8. Towards Integrated NTL Detection Frameworks

The cross-survey synthesis suggests that future NTL detection should move from isolated algorithmic solutions toward integrated detection frameworks. Physical theft, consumption anomalies, meter manipulation, and cyberattacks may produce different observable patterns and therefore require complementary detection mechanisms. A practical architecture could combine four complementary layers: (i) consumption-based anomaly detection, (ii) machine-learning or deep-learning classification, (iii) AMI and cybersecurity monitoring, and (iv) network- or physics-based consistency analysis. Such a layered architecture would reduce dependence on a single detection assumption and provide multiple sources of evidence for prioritizing suspicious customers. The role of automated detection should also be aligned with the operational inspection process. Rather than attempting to conclusively establish fraud using an automated model, the system can generate a risk score and provide supporting evidence to prioritize physical inspection. This approach is particularly relevant in regulatory contexts where physical inspection may remain necessary for formal confirmation of fraud.

Distributed and Cooperative Detection

An emerging extension of integrated NTL detection is the use of distributed and cooperative architectures. Smart meters, edge devices, or local distribution-network components could be modeled as cooperative agents that process local consumption information and exchange selected information with neighboring agents. Instead of transmitting all raw measurements to a central control center, agents could exchange aggregated features, anomaly scores, model updates, or state estimates through local communication. Consensus or distributed-optimization mechanisms could then be used to derive a coordinated assessment of abnormal consumption. Such architectures could improve scalability, reduce communication requirements, and provide additional privacy by retaining raw consumption data locally. They may also improve resilience to failures of a centralized processing architecture.
However, distributed cooperation introduces challenges including communication delays, packet loss, heterogeneous device capabilities, asynchronous updates, unreliable or malicious agents, privacy leakage, consensus convergence, and error propagation. Therefore, distributed multi-agent NTL detection should currently be regarded as an emerging research direction rather than an established evidence-supported methodology. Future studies should compare centralized and cooperative approaches using realistic AMI topologies and evaluate detection accuracy, communication cost, energy consumption, privacy, robustness, convergence, and resilience to compromised agents.

6.9. Research Priorities

Based on the tertiary synthesis, the following research priorities are identified:
  • Standardized benchmark datasets: Develop representative, well-documented, privacy-preserving datasets covering multiple geographical regions, customer populations, sampling resolutions, and theft scenarios.
  • Real-world validation: Increase validation using operational utility data and evaluate models under realistic changes in consumption behavior, seasonal patterns, customer populations, and theft strategies.
  • Standardized evaluation protocols: Establish common benchmarking procedures that report predictive, computational, operational, and economic measures.
  • Cross-region generalization: Evaluate models across geographical regions, utilities, tariff structures, and customer populations to determine whether learned patterns generalize beyond the training environment.
  • Cyber-physical resilience: Integrate physical theft detection with measurement-integrity, communication-security, smart-meter-security, and false-data injection detection.
  • Explainable and trustworthy detection: Develop models that provide interpretable evidence, quantify uncertainty, and remain robust under distribution shifts and adversarial conditions.
  • Responsible use of Generative AI: Investigate generative models primarily for data augmentation, rare-event generation, missing-data reconstruction, and stress-testing, with explicit validation against independent real-world data.
  • Operational and economic evaluation: Incorporate inspection workload, false-positive costs, deployment cost, computational requirements, response time, and return on investment into NTL detection evaluation.
  • Distributed and cooperative detection: Investigate privacy-preserving distributed architectures in which neighboring AMI devices or edge nodes collaboratively detect anomalies without requiring all raw consumption data to be transmitted to a central server.
  • Human-in-the-loop decision support: Design detection systems that support utility personnel in prioritizing inspections while retaining appropriate human and regulatory oversight.
Collectively, these priorities indicate that future progress in NTL detection should be measured not only by improvements in predictive accuracy but also by improvements in reproducibility, robustness, security, generalizability, explainability, and operational value.
The tertiary synthesis shows that NTL detection is evolving from consumption-based analysis toward integrated, secure, explainable, and deployment-oriented frameworks. Although ML/AI and AMI-based methods are increasingly prominent, heterogeneous datasets, evaluation protocols, and limited real-world validation prevent direct comparison of approaches. Future systems should jointly address data quality, standardized evaluation, cyber-physical security, privacy, interpretability, and operational cost. Generative AI and distributed cooperative detection are promising but remain emerging directions requiring rigorous validation.

7. Threats to Validity and Limitations

Although the present study follows a structured PRISMA-based procedure and applies a predefined framework for the synthesis of secondary studies, several limitations should be considered when interpreting the findings. These limitations arise from the characteristics of the available secondary literature, the review process, the heterogeneity of the evidence, and the scope of the quantitative synthesis.

7.1. Selection and Publication Bias

The identification and selection of secondary studies may be affected by publication and indexing biases. The search strategy was designed to identify relevant surveys, systematic reviews, comprehensive reviews, and related secondary studies; however, no search strategy can guarantee complete coverage of all the published or unpublished literature. In particular, secondary studies that are poorly indexed, published in less visible venues, or use terminology different from the search terms may not have been identified. In addition, published reviews may overrepresent research areas that have received greater attention, while emerging or negative findings may be comparatively underrepresented. To reduce this risk, the study employed predefined eligibility criteria and a structured screening procedure. Nevertheless, the possibility of selection and publication bias cannot be completely excluded.

7.2. Heterogeneity of the Included Secondary Studies

The 28 included secondary studies differ substantially in their scope, objectives, terminology, methodological approaches, and levels of coverage. The evidence base includes surveys, systematic reviews, comprehensive reviews, and state-of-the-art or narrative reviews. This heterogeneity makes direct comparison of individual conclusions difficult. For example, some reviews primarily address physical electricity theft and consumption anomalies, whereas others focus on machine learning, deep learning, AMI security, privacy, adversarial attacks, or broader smart-grid applications. The common analytical framework developed in this study was therefore used to harmonize these heterogeneous perspectives. However, such harmonization necessarily involves abstraction and may not capture every methodological distinction made by the original secondary studies.

7.3. Limitations of the Quality Assessment

The methodological quality of the included secondary studies was assessed using six predefined criteria covering clarity of objectives, search-strategy transparency, eligibility criteria, study-selection transparency, systematicity of synthesis, and reporting of limitations or research gaps. Although this checklist provides a common basis for assessing heterogeneous secondary studies, it is not intended to replace a specialized instrument designed for a particular review methodology. The included literature contains different types of secondary studies, and therefore a single instrument cannot fully capture all aspects of methodological quality.
Furthermore, the quality score is used to characterize the strength of the evidence and inform interpretation rather than to automatically exclude studies. Consequently, conclusions should be interpreted in light of both the reported findings and the methodological quality of the underlying secondary studies.

7.4. Screening and Data-Extraction Reliability

The screening and data-extraction processes were conducted independently by two researchers using the predefined eligibility criteria and structured data-extraction framework. Each researcher independently assessed the identified studies and extracted the information required for the tertiary synthesis, including study characteristics, attack types, detection approaches, data and AMI requirements, evaluation practices, and reported challenges and research gaps. The results of the two researchers were subsequently compared to identify disagreements in study selection, classification, or data extraction. Disagreements were resolved through discussion and consensus between the researchers. Where necessary, the original publication was re-examined to determine the appropriate classification.
The use of independent screening and data extraction by two researchers was intended to reduce individual selection and interpretation bias and to improve the consistency and reliability of the tertiary synthesis. The same procedure was applied to the mapping of heterogeneous findings into the common analytical dimensions corresponding to RQ1–RQ5.

7.5. Dataset and Evaluation Heterogeneity

The secondary literature relies on highly heterogeneous datasets, including datasets that differ in customer population, geographical origin, sampling frequency, observation period, feature availability, and data accessibility. This heterogeneity prevents direct comparison of reported predictive performance across studies. Similarly, the reviewed studies employ different evaluation metrics, experimental protocols, attack assumptions, class distributions, and validation strategies. Consequently, the present study does not interpret a reported accuracy, precision, recall, or F1-score from one secondary study as directly comparable with the corresponding value from another study unless the underlying experimental conditions are sufficiently aligned. These limitations also explain why the quantitative synthesis focuses on research-dimension coverage rather than statistical pooling of predictive performance measures.

7.6. Limitations of the Quantitative Cross-Survey Analysis

The quantitative cross-survey analysis measures the extent to which specific research dimensions are represented across the 28 secondary studies. The resulting coverage percentages should therefore be interpreted as descriptive evidence maps rather than measures of research quality, methodological effectiveness, or importance. A higher coverage percentage indicates that a dimension is explicitly addressed by more secondary studies; it does not imply that the corresponding methods are more accurate or practically superior. Similarly, a low coverage percentage does not necessarily indicate that a research direction is unimportant. It may instead reflect its emerging nature or limited treatment in the existing secondary literature. Furthermore, several dimensions can be discussed within the same secondary study. Therefore, the coverage percentages within an RQ are not necessarily mutually exclusive and should not be interpreted as parts of a single probability distribution.

7.7. Limited Evidence for Emerging Research Directions

Several emerging directions discussed in this study, including Generative AI, cyber-physical attacks, false-data injection, and distributed cooperative detection, are represented relatively infrequently in the existing secondary literature. Consequently, these topics should not be interpreted as established dominant paradigms for NTL detection. Their inclusion in the discussion reflects their potential relevance to the future evolution of AMI-enabled electricity-theft detection rather than the existence of strong and mature evidence. In particular, the tertiary synthesis does not provide sufficient evidence to conclude that cyberattacks will replace physical electricity theft or that Generative AI will outperform established detection approaches. These directions require further empirical validation using realistic datasets, operational AMI architectures, and reproducible evaluation protocols.

7.8. Scope of the Tertiary Review

The present study focuses on the secondary literature concerning non-technical losses and electricity-theft detection. It is therefore not intended to provide an exhaustive systematic synthesis of every primary NTL detection algorithm, dataset, or experimental result. The conclusions are instead derived from how existing secondary studies characterize the field. Consequently, individual primary studies may contain more recent, detailed, or contradictory evidence that is not fully represented in the secondary literature included in the present tertiary synthesis. This limitation is inherent to the tertiary-review design and should be considered when interpreting conclusions concerning rapidly evolving research directions.

7.9. Overall Threats to Validity

Taken together, the principal threats to validity arise from heterogeneity in the secondary evidence, potential selection and publication bias, single-researcher screening and extraction, differences in dataset and evaluation practices, and the limited maturity of several emerging research directions. These limitations do not invalidate the tertiary synthesis but define the appropriate interpretation of its findings. The quantitative coverage analysis should be understood as a descriptive representation of research attention, while the qualitative synthesis should be interpreted as a structured comparison of heterogeneous secondary evidence rather than as a statistical estimate of the effectiveness of NTL detection methods.
Future tertiary studies can strengthen the evidence base by employing independent duplicate screening and extraction, broader database coverage, standardized quality-assessment instruments, preregistered protocols, and continued updates as the secondary literature evolves.

8. Conclusions

This study presented a tertiary synthesis of the secondary literature on NTL and electricity-theft detection. By analyzing 28 eligible surveys, systematic reviews, comprehensive reviews, and related secondary studies, the study examined the evolution of NTL detection from conventional and statistical approaches toward machine-learning, deep-learning, hybrid, anomaly-detection, and emerging AI-based approaches. The synthesis shows that machine-learning and AI-based methods represent the most widely covered methodological direction in the secondary literature. However, the evidence does not support the existence of a universally superior detection approach. The effectiveness of a method depends strongly on the characteristics of the available data, the assumed theft mechanism, feature representation, class distribution, and evaluation protocol. Meter tampering, unauthorized connections, and consumption-based anomalies remain the most frequently represented attack categories, whereas cyber and data attacks receive comparatively limited attention.
A major finding of the tertiary analysis is the substantial heterogeneity of datasets and evaluation practices. Differences in customer populations, sampling frequencies, geographical settings, observation periods, and data availability limit direct comparison and generalization of reported results. The limited availability of representative real-world theft datasets further constrains reproducibility and benchmarking. The analysis therefore highlights the need for standardized datasets, cross-dataset validation, and common evaluation protocols that consider not only predictive performance but also computational requirements, operational cost, scalability, and practical deployment.
The study also identifies an important transition in the threat landscape associated with the increasing digitalization of distribution networks. AMI provides valuable fine-grained information for NTL detection, but simultaneously introduces additional vulnerabilities related to measurement integrity, communication security, privacy, and smart-meter compromise. Future NTL detection systems should therefore consider physical and cyber-physical threats jointly rather than relying exclusively on consumption-profile analysis.
Emerging technologies such as Generative AI, synthetic data generation, and distributed cooperative detection provide promising opportunities, particularly for addressing data scarcity, class imbalance, missing observations, and decentralized processing. However, the tertiary evidence indicates that these directions remain relatively immature. Their potential should therefore be assessed through rigorous experimentation using realistic and independently validated datasets, rather than being considered established alternatives to existing detection approaches.
Beyond summarizing individual surveys, this study provides scientific information that emerges from the comparison of the secondary literature as a whole. First, it harmonizes heterogeneous taxonomies into a common analytical framework covering attack types, detection approaches, data and feature requirements, AMI dependence, evaluation practices, and research gaps. Second, it quantitatively measures the coverage of these dimensions across the 28 secondary studies. Third, it distinguishes recurring research challenges from observations reported only by individual reviews. Finally, it connects methodological choices with data availability and infrastructure digitalization, providing a broader context for understanding the applicability and limitations of different NTL detection strategies.
The resulting evidence indicates that the next stage of NTL detection research should move beyond isolated improvements in classification accuracy toward reliable, explainable, secure, generalizable, and deployment-oriented solutions. Future research should prioritize standardized benchmark datasets, real-world validation, cross-geographical evaluation, cyber-physical resilience, privacy-preserving mechanisms, trustworthy AI, and comprehensive operational and economic assessment.
Overall, the contribution of this work is therefore not the proposal of another NTL detection algorithm, but the construction of a cross-survey evidence synthesis that reveals patterns, divergences, structural gaps, and emerging research directions across the existing secondary literature. Such a perspective can support researchers in identifying well-grounded research opportunities and assist utilities and other stakeholders in understanding the requirements for more robust and sustainable NTL detection systems.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/bdcc10090297/s1. Supplementary Material S1: Completed PRISMA 2020 checklist.

Author Contributions

Conceptualization, D.L. and D.T.; methodology, D.L. and D.T.; software, D.L. and D.T.; validation, D.L., D.T., A.G. and G.F.; formal analysis, D.L. and D.T.; investigation, D.L.; resources, D.L. and D.T.; data curation, D.L.; writing—original draft preparation, D.L.; writing—review and editing, D.L., D.T. and A.G.; visualization, D.L. and D.T.; supervision, D.T., A.G. and G.F.; project administration, A.G. and G.F.; funding acquisition, A.G. and G.F. All authors have read and agreed to the published version of the manuscript.

Funding

This work contributes to the basic research activities of the PNRR project FAIR—Future AI Research (PE00000013), Spoke 9—Green-aware AI, under the NRRP MUR program funded by the NextGenerationEU.

Institutional Review Board Statement

Institutional Review Board approval was not required for this study, as the research did not involve human participants or the collection or use of identifiable human data.

Informed Consent Statement

Informed consent was not applicable because this study did not involve human participants or identifiable personal data.

Data Availability Statement

The data used in this study are publicly available and can be accessed from the sources cited in the manuscript.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
NTLsNon-technical losses
AMIAdvanced Metering Infrastructure
DRDetection Rate
SGsSmart Grids
PLCPower Line Communication
SGCCSmart Grid Corporation of China
LLMsLarge Language Models
ANNArtificial Neural Network
MLPMulti-Layer Perceptron
CNNConvolutional Neural Network
LSTMLong Short-Term Memory
GRUsGated Recurrent Units
RNNsRecurrent Neural Networks
PRECONPakistan Residential Electricity Consumption
CERCommission for Energy Regulation
RUSRandom Under Sampling
ROSRandom Oversampling
SMOTESynthetic Minority Oversampling Technique
CBOSCluster-based Oversampling
VAEsVariational Autoencoders
GANsGenerative Adversarial Networks
SVMSupport Vector Machine
PCAPrincipal Component Analysis
ARERA   Regulatory Authority for Energy, Networks and Environment
AUCArea Under the Curve
FPR          False Positive Rate
BDRBayesian Detection Rate
DSODistribution System Operator

References

  1. Labate, D.; Thakur, D.; Fortino, G. Adaptive Sensitivity-Aware Differential Privacy Accounting for Federated Smart-Meter Theft Detection. Big Data Cogn. Comput. 2026, 10, 113. [Google Scholar] [CrossRef] [Scilit]
  2. Dhupia, B.; Usha Rani, M.; Alameen, A. The Role of Big Data Analytics in Smart Grid Management. In Emerging Research in Data Engineering Systems and Computer Communications; Venkata Krishna, P., Obaidat, M.S., Eds.; Springer: Singapore, 2020; pp. 403–412. [Google Scholar] [CrossRef] [Scilit]
  3. Theron-Ord, A. Electricity Theft and Non-Technical Losses Total $96bn Annually—Report. Available online: https://www.smart-energy.com/regional-news/africa-middle-east/electricity-theft-96bn-annually/ (accessed on 16 December 2024).
  4. Pereira, J.; Saraiva, F. Convolutional Neural Network Applied to Detect Electricity Theft: A Comparative Study on Unbalanced Data Handling Techniques. Int. J. Electr. Power Energy Syst. 2021, 131, 107085. [Google Scholar] [CrossRef] [Scilit]
  5. Smith, T.B. Electricity theft: A comparative analysis. Energy Policy 2004, 32, 2067–2076. [Google Scholar] [CrossRef] [Scilit]
  6. Guerrero, J.I.; Monedero, I.; Biscarri, F.; Biscarri, J.; Millán, R.; León, C. Non-Technical Losses Reduction by Improving the Inspections Accuracy in a Power Utility. IEEE Trans. Power Syst. 2018, 33, 1209–1218. [Google Scholar] [CrossRef] [Scilit]
  7. Labate, D.; Giubbini, P.; Chicco, G.; Piglione, F.; Italy, P. Shape: The load prediction and non-technical losses modules. In Proceedings of the 23rd International Conference on Electricity Distribution (CIRED 2015), Lyon, France, 15–18 June 2015. [Google Scholar]
  8. Xia, X.; Yang, X.; Wei, L.; Cui, J. Detection Methods in Smart Meters for Electricity Thefts: A Survey. Proc. IEEE 2022, 110, 273–319. [Google Scholar] [CrossRef] [Scilit]
  9. Transforma Insights IoT Forecast Database. Available online: https://transformainsights.com/news/global-smart-meters-2033 (accessed on 17 December 2024).
  10. Jiang, R.; Lu, R.; Wang, Y.; Luo, J.; Shen, C.; Shen, X. Energy-theft detection issues for advanced metering infrastructure in smart grid. Tsinghua Sci. Technol. 2014, 19, 105–120. [Google Scholar] [CrossRef] [Scilit]
  11. Barai, G.R.; Krishnan, S.; Venkatesh, B. Smart metering and functionalities of smart meters in smart grid—A review. In 2015 IEEE Electrical Power and Energy Conference (EPEC), London, ON, Canada; IEEE: New York, NY, USA, 2015; pp. 138–145. [Google Scholar] [CrossRef] [Scilit]
  12. Shokry, M.; Awad, A.I.; Abd-Ellah, M.K.; Khalaf, A.A.M. Systematic survey of advanced metering infrastructure security: Vulnerabilities, attacks, countermeasures, and future vision. Future Gener. Comput. Syst. 2022, 136, 358–377. [Google Scholar] [CrossRef] [Scilit]
  13. Shehzad, F.; Asif, M.; Aslam, Z.; Anwar, S.; Rashid, H.; Ilyas, M.; Javaid, N. Comparative Study of Data Driven Approaches Towards Efficient Electricity Theft Detection in Micro Grids. In Innovative Mobile and Internet Services in Ubiquitous Computing; Barolli, L., Yim, K., Chen, H.-C., Eds.; Lecture Notes in Networks and Systems; Springer International Publishing: Cham, Switzerland, 2022; Volume 279, pp. 120–131. [Google Scholar] [CrossRef] [Scilit]
  14. Glauner, P.; Meira, J.A.; Valtchev, P.; State, R.; Bettinger, F. The Challenge of Non-Technical Loss Detection Using Artificial Intelligence: A Survey. Int. J. Comput. Intell. Syst. 2017, 10, 760. [Google Scholar] [CrossRef] [Scilit]
  15. Pal, K. Review of Non-Technical Losses Identification Techniques. Int. J. Recent Innov. Trends Comput. Commun. 2021, 9, 7–22. [Google Scholar] [CrossRef] [Scilit]
  16. Pereira, J.; Saraiva, F. A Comparative Analysis of Unbalanced Data Handling Techniques for Machine Learning Algorithms to Electricity Theft Detection. In 2020 IEEE Congress on Evolutionary Computation (CEC), Glasgow, UK; IEEE: New York, NY, USA, 2020; pp. 1–8. [Google Scholar] [CrossRef] [Scilit]
  17. Pazderin, A.; Kamalov, F.; Gubin, P.Y.; Safaraliev, M.; Samoylenko, V.; Mukhlynin, N.; Odinaev, I.; Zicmane, I. Data-Driven Machine Learning Methods for Nontechnical Losses of Electrical Energy Detection: A State-of-the-Art Review. Energies 2023, 16, 7460. [Google Scholar] [CrossRef] [Scilit]
  18. Kolade, A.O.; Adetokun, B.B.; Oghorada, O. Energy Theft Detection in Power System Network: Reviews of Studies on Machine Learning Based Solutions. In 2023 2nd International Conference on Multidisciplinary Engineering and Applied Science (ICMEAS), Abuja, Nigeria; IEEE: New York, NY, USA, 2023; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  19. Hashim, M.; Khan, L.; Javaid, N.; Ullah, Z.; Javed, A. Stacked Machine Learning Models for Non-Technical Loss Detection in Smart Grid: A Comparative Analysis. Energy Rep. 2024, 12, 1235–1253. [Google Scholar] [CrossRef] [Scilit]
  20. Gunturi, S.K.; Sarkar, D. Ensemble Machine Learning Models for Energy Theft Detection. Electr. Power Syst. Res. 2021, 191, 106904. [Google Scholar] [CrossRef] [Scilit]
  21. Kabir, B.; Qasim, U.; Javaid, N.; Khan, F.A.; Mansoor, B.; Saudagar, A.K.J.; AlSagri, H.S.; Saqib, M.N. Detecting Electricity Theft in Smart Grids Using Optimized Machine Learning Ensemble Techniques and eXplainable AI. Energy Rep. 2025, 14, 3142–3162. [Google Scholar] [CrossRef] [Scilit]
  22. Abro, S.A.; Laghari, J.A.; Memon, S.A.; Khan, T.A.; Memon, I.; Nasir, H.; Fatima, K. Non-Technical Loss Detection in Power Distribution Networks Using Machine Learning. Sci. Rep. 2025, 15, 36189. [Google Scholar] [CrossRef] [Scilit]
  23. Althobaiti, A.; Jindal, A.; Marnerides, A.K.; Roedig, U. Energy Theft in Smart Grids: A Survey on Data-Driven Attack Strategies and Detection Methods. IEEE Access 2021, 9, 159291–159312. [Google Scholar] [CrossRef] [Scilit]
  24. Xia, X.; Lin, J.; Jia, Q.; Wang, X.; Ma, C.; Cui, J.; Liang, W. ETD-ConvLSTM: A Deep Learning Approach for Electricity Theft Detection in Smart Grids. IEEE Trans. Inf. Forensics Secur. 2023, 18, 2553–2568. [Google Scholar] [CrossRef] [Scilit]
  25. Tursunboev, J.; Palakonda, V.; Kang, J.-M. Multi-Objective Evolutionary Hybrid Deep Learning for Energy Theft Detection. Appl. Energy 2024, 363, 122847. [Google Scholar] [CrossRef] [Scilit]
  26. Huang, Q.; Tang, Z.; Weng, X.; He, M.; Liu, F.; Yang, M.; Jin, T. A Novel Electricity Theft Detection Strategy Based on Dual-Time Feature Fusion and Deep Learning Methods. Energies 2024, 17, 275. [Google Scholar] [CrossRef] [Scilit]
  27. Elshennawy, N.M.; Ibrahim, D.M.; Gab Allah, A.M. An Efficient Electricity Theft Detection Based on Deep Learning. Sci. Rep. 2025, 15, 12866. [Google Scholar] [CrossRef] [Scilit]
  28. Gunduz, M.Z.; Das, R. Smart Grid Security: An Effective Hybrid CNN-Based Approach for Detecting Energy Theft Using Consumption Patterns. Sensors 2024, 24, 1148. [Google Scholar] [CrossRef] [Scilit]
  29. Elgarhy, I.; Badr, M.M.E.A.; Mahmoud, M.M.E.A.; Fouda, M.M.; Alsabaan, M.; Kholidy, H.A. Clustering and Ensemble-Based Approach for Securing Electricity Theft Detectors Against Evasion Attacks. IEEE Access 2023, 11, 112147–112164. [Google Scholar] [CrossRef] [Scilit]
  30. Massarani, A.H.; Badr, M.M.; Baza, M.; Alshahrani, H.; Alshehri, A. Efficient and Accurate Zero-Day Electricity Theft Detection from Smart Meter Sensor Data Using Prototype and Ensemble Learning. Sensors 2025, 25, 4111. [Google Scholar] [CrossRef] [Scilit]
  31. Badr, M.M.; Baza, M.; Rasheed, A.; Kholidy, H.; Abdelfattah, S.; Zaman, T.S. Comparative Analysis between Supervised and Anomaly Detectors Against Electricity Theft Zero-Day Attacks. In 2024 International Telecommunications Conference (ITC-Egypt), Cairo, Egypt; IEEE: New York, NY, USA, 2024; pp. 706–711. [Google Scholar] [CrossRef] [Scilit]
  32. Yang, L.; Chen, Z.; Wu, T. Multi-Step Diffusion Model with Self-Supervised Pretraining for Electricity Theft Detection. IEEE Trans. Smart Grid 2025, 16, 2439–2450. [Google Scholar] [CrossRef] [Scilit]
  33. Kim, S.; Sun, Y.; Lee, S.; Seon, J.; Hwang, B.; Kim, J.; Kim, J.; Kim, K.; Kim, J. Data-Driven Approaches for Energy Theft Detection: A Comprehensive Review. Energies 2024, 17, 3057. [Google Scholar] [CrossRef] [Scilit]
  34. Kgaphola, P.M.; Marebane, S.M.; Hans, R.T. Electricity Theft Detection and Prevention Using Technology-Based Models: A Systematic Literature Review. Electricity 2024, 5, 334–350. [Google Scholar] [CrossRef] [Scilit]
  35. Naidji, I.; Choucha, C.E.; Ramdani, M. Electricity Theft Detection Techniques Using Artificial Intelligence: A Survey. In 2024 IEEE International Conference on Advanced Systems and Emergent Technologies (ICASET), Hammamet, Tunisia; IEEE: New York, NY, USA, 2024; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  36. Badr, M.; Ibrahem, M.; Kholidy, H.; Fouda, M.; Ismail, M. Review of the Data-Driven Methods for Electricity Fraud Detection in Smart Metering Systems. Energies 2023, 16, 2852. [Google Scholar] [CrossRef] [Scilit]
  37. Ahmed, M.; Khan, A.; Ahmed, M.; Tahir, M.; Jeon, G.; Fortino, G.; Piccialli, F. Energy Theft Detection in Smart Grids: Taxonomy, Comparative Analysis, Challenges, and Future Research Directions. IEEE/CAA J. Autom. Sin. 2022, 9, 578–600. [Google Scholar] [CrossRef] [Scilit]
  38. Savian, F.D.S.; Siluk, J.C.M.; Garlet, T.B.; Nascimento, F.M.D.; Pinheiro, J.R.; Vale, Z. Non-Technical Losses: A Systematic Contemporary Article Review. Renew. Sustain. Energy Rev. 2021, 147, 111205. [Google Scholar] [CrossRef] [Scilit]
  39. Stracqualursi, E.; Rosato, A.; Lorenzo, G.D.; Panella, M.; Araneo, R. Systematic Review of Energy Theft Practices and Autonomous Detection through Artificial Intelligence Methods. Renew. Sustain. Energy Rev. 2023, 184, 113544. [Google Scholar] [CrossRef] [Scilit]
  40. Saeed, M.S.; Mustafa, M.W.; Hamadneh, N.N.; Alshammari, N.A.; Sheikh, U.U.; Jumani, T.A.; Khalid, S.B.A.; Khan, I. Detection of Non-Technical Losses in Power Utilities—A Comprehensive Systematic Review. Energies 2020, 13, 4727. [Google Scholar] [CrossRef] [Scilit]
  41. Chauhan; Abhishek; Rajvanshi, S. Non-Technical Losses in Power System: A Review. In 2013 International Conference on Power, Energy and Control (ICPEC), Dindigul, India; IEEE: New York, NY, USA, 2013; pp. 558–561. [Google Scholar] [CrossRef] [Scilit]
  42. Viegas, J.L.; Esteves, P.R.; Melício, R.; Mendes, V.M.F.; Vieira, S.M. Solutions for Detection of Non-Technical Losses in the Electricity Grid: A Review. Renew. Sustain. Energy Rev. 2017, 80, 1256–1268. [Google Scholar] [CrossRef] [Scilit]
  43. Messinis, G.M.; Hatziargyriou, N.D. Review of Non-Technical Loss Detection Methods. Electr. Power Syst. Res. 2018, 158, 250–266. [Google Scholar] [CrossRef] [Scilit]
  44. Ahmad, T.; Chen, H.; Wang, J.; Guo, Y. Review of Various Modeling Techniques for the Detection of Electricity Theft in Smart Grid Environment. Renew. Sustain. Energy Rev. 2018, 82, 2916–2933. [Google Scholar] [CrossRef] [Scilit]
  45. Hammerschmitt, B.K.; Abaide, A.D.R.; Lucchese, F.C.; Martins, C.C.; Da Silveira, A.S.; Rigodanzo, J.; Castro, J.V.M.B.; Rohr, J.A.D.A. Non-Technical Losses Review and Possible Methodology Solutions. In 2020 6th International Conference on Electric Power and Energy Conversion Systems (EPECS), Istanbul, Turkey; IEEE: New York, NY, USA, 2020; pp. 64–68. [Google Scholar] [CrossRef] [Scilit]
  46. Chuwa, M.G.; Wang, F. A Review of Non-Technical Loss Attack Models and Detection Methods in the Smart Grid. Electr. Power Syst. Res. 2021, 199, 107415. [Google Scholar] [CrossRef] [Scilit]
  47. Pealy, S.; Matin, M.A. Tackling Energy Theft in Smart Grid-A Comprehensive Review and Framework. In 2021 International Conference on Control, Automation, Power and Signal Processing (CAPS), Jabalpur, India; IEEE: New York, NY, USA, 2021; pp. 1–5. [Google Scholar] [CrossRef] [Scilit]
  48. Yadav, R.; Kumar, Y. The Detection of Non-Technical Losses and Electricity Theft by Smart Meter Data and Artificial Intelligence in the Context of Electric Distribution Utilities: A Comprehensive Review. Int. J. Comput. Digit. Syst. 2022, 12, 731–740. [Google Scholar] [CrossRef] [Scilit]
  49. Guarda, F.; Hammerschmitt, B.; Capeletti, M.; Neto, N.; Santos, L.D.; Prade, L.; Abaide, A. Non-Hardware-Based Non-Technical Losses Detection Methods: A Review. Energies 2023, 16, 2054. [Google Scholar] [CrossRef] [Scilit]
  50. Haruna, U.; Pal, B.L.; Dhabariya, A.S.; Rasheed, F.; Shah, A.F.; Sani, A.; Mu’azu, B.S.; Yahya, A.A. Review on Temporal Convolutional Networks for Electricity Theft Detection with Limited Data. J. Br. J. Comput. Netw. Inf. Technol. (BJCNIT) 2024, 7, 94–106. [Google Scholar] [CrossRef] [Scilit]
  51. Nayak, R.; C D, J. Data-driven models for electricity theft and anomalous power consumption detection: A systematic review. Appl. Intell. 2025, 55, 794. [Google Scholar] [CrossRef] [Scilit]
  52. Iqbal, M.S.; Munawar, S.; Adnan, M.; Raza, A.; Akbar, M.A.; Bermak, A. A critical review of technical case studies for electricity theft detection in smart grids: A new paradigm based transformative approach. Energy Convers. Manag. X 2025, 26, 100965. [Google Scholar] [CrossRef] [Scilit]
  53. Morgoev, I.D.; Klyuev, R.V.; Morgoeva, A.D. Literature Review of Methods for Detecting Non-Technical Electricity Losses in Distribution Grids. Artif. Intell. Appl. 2026. [Google Scholar] [CrossRef] [Scilit]
  54. Ghori, K.M.; Awais, M.; Khattak, A.S.; Imran, M.; Abbasi, R.A.; Szathmary, L. A Review on Latest Trends in Non-Technical Loss Detection. In Proceedings of the 1st Conference on Information Technology and Data Science, Debrecen, Hungary, 6–8 November 2020; Available online: https://ceur-ws.org/Vol-2874/short12.pdf (accessed on 2 January 2025).
  55. Razavi, R.; Gharavi, H.; Ghafouri-Fard, H. Practical Feature Engineering Framework for Electricity Theft Detection in Smart Grids. Appl. Energy 2019, 238, 107–120. [Google Scholar] [CrossRef] [Scilit]
  56. Zhang, W.; Dai, Y. A Multiscale Electricity Theft Detection Model Based on Feature Engineering. Big Data Res. 2024, 36, 100457. [Google Scholar] [CrossRef] [Scilit]
  57. Massaferro, P.; Parisio, A.; Strbac, G. Maximizing Economic Return in Electricity Theft Detection Using Cost-Sensitive Learning. IEEE Trans. Power Syst. 2020, 35, 1342–1352. [Google Scholar] [CrossRef] [Scilit]
  58. Ge, L.; Li, J.; Du, T.; Hou, L. Double-layer stacking optimization for electricity theft detection considering data incompleteness and intra-class imbalance. Int. J. Electr. Power Energy Syst. 2025, 165, 110461. [Google Scholar] [CrossRef] [Scilit]
  59. Sebastian, P.K.; Deepa, K.; Neelima, N.; Paul, R.; Özer, T. A Comparative Analysis of Deep Neural Network Models in IoT-based Smart Systems for Energy Prediction and Theft Detection. IET Renew. Power Gener. 2024, 18, 398–411. [Google Scholar] [CrossRef] [Scilit]
  60. Yan, Z.; Wen, H. Comparative Study of Electricity-Theft Detection Based on Gradient Boosting Machine. In 2021 IEEE International Instrumentation and Measurement Technology Conference (I2MTC), Glasgow, UK; IEEE: New York, NY, USA, 2021; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  61. Vinoth Kumar, K.; Aishwarya, P.; Charishma, A.; Gautammee, K.K.; Deepthi, K. A Review of Theft Diagnosis from Smart Energy Meter Using IoT. In 2022 6th International Conference on Electronics, Communication and Aerospace Technology, Coimbatore, India; IEEE: New York, NY, USA, 2022; pp. 526–529. [Google Scholar] [CrossRef] [Scilit]
  62. Takiddin, A.; Ismail, M.; Zainal, A. Robust Electricity Theft Detection Against Data Poisoning Attacks in AMI Systems. IEEE Trans. Smart Grid 2021, 12, 2214–2225. [Google Scholar] [CrossRef] [Scilit]
  63. Takiddin, A.; Ismail, M.; Serpedin, E. Robust Data-Driven Detection of Electricity Theft Adversarial Evasion Attacks in Smart Grids. IEEE Trans. Smart Grid 2023, 14, 663–676. [Google Scholar] [CrossRef] [Scilit]
  64. Fu, C.; Kazmi, H.; Quintana, M.; Miller, C. Creating synthetic energy meter data using conditional diffusion and building metadata. Energy Build. 2024, 312, 114216. [Google Scholar] [CrossRef] [Scilit]
  65. Shi, X.; Dong, L.; Cai, Y.; Tian, C.; Kalathil, D.; Ding, K.; Li, N.; Xie, L. Review of the Opportunities and Challenges to Accelerate Mass-Scale Application of Smart Grids with Large Language Models. IET Smart Grid 2024, 7, 737–759. [Google Scholar] [CrossRef] [Scilit]
  66. Majumder, S.; Dong, L.; Doudi, F.; Cai, Y.; Tian, C.; Kalathil, D.; Ding, K.; Thatte, A.A.; Li, N.; Xie, L. Exploring the Capabilities and Limitations of Large Language Models in the Electric Energy Sector. Joule 2024, 8, 1544–1549. [Google Scholar] [CrossRef] [Scilit]
  67. Madani, K.; Smith, J.; Rahman, A. Large Language Model Integration in Smart Grids: Use Cases and Reliability Considerations. Energy Rep. 2025, 14, 1562–1577. [Google Scholar] [CrossRef] [Scilit]
  68. Liao, W.; Zhu, R.; Ge, L.; Cao, D.; Yang, Z. Mitigating Class Imbalance Issues in Electricity Theft Detection via a Sample-Weighted Loss. IEEE Trans. Ind. Inform. 2025, 21, 1754–1763. [Google Scholar] [CrossRef] [Scilit]
  69. Wang, Y.; Li, X.; Zhou, Q. A Two-Stage Generalizable Electricity Theft Detection Framework for Cross-Regional Deployment. Appl. Energy 2024, 367, 123228. [Google Scholar] [CrossRef] [Scilit]
  70. Liao, W.; Takiddin, A.; Tariq, M.; Chen, S.; Ge, L.; Yang, Z. Sample Adaptive Transfer for Electricity Theft Detection With Distribution Shifts. IEEE Trans. Power Syst. 2024, 39, 7012–7024. [Google Scholar] [CrossRef] [Scilit]
  71. Zheng, Z.; Yang, Y.; Niu, X.; Dai, H.; Zhou, Y. Wide and Deep Convolutional Neural Network for Electricity Theft Detection. IEEE Trans. Ind. Inform. 2018, 14, 1609–1618. [Google Scholar] [CrossRef] [Scilit]
  72. Liao, W.; Zhu, R.; Yang, Z.; Liu, K.; Zhang, B.; Zhu, S.; Feng, B. Electricity Theft Detection Using Dynamic Graph Construction and Graph Attention Networks. IEEE Trans. Ind. Inform. 2024, 20, 5074–5086. [Google Scholar] [CrossRef] [Scilit]
  73. Nandhini, N.; Manikandan, V.; Elango, S. An Interpretable Generalized Additive Neural Networks for Electricity Theft Detection in Smart Cities Using Balanced Data and Intelligent Grid Management. Energy Build. 2025, 346, 116123. [Google Scholar] [CrossRef] [Scilit]
  74. Kawoosa, A.I.; Prashar, D.; Anantha Raman, G.R.; Bijalwan, A.; Haq, M.A.; Aleisa, M.; Alenizi, A. Improving Electricity Theft Detection Using Electricity Information Collection System and Customers’ Consumption Patterns. Energy Explor. Exploit. 2024, 42, 1684–1714. [Google Scholar] [CrossRef] [Scilit]
  75. Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ 2021, 372, n71. [Google Scholar] [CrossRef] [Scilit]
  76. Fragkioudaki, A.; Cruz-Romero, P.; Gómez-Expósito, A.; Biscarri, J.; De Tellechea, M.J.; Arcos, Á. Detection of Non-Technical Losses in Smart Distribution Networks: A Review. In Trends in Practical Applications of Scalable Multi-Agent Systems, the PAAMS Collection; De La Prieta, F., Escalona, M.J., Corchuelo, R., Mathieu, P., Vale, Z., Campbell, A.T., Rossi, S., Adam, E., Jiménez-López, M.D., Navarro, E.M., et al., Eds.; Advances in Intelligent Systems and Computing; Springer International Publishing: Cham, Switzerland, 2016; Volume 473, pp. 43–54. [Google Scholar] [CrossRef] [Scilit]
  77. De Faria, R.A.; Fonseca, K.V.O.; Schneider, B.; Nguang, S.K. Collusion and Fraud Detection on Electronic Energy Meters—A Use Case of Forensics Investigation Procedures. In 2014 IEEE Security and Privacy Workshops. Presented at the 2014 IEEE Security and Privacy Workshops (SPW); IEEE: San Jose, CA, USA, 2014; pp. 65–68. [Google Scholar] [CrossRef] [Scilit]
  78. Wen, H.; Liu, X.; Lei, B.; Yang, M.; Cheng, X.; Chen, Z. A Privacy-Preserving Heterogeneous Federated Learning Framework with Class Imbalance Learning for Electricity Theft Detection. Appl. Energy 2025, 378, 124789. [Google Scholar] [CrossRef] [Scilit]
  79. Available online: https://github.com/henryRDlab/ElectricityTheftDetection (accessed on 2 January 2025).
  80. Nadeem, A.; Arshad, N. PRECON: Pakistan Residential Electricity Consumption Dataset, in: Proceedings of the Tenth ACM International Conference on Future Energy Systems. In e-Energy ’19: The Tenth ACM International Conference on Future Energy Systems, Phoenix, AZ, USA; ACM: New York, NY, USA, 2019; pp. 52–57. [Google Scholar] [CrossRef] [Scilit]
  81. Kelly, J.; Knottenbelt, W. The UK-DALE dataset, domestic appliance-level electricity demand and whole-house demand from five UK homes. Sci. Data 2015, 2, 150007. [Google Scholar] [CrossRef] [Scilit]
  82. Ratnam, E.L.; Weller, S.R.; Kellett, C.M.; Murray, A.T. Residential load and rooftop PV generation: An Australian distribution network dataset. Int. J. Sustain. Energy 2015, 36, 787–806. [Google Scholar] [CrossRef] [Scilit]
  83. Chavat, J.; Nesmachnow, S.; Graneri, J.; Alvez, G. ECD-UY, detailed household electricity consumption dataset of Uruguay. Sci. Data 2022, 9, 21. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Global smart meters forecast [9].
Figure 1. Global smart meters forecast [9].
Bdcc 10 00297 g001
Figure 2. Flow chart of the literature review selection according to the PRISMA guidelines.
Figure 2. Flow chart of the literature review selection according to the PRISMA guidelines.
Bdcc 10 00297 g002
Figure 3. Smart meter attacks.
Figure 3. Smart meter attacks.
Bdcc 10 00297 g003
Figure 4. Real case NTL [7].
Figure 4. Real case NTL [7].
Bdcc 10 00297 g004
Figure 5. Methods for NTL reduction and detection.
Figure 5. Methods for NTL reduction and detection.
Bdcc 10 00297 g005
Table 1. Comparison of representative existing secondary studies with the present tertiary review.
Table 1. Comparison of representative existing secondary studies with the present tertiary review.
StudyYearMain ScopeTaxonomy/Methodological FocusPrimary EmphasisCross-Review SynthesisQuantitative Cross-Survey Analysis
Chauhan and Rajvanshi [41]2013NTL sources, impacts, estimation, and diagnostic techniquesNTL sources, meter tampering, illegal connections, unmetered supply, and load-profile-based detectionGeneral NTLNoNo
Viegas et al. [42]2017Solutions for detection of NTL in electricity gridsTypology of NTL detection solutions, attack/vulnerability points, and hardware/non-hardware approachesMethods, data, and limitationsNoNo
Glauner et al. [14]2017NTL detection using artificial intelligenceExpert systems and machine learning; algorithms, features, datasets, and scientific/engineering challengesAI-based NTL detectionNoNo
Messinis and Hatziargyriou [43]2018NTL detection methodsData-oriented, network-oriented, and hybrid detection methods; algorithms, features, datasets, metrics, and response timeDetection methodology and evaluationNoNo
Ahmad et al. [44]2018Modeling techniques for electricity-theft and NTL detection in smart-grid environmentsData mining, SVM, genetic-SVM, optimum-path forest, decision trees, Bayesian networks, real-time state estimation, and hybrid modelsModeling techniques, smart-meter data, consumption profiles, and NTL detectionNoNo
Saeed et al. [40]2020Detection of NTL in power utilitiesSocial/economic, hardware-based, and non-hardware-based approachesAlgorithms, features, metrics, cost, and response timeNoNo
Hammerschmitt et al. [45]2020NTL characterization and methodological solutionsRegulatory characterization, NTL estimation, fraud, theft, and detection methodologiesRegulatory and methodological aspectsNoNo
Chuwa and Wang [46]2021NTL attack models and detection in smart gridsAMI attack models, feature engineering, learning models, and attack-oriented detectionAMI and attack modelsNoNo
Althobaiti et al. [23]2021Data-driven energy-theft attacks and detection methodsEnergy-theft attacks across smart-grid demand, supply, and control layers; data-driven detection modelsCyber/data attacks and detectionNoNo
Pealy and Matin [47]2021Energy-theft detection and control in smart gridsSmart-meter monitoring, SVM, fuzzy classification, visualization, game theory, and preventive measuresSmart-meter-based theft detection and controlNoNo
Pal [15]2021Identification of non-technical lossesData-oriented, network-oriented, and hybrid methods; supervised and unsupervised learning, network analysis, and intrusion detectionData, network, and hybrid methodsNoNo
Savian et al. [38]2021Worldwide panorama of non-technical losses, impacts, barriers, strategies, and regulationsSystematic review of NTL definitions, identification barriers, mitigation strategies, and regulatory/policy perspectivesNTL characterization, mitigation, and regulationNoNo
Ahmed et al. [37]2022Energy-theft detection in smart gridsThree-level taxonomy covering data mining, state/network, and game-theoretic approachesTaxonomy and comparative analysisNoNo
Xia et al. [8]2022Electricity-theft detection in smart metersMeter technologies, vulnerabilities, cyber/physical attacks, and detection methodsSmart meters and attack mechanismsNoNo
Shokry et al. [12]2022Security of advanced metering infrastructureAMI vulnerabilities, attacks, countermeasures, security perimeters, and future directionsAMI security and attack mitigationNoNo
Yadav and Kumar [48]2022NTL and electricity-theft detection using smart-meter data and AIAI, expert systems, SVM, genetic algorithms, ANN, CNN, RF, SVDD, and related techniquesAI-based detection and evaluationNoNo
Guarda et al. [49]2023Non-hardware-based NTL detectionNetwork-based, data-based, and hybrid non-hardware approachesData-driven and network-based methodsNoNo
Stracqualursi et al. [39]2023Energy-theft practices and autonomous AI-based detectionTheft practices, ML, DL, neural networks, smart meters, and generalized AI detectionAI-based detectionNoNo
Badr et al. [36]2023Data-driven electricity-fraud detection in smart metering systemsSupervised, unsupervised, deep learning, privacy-preserving, and adversarially robust detectionData-driven methods, privacy, and adversarial robustnessNoNo
Pazderin et al. [17]2023Data-driven ML methods for NTL detectionMachine-learning and neural-network approaches for anomaly detectionML/DL and computational methodsNoNo
Kolade et al. [18]2023ML-based energy-theft detectionClassification, anomaly detection, time-series, deep reinforcement, and ensemble approachesMachine learningNoNo
Haruna et al. [50]2024Electricity-theft detection with limited dataState-based/hardware and data-driven methods; supervised, unsupervised, TCN, LSTM, DCNN, MLP, GRU, ANN, and related AI approachesLimited-data detection, computational complexity, data requirements, overfitting, scalability, and generalizabilityNoNo
Kgaphola et al. [34]2024Technology-based electricity-theft detection and preventionConventional, government, and technology-based solutionsTechnology solutions and effectivenessNoNo
Kim et al. [33]2024Data-driven approaches for energy-theft detectionSupervised, unsupervised, deep learning, datasets, privacy, and Generative AIData-driven ETD and Generative AINoNo
Naidji et al. [35]2024AI-based electricity-theft detection in smart gridsMachine learning, deep learning, data mining, data analytics, privacy-preserving and federated-learning approachesAI-based ETD, privacy, robustness, scalability, and real-time processingNoNo
Nayak and Jaidhar [51]2025Electricity theft and anomalous power consumptionML, DL, hybrid, statistical, privacy- preserving, and dataset-oriented approachesTheft and anomalous consumptionNoNo
Iqbal et al. [52]2025Technical case studies for electricity-theft detection in smart gridsSynthetic-data detection, sequential data, non-sequential data, neighborhood area networks, and IoT/hardware solutionsTechnical case studies and performance comparisonNoNo
Morgoev et al. [53]2026Data-driven NTL detection in distribution gridsAnalytical paradigm, input-data structure, and grid digitalizationUtility-centric method selectionNoNo
Present study2026Tertiary synthesis of the secondary literature on NTL/ETDCommon framework covering attack types, detection approaches, data/features, AMI, evaluation, challenges, and research gapsCross-survey evidence synthesisYesYes
Table 2. Tertiary evidence matrix and comparative characteristics of the secondary studies included in the systematic tertiary review.
Table 2. Tertiary evidence matrix and comparative characteristics of the secondary studies included in the systematic tertiary review.
Secondary StudyYear#Cit.Main ContributionRQ1: Attack TypesRQ2: Detection ApproachesRQ3: Data, Features & AMIRQ4: EvaluationRQ5:
Challenges/Gaps
Detection Methods in Smart Meters for Electricity Thefts: A Survey [8]202269Examines the transition from conventional to smart meters, electricity-theft motivations, meter technologies and associated vulnerabilities.Meter manipulation; smart-meter attacksMachine learning; inspection-based methodsSmart-meter data and featuresPerformance metrics discussedCybersecurity; smart-meter vulnerabilities
Review of Non-Technical Loss Detection Methods [43]2018179Provides a classification of NTL detection methods into data-oriented, network-oriented and hybrid approaches.NTL; electricity theftData-oriented; network-oriented; hybridData types and featuresClassification metricsClass imbalance
Non-Hardware-Based Non-Technical Losses Detection Methods: A Review [49]20235Reviews non-hardware NTL detection methods and compares data-oriented, network-oriented and hybrid approaches.NTLData-oriented; network-oriented; hybridLimited emphasisComparative analysisLimited coverage of practical evaluation
Data-Driven Approaches for Energy Theft Detection: A Comprehensive Review [33]20243Reviews supervised and semi-supervised data-driven methods and discusses generative AI as an emerging direction.Electricity theftSupervised; semi-supervised; generative AIHigh-dimensional and limited-label dataComparative discussionHigh dimensionality; lack of labels
A Review of Non-Technical Loss Attack Models and Detection Methods in the Smart Grid [46]202133Examines attack models based on consumption patterns and constructs malicious load profiles from real-world data.Consumption-pattern attacksMultiple NTL detection techniquesReal-world consumption data; featuresRobustness comparisonRobustness against attack profiles
Detection of Non-Technical Losses in Power Utilities—A Comprehensive Systematic Review [40]202042Classifies NTL detection into social/economic, hardware-based and non-hardware-based approaches.NTL; electricity theftData-based; network-based; hybrid; hardwareData and network informationPerformance; cost; response timeDeployment cost; practical applicability
Electricity Theft Detection and Prevention Using Technology-Based Models: A Systematic Literature Review [34]20242Classifies detection methods into conventional, government and technology-based approaches and applies quality-assessment criteria.Electricity theftConventional; government; technology-basedLimited emphasisMetrics; classification results; quality assessmentMethodological quality
Electricity Theft Detection Techniques Using Artificial Intelligence: A Survey [35]20240Examines AI-based data-driven electricity-theft detection with emphasis on privacy-preservation techniques.Electricity theftArtificial intelligence; data-driven methodsSmart-meter/data-driven contextLimited emphasisPrivacy preservation
Review of the Data-Driven Methods for Electricity Fraud Detection in Smart Metering Systems [36]202332Reviews data-driven electricity-fraud detection with emphasis on privacy and adversarial attacks.Electricity fraud; adversarial attacksData-driven approachesSmart-meter dataLimited emphasisPrivacy; adversarial attacks; defensive mechanisms
Systematic Review of Energy Theft Practices and Autonomous Detection through Artificial Intelligence Methods [39]202318Examines illegal tapping, meter tampering and physical attack practices in low- and medium-voltage networks.Illegal tapping; meter tampering; magnetic attacksAI-based autonomous detectionElectrical measurements; attack characteristicsLimited emphasisPhysical attacks; practical detection
Solutions for Detection of Non-Technical Losses in the Electricity Grid: A Review [42]2017143Categorizes NTL solutions into social/economic, hardware-based and non-hardware-based approaches.NTL; electricity theftHardware; non-hardware; socio-economicLimited emphasisLimited emphasisHardware/non-hardware trade-offs
Energy Theft Detection in Smart Grids: Taxonomy, Comparative Analysis, Challenges, and Future Research Directions [37]202227Provides a taxonomy based on data mining, state/network and game-theoretic approaches and compares detection methods.Energy theft; NTLData mining; state/network; game theory; classification; clusteringSmart-grid contextMetrics; comparative analysisChallenges; future research
A Critical Review of Technical Case Studies for Electricity Theft Detection in Smart Grids: A New Paradigm Based Transformative Approach [52]20258Converts technical electricity-theft detection literature into case-study-oriented evidence and organizes approaches into synthetic-data, sequential-data, non-sequential-data, NAN, and IoT/hardware categories.Theft cases; false-data injection; meter/data manipulationSynthetic-data; sequential; non-sequential; NAN; IoT/hardwareSmart-meter data; sequential and non-sequential data; AMI/NANF1-score and multiple evaluation metricsData integrity; privacy; false positives; deployment and hardware constraints
Data-Driven Models for Electricity Theft and Anomalous Power Consumption Detection: A Systematic Review [51]2025Systematically reviews electricity-theft and anomalous power-consumption detection and classifies studies into ML, DL and hybrid models.Electricity theft; anomalous consumption; meter manipulation; feeder bypassMachine learning; deep learning; hybrid modelsDatasets; smart-grid consumption data; privacy-preserving dataDetection performance and reported metricsPrivacy; data availability; scalability; generalization
Review on Temporal Convolutional Networks for Electricity Theft Detection with Limited Data [50]2024Reviews AI/ML approaches for electricity-theft detection under limited-data conditions, emphasizing computational complexity, overfitting and generalizability.Electricity theft; meter tampering; meter bypassing; false readingsTCN; LSTM; DCNN; MLP; GRU; ANNLimited electricity-consumption dataPerformance discussed across reviewed modelsLimited data; computational complexity; overfitting; scalability; generalizability
Data-Driven Machine Learning Methods for Nontechnical Losses of Electrical Energy Detection: A State-of-the-Art Review [17]202317Provides a state-of-the-art review of computational methods for locating and identifying NTL sources, with emphasis on neural-network-based methods.NTL; electricity theft; abnormal consumptionMachine learning; neural networks; CNN; autoencodersInitial data sources; data composition; consumption dataTraining/testing metrics and effectiveness criteriaData characteristics; algorithm selection; method effectiveness
Energy Theft Detection in Power System Network: Reviews of Studies on Machine Learning Based Solutions [18]20235Reviews ML-based electricity-theft detection and classifies methods into classification, anomaly detection, time-series, deep reinforcement and ensemble approaches.Energy theft; NTLSupervised; unsupervised; reinforcement; ensemble; deep learningEnergy-consumption data; smart-grid contextAccuracy and other reported metricsData quality; computational resources; privacy; security; algorithm tuning
Literature Review of Methods for Detecting Non-Technical Electricity Losses in Distribution Grids [53]2026Systematically reviews data-driven NTL detection studies and proposes a utility-centric classification based on analytical paradigm, input-data structure and grid digitalization.NTL; electricity theftSupervised classification; unsupervised clustering; forecasting/regression; scenario modelingAMI; high-resolution consumption data; input-data structure; grid digitalizationF1-score and comparative performance analysisData imbalance; interpretability; real-world deployment; digitalization constraints
Energy Theft in Smart Grids: A Survey on Data-Driven Attack Strategies and Detection Methods [23]202139Survey of data-driven energy-theft strategies and detection methods across smart-grid operational layers.Energy theft; fraud; data-driven attacksML; DL; anomaly detection; data-driven detection modelsAMI; demand, supply, and control-chain dataCategorization and comparative assessment of detection modelsCybersecurity; data quality; distributed-grid complexity; open research issues
Non-Technical Losses in Power System: A Review [41]201367Review of NTL sources, impacts, estimation techniques, and diagnostic approaches used by utilities.Meter tampering; illegal connections; false readings; unmetered supplyClassification; load profiling; diagnostic and detection techniquesLoad profiles; distribution-system informationQualitative review of NTL estimation and diagnostic techniquesHigh operational cost; difficulty of NTL estimation; manual inspection dependence
The Challenge of Non-Technical Loss Detection Using Artificial Intelligence: A Survey [14]2017306AI-oriented survey covering NTL definitions, algorithms, features, datasets, and scientific/engineering challenges.Electricity theft; meter tampering; bypassing; faulty meters; billing errorsExpert systems; ML; statistical methods; AI-based detectionMonthly consumption; load profiles; smart-meter/customer dataComparison of algorithms, features, and datasetsCovariate shift; limited labeled data; generalization; deployment challenges
Tackling Energy Theft in Smart Grid-A Comprehensive Review and Framework [47]2021Comprehensive review and framework for energy-theft detection and control using smart-meter technology.Meter tampering; meter bypass; billing anomalies; unpaid billsSVM; fuzzy classification; visualization; AI-based approachesSmart-meter consumption data; smart-grid infrastructureComparative discussion of existing detection techniquesPrivacy; cybersecurity; manual inspection; implementation challenges
Non-Technical Losses Review and Possible Methodology Solutions [45]202022Review of NTL characterization, regulatory aspects, and methodological solutions for NTL reduction.Fraud; theft; meter adulteration; clandestine connectionsNTL estimation and detection methodologiesUtility and regulatory data; distribution-system lossesDiscussion of Brazilian NTL statistics and methodological approachesRegulatory limitations; persistent NTL; effectiveness of existing measures
Review of various modeling techniques for the detection of electricity theft in smart grid environment [44]2018128Compares major modeling strategies and identifies their strengths, limitations, and applicability for NTL detection.Electricity theft; NTL; irregular consumptionSVM; ANN; OPF; clustering; state estimation; hybrid models; decision trees; Bayesian methodsSmart-meter data; customer consumption profiles; AMIComparative discussion of detection/modeling performanceData availability; model selection; practical deployment limitations
The Detection of Non-Technical Losses and Electricity Theft by Smart Meter Data and Artificial Intelligence in the Context of Electric Distribution Utilities: A Comprehensive Review [48]20228Comprehensive review of AI-based NTL and electricity-theft detection using smart-meter data.Electricity theft; NTL; meter-related anomaliesSVM; GA-SVM; expert systems; CNN; ANN; RF; image-based learning; SVDDSmart-meter data; consumption profiles; AI-based datasetsComparison of AI techniques, tools, and environmentsData quality; implementation complexity; model limitations; future research needs
Review of Non-Technical Losses Identification Techniques [15]2021Review of NTL identification techniques with qualitative comparison based on performance, cost, data handling, quality control, and execution time.Meter tampering; illegal connections; billing irregularities; faulty metersStatistical methods; decision trees; ANN; SVM; graph-based methods; clusteringCustomer databases; consumption patterns; smart-meter dataQualitative comparison of accuracy, cost, data handling, and execution timeHigh inspection cost; data complexity; scalability and practical limitations
Non-technical losses: A systematic contemporary article review [38]202186Systematic review providing a global overview of NTL, impacts, barriers, mitigation strategies, policies, and regulations.Electricity theft; illegal connections; meter errors; billing errorsDetection and mitigation strategies; technological and organizational approachesDistribution-system data; consumption information; smart-meter dataPRISMA-based synthesis of 121 journal articlesRegulatory barriers; technological limitations; implementation and policy gaps
Systematic survey of advanced metering infrastructure security: Vulnerabilities, attacks, countermeasures, and future vision [12]202268Provides a systematic security perspective on AMI by mapping vulnerabilities, attacks, countermeasures, and open research challenges across the AMI architecture.Meter tampering; fraudulent data manipulation; data theft; impersonation; DoS/DDoS; MITMSecurity monitoring; intrusion detection; authentication; encryption; countermeasure-based approachesAMI hardware, data and communication layers; smart meters; data concentrators; utility center; consumption dataComparative analysis of attacks, vulnerabilities, impacts, and countermeasuresAMI security; privacy; data integrity; communication vulnerabilities; resource constraints; deployment challenges
Table 3. Quantitative coverage of research dimensions across the 28 secondary studies.
Table 3. Quantitative coverage of research dimensions across the 28 secondary studies.
RQResearch DimensionStudiesCoverage (%)
RQ1Meter tampering/manipulation1139.3
Unauthorized connections/meter bypass725.0
Cyber/data attacks310.7
Billing/fault/anomalous-consumption issues621.4
RQ2ML/AI-based approaches1657.1
Deep learning architectures725.0
Statistical/conventional methods725.0
Hybrid/ensemble approaches621.4
Unsupervised/anomaly detection/clustering517.9
Security/intrusion-detection approaches13.6
Generative AI/generative models414.3
RQ3Smart-meter/AMI data1139.3
Consumption/load-profile data932.1
Explicit feature representation27.1
Real-world data explicitly identified13.6
Limited/synthetic/constrained data517.9
RQ4Performance metrics/effectiveness1242.9
Comparative evaluation932.1
Cost/response/execution-time evaluation27.1
RQ5Data availability/quality/complexity828.6
Privacy/cybersecurity932.1
Class imbalance27.1
Scalability/generalizability414.3
Deployment/practical implementation1035.7
Computational resources/complexity621.4
Interpretability13.6
Regulatory/policy issues27.1
Table 4. Representative datasets reported in the reviewed literature.
Table 4. Representative datasets reported in the reviewed literature.
DatasetCustomersSample RateDurationCountryAppliances
SGCC [33,79]42,3721 dayJanuary 2014–October 2016ChinaNo
PRECON [34,80]421 minJune 2018–May 2019PakistanYes
UK-DALE [39,81]56 sMax. 2014–2017UKYes
CER [36,46]500030 minJanuary 2009–December 2010IrelandNo
Ausgrid [36,82]30030 minJuly 2010.7–June 2013AustraliaNo
ECD-UY [33,83]110,9531–15 minJanuary 2019–November 2020UruguayYes
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Labate, D.; Thakur, D.; Guzzo, A.; Fortino, G. A Comprehensive Systematic Meta-Survey of Energy Theft Detection: From Traditional Methods to Generative AI. Big Data Cogn. Comput. 2026, 10, 297. https://doi.org/10.3390/bdcc10090297

AMA Style

Labate D, Thakur D, Guzzo A, Fortino G. A Comprehensive Systematic Meta-Survey of Energy Theft Detection: From Traditional Methods to Generative AI. Big Data and Cognitive Computing. 2026; 10(9):297. https://doi.org/10.3390/bdcc10090297

Chicago/Turabian Style

Labate, Diego, Dipanwita Thakur, Antonella Guzzo, and Giancarlo Fortino. 2026. "A Comprehensive Systematic Meta-Survey of Energy Theft Detection: From Traditional Methods to Generative AI" Big Data and Cognitive Computing 10, no. 9: 297. https://doi.org/10.3390/bdcc10090297

APA Style

Labate, D., Thakur, D., Guzzo, A., & Fortino, G. (2026). A Comprehensive Systematic Meta-Survey of Energy Theft Detection: From Traditional Methods to Generative AI. Big Data and Cognitive Computing, 10(9), 297. https://doi.org/10.3390/bdcc10090297

Article Metrics

Back to TopTop