Figure 1.
The geometry of multiplicative under-registration vs. lawful smart-city confounders. (a) A real LCL household (MAC000009) with a synthetic multiplicative attack () injected at : the post-event level drops but the daily structure is preserved. (b) Daily-profile view of the same attack: the slotwise ratio is approximately flat at . (c) Lawful PV onset (post-event profile built from the same recipient minus a real OPSD PV donor feed): the ratio drops at midday when PV cancels demand and is ∼1 at night. (d) Lawful vacation: the post-event profile collapses to zero, so the slotwise ratio is degenerate. Our detector targets geometry (b) and rejects (c,d) via spectral and post-activity gates. In all panels, blue marks the pre-event daily profile; the post-event profile is shown in red for the attack (panel b), green for the PV onset (panel c), and purple for the vacation (panel d); and the dotted gray curve is the slotwise ratio on the right-hand axis.
Figure 1.
The geometry of multiplicative under-registration vs. lawful smart-city confounders. (a) A real LCL household (MAC000009) with a synthetic multiplicative attack () injected at : the post-event level drops but the daily structure is preserved. (b) Daily-profile view of the same attack: the slotwise ratio is approximately flat at . (c) Lawful PV onset (post-event profile built from the same recipient minus a real OPSD PV donor feed): the ratio drops at midday when PV cancels demand and is ∼1 at night. (d) Lawful vacation: the post-event profile collapses to zero, so the slotwise ratio is degenerate. Our detector targets geometry (b) and rejects (c,d) via spectral and post-activity gates. In all panels, blue marks the pre-event daily profile; the post-event profile is shown in red for the attack (panel b), green for the PV onset (panel c), and purple for the vacation (panel d); and the dotted gray curve is the slotwise ratio on the right-hand axis.
Figure 2.
versus scaling factor (smaller = stronger attack) for the preferred profile_gls detector (blue) and the relaxed robust variant profile_rgls (green) on the two QC-filtered internal industrial subsets, at attack prevalence .
Figure 2.
versus scaling factor (smaller = stronger attack) for the preferred profile_gls detector (blue) and the relaxed robust variant profile_rgls (green) on the two QC-filtered internal industrial subsets, at attack prevalence .
Figure 3.
Public Low Carbon London sample512 device-split results. (Left panel): Precision, recall, and for the three detectors at the recommended operating point (, ). (Right panel): How the preferred detector’s score degrades as the scaling factor grows (i.e., as the attack weakens) at the two tested prevalence levels.
Figure 3.
Public Low Carbon London sample512 device-split results. (Left panel): Precision, recall, and for the three detectors at the recommended operating point (, ). (Right panel): How the preferred detector’s score degrades as the scaling factor grows (i.e., as the attack weakens) at the two tested prevalence levels.
Figure 4.
Public lawful-change specificity benchmark on the LCL sample512 device split at event prevalence . The preferred GLS detector (blue, almost invisible at this scale because the values are below on every event group) remains selective on non-event baseline devices, profile-change events, and vacation events, whereas the robust variant (green) flags roughly one in four non-theft devices across every event category.
Figure 4.
Public lawful-change specificity benchmark on the LCL sample512 device split at event prevalence . The preferred GLS detector (blue, almost invisible at this scale because the values are below on every event group) remains selective on non-event baseline devices, profile-change events, and vacation events, whereas the robust variant (green) flags roughly one in four non-theft devices across every event category.
Figure 5.
Reactive-gate effect on the two public reactive-capable benchmarks at , . (Left panel): WPuQ; (right panel): Mendeley.
Figure 5.
Reactive-gate effect on the two public reactive-capable benchmarks at , . (Left panel): WPuQ; (right panel): Mendeley.
Figure 6.
External-baseline comparison. (a) Pooled attack average precision on Low Carbon London with 95% bootstrap confidence intervals; the dotted line marks the theft prevalence. (b) The selectivity frontier: in-distribution ranking power (attack AP, vertical) versus the number of false alarms on the real OPSD PV/EV/heat-pump overlays under frozen cross-dataset transfer (horizontal, out of 72 cases). The proposed profile_gls detector is the only method in the high-precision, zero-transfer-false-alarm region; the supervised gradient-boosting model attains higher in-distribution AP but false alarms on many real overlays, and the generic detectors fail on both axes. In both panels, the proposed physics-guided detectors are drawn in color (the preferred GLS detector in blue and its two profile-based variants in green), while the external baselines are in gray; in panel (b) the proposed detectors are filled circles and the baselines are gray crosses, and the gray arrow indicates the direction of better performance (higher average precision with fewer transfer false alarms).
Figure 6.
External-baseline comparison. (a) Pooled attack average precision on Low Carbon London with 95% bootstrap confidence intervals; the dotted line marks the theft prevalence. (b) The selectivity frontier: in-distribution ranking power (attack AP, vertical) versus the number of false alarms on the real OPSD PV/EV/heat-pump overlays under frozen cross-dataset transfer (horizontal, out of 72 cases). The proposed profile_gls detector is the only method in the high-precision, zero-transfer-false-alarm region; the supervised gradient-boosting model attains higher in-distribution AP but false alarms on many real overlays, and the generic detectors fail on both axes. In both panels, the proposed physics-guided detectors are drawn in color (the preferred GLS detector in blue and its two profile-based variants in green), while the external baselines are in gray; in panel (b) the proposed detectors are filled circles and the baselines are gray crosses, and the gray arrow indicates the direction of better performance (higher average precision with fewer transfer false alarms).
![Smartcities 09 00110 g006 Smartcities 09 00110 g006]()
Figure 7.
Pooled score-sweep PR curves reconstructed directly from per-device detector scores, aggregated across all settings in each benchmark family. (Left): LCL sample512; (right): Mendeley sample512.
Figure 7.
Pooled score-sweep PR curves reconstructed directly from per-device detector scores, aggregated across all settings in each benchmark family. (Left): LCL sample512; (right): Mendeley sample512.
Figure 8.
versus smooth slotwise distortion strength on the LCL sample512 evaluation split at , with separate curves for (solid) and (dashed).
Figure 8.
versus smooth slotwise distortion strength on the LCL sample512 evaluation split at , with separate curves for (solid) and (dashed).
Figure 9.
Recall on the injected meters versus within-day distortion strength for the three distortion families on the LCL sample512 split at : smooth (low-frequency), non-smooth (independent per-slot), and time-selective (evening-peak suppression).
Figure 9.
Recall on the injected meters versus within-day distortion strength for the three distortion families on the LCL sample512 split at : smooth (low-frequency), non-smooth (independent per-slot), and time-selective (evening-peak suppression).
Table 1.
QC-filtered industrial subsets group10 (stable) and group1 (noisy) at attack prevalence . Best per metric within each (subset, ) row pair is shown in bold. At the RGLS rows would coincide with the values and are omitted.
Table 1.
QC-filtered industrial subsets group10 (stable) and group1 (noisy) at attack prevalence . Best per metric within each (subset, ) row pair is shown in bold. At the RGLS rows would coincide with the values and are omitted.
| Method | | Precision | Recall | | FP |
|---|
| Subset
group10
(stable) |
| profile_gls | 0.10 | 1.000 | 1.000 | 1.000 | 0 |
| profile_gls | 0.20 | 1.000 | 1.000 | 1.000 | 0 |
| profile_gls | 0.30 | 0.500 | 0.167 | 0.250 | 0 |
| profile_gls | 0.40 | 0.000 | 0.000 | 0.000 | 0 |
| profile_rgls | 0.10 | 0.354 | 1.000 | 0.523 | 11 |
| profile_rgls | 0.20 | 0.354 | 1.000 | 0.523 | 11 |
| Subset
group1
(noisy) |
| profile_gls | 0.10 | 1.000 | 0.500 | 0.667 | 0 |
| profile_gls | 0.20 | 1.000 | 0.500 | 0.667 | 0 |
| profile_gls | 0.30 | 0.500 | 0.250 | 0.333 | 0 |
| profile_gls | 0.40 | 0.000 | 0.000 | 0.000 | 0 |
| profile_rgls | 0.10 | 0.037 | 0.500 | 0.069 | 52 |
| profile_rgls | 0.20 | 0.037 | 0.500 | 0.069 | 52 |
Table 2.
Public device-split benchmark on the LCL sample512 evaluation pool (243 households). Thresholds are calibrated once on the disjoint 244-household calibration split and frozen for this table. Best per metric within each row group is shown in bold.
Table 2.
Public device-split benchmark on the LCL sample512 evaluation pool (243 households). Thresholds are calibrated once on the disjoint 244-household calibration split and frozen for this table. Best per metric within each row group is shown in bold.
| Method | | Precision | Recall | | FP |
|---|
| Theft prevalence |
| profile | 0.10 | 0.733 | 1.000 | 0.846 | 8 |
| profile_gls | 0.10 | 0.846 | 1.000 | 0.917 | 4 |
| profile_rgls | 0.10 | 0.144 | 1.000 | 0.251 | 131 |
| profile | 0.20 | 0.733 | 1.000 | 0.846 | 8 |
| profile_gls | 0.20 | 0.840 | 0.955 | 0.893 | 4 |
| profile_rgls | 0.20 | 0.144 | 1.000 | 0.251 | 131 |
| Theft prevalence |
| profile | 0.10 | 0.846 | 1.000 | 0.917 | 8 |
| profile_gls | 0.10 | 0.915 | 0.978 | 0.945 | 4 |
| profile_rgls | 0.10 | 0.263 | 1.000 | 0.417 | 123 |
| profile | 0.20 | 0.837 | 0.935 | 0.882 | 8 |
| profile_gls | 0.20 | 0.913 | 0.957 | 0.934 | 4 |
| profile_rgls | 0.20 | 0.263 | 1.000 | 0.417 | 123 |
| Failure regime (, GLS only) |
| profile_gls | 0.30 | 0.657 | 0.184 | 0.286 | 4 |
| profile_gls | 0.40 | 0.167 | 0.024 | 0.042 | 4 |
Table 3.
Results on the Mendeley active/reactive sample512 split (255-household evaluation pool). Best per metric within each (variant, ) row group is shown in bold.
Table 3.
Results on the Mendeley active/reactive sample512 split (255-household evaluation pool). Best per metric within each (variant, ) row group is shown in bold.
| Method | | Precision | Recall | | FP |
|---|
| Variant
main
(no reactive gate) |
| profile | 0.10 | 0.727 | 0.941 | 0.821 | 6 |
| profile_gls | 0.10 | 0.696 | 0.941 | 0.800 | 7 |
| profile_rgls | 0.10 | 0.145 | 0.941 | 0.252 | 94 |
| profile | 0.20 | 0.727 | 0.941 | 0.821 | 6 |
| profile_gls | 0.20 | 0.708 | 1.000 | 0.829 | 7 |
| profile_rgls | 0.20 | 0.153 | 1.000 | 0.266 | 94 |
| Variant
reactive_gate
(with reactive-channel agreement) |
| profile | 0.10 | 0.867 | 0.765 | 0.812 | 2 |
| profile_gls | 0.10 | 0.696 | 0.941 | 0.800 | 7 |
| profile_rgls | 0.10 | 0.145 | 0.941 | 0.252 | 94 |
| profile | 0.20 | 0.857 | 0.706 | 0.774 | 2 |
| profile_gls | 0.20 | 0.708 | 1.000 | 0.829 | 7 |
| profile_rgls | 0.20 | 0.153 | 1.000 | 0.266 | 94 |
Table 4.
Lawful-change specificity on the LCL sample512 evaluation pool with zero theft injection. Lower is better; best per metric within each prevalence setting is shown in bold (ties bolded jointly).
Table 4.
Lawful-change specificity on the LCL sample512 evaluation pool with zero theft injection. Lower is better; best per metric within each prevalence setting is shown in bold (ties bolded jointly).
| Method | Non-Theft Suspected Rate | Lawful-Event Suspected Rate |
|---|
| Event prevalence |
| profile_gls | 0.8% (4/486) | 0.0% (0/48) |
| profile | 1.9% (9/486) | 2.1% (1/48) |
| profile_rgls | 27.4% (133/486) | 18.8% (9/48) |
| Event prevalence |
| profile_gls | 1.0% (5/486) | 1.0% (1/96) |
| profile | 1.6% (8/486) | 1.0% (1/96) |
| profile_rgls | 27.4% (133/486) | 24.0% (23/96) |
Table 5.
OPSD cross-dataset specificity benchmark with real PV, EV, and heat-pump overlays and frozen LCL thresholds. Lower is better; best per metric within each prevalence setting is shown in bold.
Table 5.
OPSD cross-dataset specificity benchmark with real PV, EV, and heat-pump overlays and frozen LCL thresholds. Lower is better; best per metric within each prevalence setting is shown in bold.
| Method | Non-Theft Suspected Rate | Lawful Suspected Rate |
|---|
| Event prevalence |
| profile_gls | 0.0% (0/36) | 0.0% (0/18) |
| profile | 11.1% (4/36) | 22.2% (4/18) |
| profile_rgls | 11.1% (4/36) | 22.2% (4/18) |
| Event prevalence |
| profile_gls | 0.0% (0/36) | 0.0% (0/24) |
| profile | 11.1% (4/36) | 16.7% (4/24) |
| profile_rgls | 11.1% (4/36) | 16.7% (4/24) |
Table 6.
Three-year WPuQ household reactive-channel benchmark, 17 evaluation households. The RGLS rows at are omitted because they coincide with the values. Best per metric within each (variant, ) row group is shown in bold; ties are bolded jointly.
Table 6.
Three-year WPuQ household reactive-channel benchmark, 17 evaluation households. The RGLS rows at are omitted because they coincide with the values. Best per metric within each (variant, ) row group is shown in bold; ties are bolded jointly.
| Method | | Precision | Recall | | FP |
|---|
| Variant
main
(no reactive gate) |
| profile | 0.10 | 0.500 | 1.000 | 0.667 | 1 |
| profile | 0.20 | 0.500 | 1.000 | 0.667 | 1 |
| profile_gls | 0.10 | 1.000 | 1.000 | 1.000 | 0 |
| profile_gls | 0.20 | 1.000 | 1.000 | 1.000 | 0 |
| profile_rgls | 0.10 | 0.250 | 1.000 | 0.400 | 3 |
| Variant
reactive_gate
(with reactive-channel agreement) |
| profile | 0.10 | 1.000 | 1.000 | 1.000 | 0 |
| profile | 0.20 | 1.000 | 1.000 | 1.000 | 0 |
| profile_gls | 0.10 | 1.000 | 1.000 | 1.000 | 0 |
| profile_gls | 0.20 | 1.000 | 1.000 | 1.000 | 0 |
| profile_rgls | 0.10 | 0.250 | 1.000 | 0.400 | 3 |
Table 7.
External baseline comparison on Low Carbon London (243-household evaluation pool) and frozen cross-dataset transfer to the real OPSD PV/EV/heat-pump overlays. AP is the threshold-free pooled average precision (prevalence ); recall is reported at a matched false-positive budget on held-out clean meters; “LCL law.” is the false-alarm rate on synthetic lawful-change events at that operating point; “OPSD real” is the number of false alarms among the 72 real-overlay cases when each method is frozen on LCL and transferred unchanged. Best value per column in bold. The proposed profile_gls is the only method that combines competitive attack recall, low lawful false alarms, and zero false alarms under frozen transfer to real confounders.
Table 7.
External baseline comparison on Low Carbon London (243-household evaluation pool) and frozen cross-dataset transfer to the real OPSD PV/EV/heat-pump overlays. AP is the threshold-free pooled average precision (prevalence ); recall is reported at a matched false-positive budget on held-out clean meters; “LCL law.” is the false-alarm rate on synthetic lawful-change events at that operating point; “OPSD real” is the number of false alarms among the 72 real-overlay cases when each method is frozen on LCL and transferred unchanged. Best value per column in bold. The proposed profile_gls is the only method that combines competitive attack recall, low lawful false alarms, and zero false alarms under frozen transfer to real confounders.
| Method | Family | AP | Recall @ Matched 1% FP | False Alarms |
|---|
| (Pooled)
| | | |
LCL Law.
|
OPSD Real
|
|---|
| Proposed detector family |
| profile_gls | physics, label-free | 0.797 | 1.00 | 0.95 | 0.32 | 2.1% | 0/72 |
| profile | physics, label-free | 0.779 | 1.00 | 0.77 | 0.14 | 0.7% | 8/72 |
| profile_rgls | physics, label-free | 0.788 | 1.00 | 0.93 | 0.18 | 2.1% | 8/72 |
| External baselines |
| change-point (PELT) | change-point | 0.665 | 0.93 | 0.91 | 0.39 | 1.4% | 8/72 |
| Isolation Forest | unsup. outlier | 0.275 | 0.02 | 0.00 | 0.00 | 13.2% | 56/72 |
| autoencoder | deep repr. | 0.043 | 0.00 | 0.00 | 0.00 | 0.7% | 58/72 |
| LSTM | supervised deep | 0.526 | 0.39 | 0.30 | 0.18 | 0.7% | 0/72 |
| gradient boosting | sup. (phys. feat.) | 0.962 | 0.98 | 0.98 | 0.93 | 0.7% | 19/72 |
| Buzau-style GBM | sup. (cust. feat.) | 0.741 | 0.77 | 0.68 | 0.57 | 2.1% | 0/72 |
Table 8.
Calibrated decision thresholds for the Low Carbon London benchmark, frozen and transferred unchanged to OPSD. Quantiles are taken over the per-device statistics of the disjoint calibration split under attack injection (scale, dispersion, spectral) or on uninjected calibration meters, following the protocol of
Section 4.5.
Table 8.
Calibrated decision thresholds for the Low Carbon London benchmark, frozen and transferred unchanged to OPSD. Quantiles are taken over the per-device statistics of the disjoint calibration split under attack injection (scale, dispersion, spectral) or on uninjected calibration meters, following the protocol of
Section 4.5.
| Gate | Calibrated Value (LCL) | Quantile Basis |
|---|
| scale gate () | | of on uninjected (null) calibration meters |
| shape correlation (normal/strong) | / | calibration policy |
| residual dispersion (normal/strong) | / | of calib. |
| spectral JSD | | of calib. JSD |
| spectral log-ratio (max) | | of calib. |
| strong-regime cutoff | | fixed policy |