3.1. Physics-Informed Loss Model and Feature Construction
Let
denote the number of active phases,
the root-mean-square phase voltage,
the per-phase charging-current set-point, and
the displacement power factor. The grid-side AC active power drawn by the OBC is expressed as
Following the definition adopted in the original experimental campaign, conversion efficiency is the ratio of the DC power delivered to the battery,
, to the AC active power drawn at the measurement point,
. Equivalently, efficiency may be expressed in terms of the total converter loss,
, as
The measured efficiency values were converted from percentages to per-unit quantities before model fitting. The resulting root mean square error values were subsequently converted back to percentage points to preserve direct engineering interpretation.
The total OBC loss was represented using the classical quadratic loss model
where
represents approximately fixed losses associated with control electronics, gate drives, magnetic components, auxiliary systems, and cooling overheads. The coefficient
represents losses that scale approximately in proportion to converter throughput, including semiconductor conduction and switching contributions, whereas
represents predominantly ohmic losses associated with windings, filters, connectors, and cabling. Substituting Equation (3) into Equation (2), and applying Equation (1) under approximately constant phase voltage and displacement factor, gives an efficiency model expressed directly in terms of the controllable per-phase current set-point:
The model coefficients are related to the underlying loss terms through
The coefficient
has units of amperes,
is dimensionless, and
has units of inverse amperes. Accordingly,
captures the increasing proportional influence of fixed losses at low charging current,
represents the current-proportional loss floor, and
represents the increase in ohmic losses at higher current.
Figure 2 illustrates the physical and computational structure of the proposed loss model. The measured AC operating variables determine the input active power; the total converter loss is decomposed into fixed, proportional, and ohmic components, and constrained nonlinear parameter estimation produces vehicle-specific loss coefficients and derived efficiency-optimal operating quantities.
Equation (4) was fitted independently to the valid three-phase efficiency measurements of each vehicle using trust-region-reflective nonlinear least squares implemented through scipy.optimize.curve_fit [
21]. Non-negativity constraints were imposed on all three coefficients to preserve their physical interpretation. For a vehicle with
valid operating points, the fitted parameters were obtained from
The optimisation was initialised using
Only experimentally valid current set-points were included in the objective function. Vehicles with fewer than four valid efficiency observations were excluded from parameter fitting because the available information was considered insufficient for stable estimation of three constrained coefficients. Converged boundary estimates (e.g., b = 0) were treated as unresolved components, not confirmed zero losses.
Differentiating Equation (4) with respect to current gives
Setting the derivative equal to zero produces the interior efficiency-optimal current:
At
, the fixed-loss and ohmic-loss contributions are equal. Substitution into Equation (4) yields the corresponding model-predicted peak efficiency:
The coefficient therefore governs the depth of the low-current efficiency penalty, whereas and the geometric mean jointly determine the attainable efficiency ceiling. Where exceeded the largest experimentally supported current, the vehicle was interpreted as exhibiting a monotonically increasing efficiency curve over the measured range, and the observed optimum occurred at the maximum available current set-point.
Goodness of fit was evaluated using the coefficient of determination,
and the root mean square error,
For nearly flat efficiency curves, the denominator of Equation (9) becomes small, meaning that modest absolute residuals can produce comparatively low values. RMSE was therefore treated as the primary measure of absolute fit quality, while was retained as a complementary measure of explained variation.
A total of 24 analytical features were extracted for each test unit. These features were organised into three principal blocks: efficiency curve descriptors, physics-informed loss coefficients, and grid-side power quality descriptors. The
and RMSE values were retained as model-fit diagnostics but were not counted among the 24 characterisation features.
Figure 3 summarises the complete feature-engineering architecture.
Eleven features were extracted from each valid three-phase efficiency curve: peak measured efficiency,
; the corresponding current set-point,
; efficiencies at 6 A, 10 A, and 16 A; mean efficiency over all valid set-points; the low-current part-load penalty; the low-current efficiency slope; the minimum current required to reach 90% efficiency,
; the minimum operable current,
; and the binary
fails-at-6 A indicator. The low-current part-load penalty was defined as
For vehicles with valid observations at both 6 A and 10 A, the low-current slope was calculated as
and was reported in percentage points per ampere. The threshold feature
was defined as the smallest measured current set-point at which conversion efficiency reached or exceeded 90%.
The fitted coefficients , , and constituted three additional features. These parameters provide more physically interpretable information than curve-shape descriptors alone because they separately represent low-current fixed-loss behaviour, proportional converter losses, and high-current ohmic effects. The coefficient of determination and RMSE were retained as diagnostic quantities for evaluating the numerical adequacy of each vehicle-level fit. They were not included in the primary count of 24 behavioural features because they describe model-fit quality rather than the underlying electrical behaviour of the OBC.
Power quality features were derived from the measured power factor, denoted by
, and the measured per-phase reactive power, denoted by
. Their relationships to active and apparent power are
where
is the apparent power. Four power-factor features were extracted: power factor at 6 A, power factor at 16 A, minimum measured power factor, and the smallest current set-point at which
, denoted by
. Six reactive-power features were extracted: reactive power at 6 A, reactive power at 16 A, mean absolute reactive power, maximum absolute reactive power, the ordinary-least-squares slope of
in var/A, and a binary indicator identifying a capacitive-to-inductive or inductive-to-capacitive sign transition. Together, the 11 efficiency curve descriptors, three loss coefficients, four power-factor features, and six reactive-power features produced the 24-dimensional analytical representation.
For the 21 vehicles with matched three-phase and curtailed single-phase efficiency measurements, a same-current configuration difference was calculated over the common set of valid per-phase current set-points:
where M denotes the set of per-phase current values measured under both configurations. This within-vehicle comparison controls for vehicle identity and current set-point, but it does not hold total charging power constant: at the same per-phase current, single-phase operation transfers approximately one third of the power of three-phase operation. The statistic is therefore interpreted as a combined phase-configuration and total-throughput difference, not as a causal phase-only effect.
3.3. Exploratory Clustering, Feature Selection, and Stability Analysis
The 24 variables were retained for descriptive characterisation, but they were not all entered into K-means. The primary clustering set was reduced to eight variables chosen to represent complementary electrical dimensions while excluding identifiers, sparse threshold variables, binary indicators, direct duplicate-current measures, and an exact algebraic redundancy: peak efficiency (ηmax), the 6 A part-load difference (Δη6 = ηmax − η6), fitted coefficients a, b, and c, minimum power factor (PFmin), log10(mean |Q|), and reactive-power slope dQ/dI. The η6 descriptor was excluded because η6 = ηmax − Δη6, so retaining ηmax, η6, and Δη6 would exactly duplicate one degree of information. The logarithmic transformation was applied to mean absolute reactive power because its magnitude spans approximately two orders. Each selected variable was standardised before clustering:
where μj and σj are the sample mean and sample standard deviation of feature j. Missing power quality descriptors affected three of the 36 efficiency-characterised test units and were median-imputed after feature assembly. The remaining selected variables were available from the measured or fitted data. The Spearman screening showed that Δη6 and a remained strongly associated (ρ = 0.94); therefore, reduced-feature reruns excluding either member of this pair were included as sensitivity checks rather than claiming complete statistical independence among the eight variables.
Figure 4 illustrates the exploratory two-stage clustering and stability-analysis workflow.
Stage 1 applied K-means clustering with 50 independent initialisations and a fixed random seed of 42. For a candidate partition containing
clusters, the method minimised the within-cluster sum of squares:
where
is the standardised feature vector of vehicle
and
is the centroid of cluster
. Candidate solutions were evaluated for
. The preferred number of clusters was selected using the mean silhouette coefficient [
22]. For observation
,
where
is the mean distance between observation
and the other members of its assigned cluster, whereas
is the smallest mean distance between
and any alternative cluster.
The Stage 1 partition was independently evaluated using agglomerative hierarchical clustering with Ward linkage. At each step, the pair of clusters producing the smallest increase in total within-cluster variance was merged according to
Dendrogram fidelity was quantified using the cophenetic correlation coefficient:
where
is the original pairwise distance between observations
and
, and
is the corresponding cophenetic distance in the hierarchical tree. Agreement between the K-means and Ward partitions was quantified using the adjusted Rand index [
23]:
An ARI of 1 indicates identical partitions, whereas a value near 0 indicates agreement no greater than expected under random labelling.
The two-stage strategy was not preregistered or specified independently of the data. It was adopted after the silhouette-optimal Stage 1 solution (k = 2) isolated four electrically extreme motor-winding-integrated chargers from the remaining 32 units. Stage 2 then re-standardised and reclustered the 32-unit dedicated OBC subset for k = 2, …, 5. The resulting four subgroups, together with the Stage 1 minority group, define the five-group reference partition used for descriptive comparisons. Because this is a data-informed second step and Stage 2 separation is only moderate, the groups are not presented as universal OBC behavioural groups.
Stability was evaluated in four ways. First, Ward hierarchical clustering was compared with K-means using the adjusted Rand index (ARI) and cophenetic correlation. Second, 1000 vehicle-level bootstrap resamples were fitted with the same two-stage pipeline; labels were optimally aligned to the reference partition and ARI, assignment accuracy, and groupwise Jaccard overlap were recorded. Third, each test unit was removed in turn, and the full pipeline was rerun, with both fixed-k agreement and silhouette-selected k recorded. Fourth, sensitivity analyses removed the two-unit high-efficiency/reactive group, removed the integrated-charger group, and repeated clustering after excluding highly correlated efficiency/loss descriptors. These analyses test the dependence of the finer partition on individual vehicles, small groups, and feature choice.
Principal component analysis was used exclusively for visualisation and did not determine cluster membership. For the standardised feature matrix
, the sample covariance matrix was calculated as
The principal axes were obtained from
where
and
are the
-th eigenvalue and eigenvector. The first two eigenvectors defined the biplot axes, while the corresponding feature coefficients were displayed as loading vectors.