3.1. Formalising CBM as a Condition Monitoring and Decision-Making Problem in a Multichannel Multi-Fidelity System
CBM for a marine diesel engine is formulated as a condition monitoring and decision-making problem under (i) failure rarity, (ii) operating-condition variability, and (iii) heterogeneous measurement quality. CBM is treated not as “a single classifier model”, but as a reproducible decision logic that, based on available observations, produces a binary decision—no alarm () or alarm (A)—and, when required, triggers subsequent fault diagnosis (FD) to identify the likely cause. In the present paper, this logic is treated as a formal design language for maintenance triggering, in which the structure of decision stages and the associated parameters are explicit design variables subject to verification through CRIs.
Terminology and scope. Unless stated otherwise, terminology follows ISO 13372 and ISO 2041, while the general CM/diagnostics workflow is interpreted in line with ISO 17359 [
60,
61,
62]. In ISO 13372, predictive maintenance is deprecated in favour of condition-based maintenance; throughout the paper, “PdM” is used only as an industry synonym. We use fault for a state and failure for an event (loss of function), consistent with ISO usage. Finally, the proposed control-reliability indicators quantify correctness of monitoring decisions (TN/TP/FA/MD) and should not be confused with equipment reliability in the ISO sense.
The following definitions and notations are used.
Diagnostic variable . The diagnostic variable is the quantity used to infer the state of a component or subsystem and detect abnormalities. The variable can be (i) a physical parameter (e.g., charged-air pressure, exhaust gas temperature (EGT), turbocharger speed) or (ii) a computed diagnostic indicator (descriptor), such as a residual, reconstruction error, or anomaly score, obtained from measurement data and a model. In ML-enabled CBM/PdM, such computed indicators are typically the direct outputs of AI/ML models trained to represent normal behaviour or to estimate state from limited signals; in this case, the monitoring channel operates on a model-derived statistic rather than on a single physical measurement.
Measurement path
j. A measurement path
j is a specific route by which an estimate of
is obtained, including the sensor, signal conversion, transmission, computation, and preprocessing. The path is characterized by the resulting error
and, in the general case, by delays, missing data, and noise. The measurement outcome in path
j is described by the model:
For computed diagnostic statistics, Equation (
1) is treated as an estimation-error model:
denotes the latent (true) value of the diagnostic indicator, whereas
is an aggregate random term that represents the combined uncertainty introduced along measurement path
j. This term subsumes instrumentation and acquisition effects (measurement noise, calibration changes, time misalignment, missing-data handling) as well as processing and modelling effects (preprocessing, feature extraction, operating-condition normalisation, and model mismatch).
Accordingly, when is a directly measured physical parameter, is dominated by instrumentation and acquisition uncertainty; when is a computed diagnostic statistic (e.g., a residual of a normal-behaviour model, an autoencoder reconstruction error, or an anomaly score), also includes the approximation error of the underlying model and the variability introduced by preprocessing. In all cases, is the observable diagnostic quantity used for decision making in a monitoring channel, and a decision region is defined in the space of .
Regardless of whether is a physical measurement or an AI/ML-derived diagnostic statistic, the operational task is to map to a reproducible alarm/no-alarm decision with controlled risks. In this formulation, the CRI-based layer is placed downstream of the AI/ML module and upstream of maintenance actions, which is typical of cyber–physical PdM systems.
Monitoring channel i. A monitoring channel i aggregates one or more measurement paths j used to monitor a single diagnostic variable . A channel may include multi-fidelity measurement paths, for example, a “high-accuracy” sensor and a “coarse” or indirect source, different sensor types, or a combination of a measurement and a model-based estimator.
Proxy indicator. A proxy indicator is a quantity that does not coincide directly with but is informative about the system state. Proxies may be direct measurements (e.g., generator electrical quantities) or derived descriptors (e.g., RMS, kurtosis, spectral peaks of a vibration signal, or time-series statistics). Proxies are introduced either as separate diagnostic variables or as an additional measurement path within a channel when the proxy is used for confirmation or refinement.
Fusion. Fusion denotes a rule for combining evidence across measurement paths within a channel and/or across channels of the overall system. Fusion is used broadly, ranging from simple logical combination (e.g., “a out of m”, majority voting, averaging, or weighted fusion) to multi-level aggregation of decisions (e.g., confirmation over observation sequences, alarm escalation, and consistency checks between independent channels).
The complete list of symbols and notation is summarized in
Appendix A.
From a decision-making perspective, each monitoring channel produces a binary decision based on the observation(s) :
At the CBM level, the decision is constructed in a multi-stage manner. First, a data quality gate is applied. Next, threshold checks are performed at the measurement-path level. This is followed by confirmation using repeated measurements, repeated decision cycles, or an alternative measurement path. The workflow may then proceed to the two-stage “AD–FD” separation (abnormality detection followed by fault diagnosis). Finally, the logic aggregates evidence across channels via fusion and issues an action (observe/alert/intervene).
In the subsequent methodology, solution quality is not defined solely through model metrics; it is quantified using probabilistic indicators of decision correctness—the risks of false alarms and missed detections—in an unconditional setting, i.e., accounting for typical operating conditions and failure rarity. This provides a direct link between decision-logic design and verifiable operational requirements.
3.2. Object and Data/Signal Loop
The application context of this study is a marine diesel engine and its air-handling system, including subsystems that affect cylinder charging, thermal conditions, and turbocharging operation. The data loop comprises measurements available from the ship’s standard automation system and/or additional monitoring devices; the available channels, sampling rates, and data completeness depend on the specific vessel and interface implementation.
To ensure comparability under variable operating conditions, the diagnostic variables are interpreted in the context of operating conditions (load, speed, and thermal conditions). Consequently, within the CBM logic, some thresholds and confirmation criteria are defined not in absolute terms, but as functions of operating conditions or as deviations from expected normal behaviour at a given operating condition (e.g., via residuals of a model of normal behaviour).
Measurement quality and availability in marine operation are heterogeneous (see
Table 1 for a representative mix of channels reported in [
10]). For a practical CBM formulation, we consider the following typical constraints:
Sparse measurements (low logging rate, irregular recording);
Missing values, data gaps, and incomplete intervals;
Noise and outliers (spikes), sensor drift, and calibration shifts;
Delays and time misalignment between channels;
Unavailability of some parameters due to missing direct interfaces or constraints on sensor installation;
Unequal accuracy and noise immunity across measurement paths (multi-fidelity).
The computational experiment is used as a verification protocol for decision-logic elements rather than as a data-driven identification of an engine model. To isolate the contribution of thresholding, averaging, confirmation, and multi-fidelity fusion, we use a dimensionless Gaussian baseline with symmetric regions and a prescribed abnormal-state prior. This controlled baseline is intentionally simplified and serves as a reference for comparing CRI changes induced by individual logic stages.
3.3. Monitoring Model and CRIs
This section introduces unconditional control-reliability indicators (CRIs) in a self-contained probabilistic setting for a single monitoring channel and a single effective measurement path. The formulation is conceptually related to earlier work on decision-correctness indicators in multiparameter control, including the formulation reported by Filatov et al. [
63], which is cited as an early related source rather than as the sole basis of the method.
The true state with respect to the diagnostic variable
X is defined by a normal region
and an abnormal region
. In the simplest threshold formulation, the normal region is an interval:
where
and
are the lower and upper normal bounds.
The channel decision is formed from the observable value
Z using the no-alarm decision region
:
where
and
are the lower and upper decision bounds;
denotes the no-alarm decision and
A denotes the alarm decision.
Here,
describe the true state with respect to
X, whereas
denote the decision made from the observed
Z based on
. Using Equation (
1), the no-alarm condition
can be written as a constraint on the measurement error
Y (for a given
X):
The corresponding decision outcomes (TN/TP/FA/MD) are summarized in
Table 2.
In the CBM interpretation,
corresponds to the conditional false-alarm risk under a true normal state, i.e.,
(Equation (
6)), whereas
corresponds to the conditional missed-detection risk under a true abnormal state, i.e.,
(Equation (
7)).
Unconditional CRIs are estimated using Monte Carlo simulation. We simulate pairs of random variables according to the prescribed distributions and . The distribution is specified as a mixture over operating conditions (e.g., by load and speed), and may also depend on operating conditions. Therefore, the unconditional CRIs are integral with respect to the actual operating profile.
For each realization, we compute
, determine the true state (via whether
), and determine the decision (via whether
). For each of the four situations, we count the number of hits
and estimate the corresponding indicator as:
where
N is the number of Monte Carlo trials. The hat notation denotes a Monte Carlo estimate of the corresponding population CRI, obtained from the relative frequency of the event in
N simulated trials.
The use of unconditional CRIs is motivated by the practical difficulty of ensuring reference monitoring with negligible error (i.e., “ideal” measurement paths). As a result, directly defining Type I and Type II error probabilities is nontrivial. However, the conditional error characteristics
and
can be obtained from the unconditional CRIs as follows:
To compare decision-logic variants (different thresholds, confirmation rules, numbers of repeats, etc.), it is convenient to use an integral performance index defined as an additive aggregation of normalised CRIs:
where
,
,
, and
are normalised values of the corresponding indicators, and
are expert-defined importance weights reflecting the relative cost of false alarms and missed detections in the considered CBM loop.
For and , normalisation is performed as for indicators whose values should be increased, whereas for and it is performed as for indicators whose values should be decreased. Linear normalisation is defined over an admissible (engineering-relevant) value range:
if an increase of the indicator is desired,
if a decrease of the indicator is desired,
Here,
k is the indicator being normalised, and
and
define not the theoretical probability bounds
, but an admissible range specified relative to a baseline decision-logic variant or to target (acceptable) indicator levels for the considered CBM loop. Because these normalisation bounds are defined relative to a baseline variant or to an admissible design set, the ranking of alternatives may depend on the chosen practical range. Accordingly, the normalisation in Equations (
9) and (
10) is used here as a tool for engineering comparison within a specified set of logic variants rather than as a universal optimality statement.
This choice follows from the relationships between “correct” and “erroneous” outcomes implied by the definitions in
Table 2. For a true normal state
, the decision outcomes form a partition into two mutually exclusive events:
and for a true abnormal state
:
It follows that
and, consequently, when the decision logic is modified (thresholds, confirmation rules, numbers of repeats),
That is, reducing the false-alarm risk
by an absolute amount
automatically increases
by the same amount
, while reducing the missed-detection risk
by
increases
by
. At the same time, a multiple-fold reduction of errors in relative terms corresponds to a moderate increase in
and
in percentage points. For example, if
is reduced by a factor of
r, then
For
–3 and, for example,
, we obtain
–
, i.e., an increase of
by 3–4 percentage points while
is reduced by a factor of 2–3. The reasoning is fully analogous for the pair
and
:
For this reason, when applying the normalisation in Equations (
9) and (
10), it is advisable to set
and
as a practical range of attainable or admissible values relative to a baseline variant or to CBM requirements, rather than using the formal bounds
. Otherwise, the contributions of metrics that change “by a factor” (error probabilities) and metrics that change by a few percentage points (correct-decision probabilities) can be distorted disproportionately.
In a special case, a simplified index focusing only on errors can be used:
Within the normalised-criterion formulation, the decision-logic parameters (thresholds, confirmation rules, and numbers of repeats) can be selected by maximising
or, in the simplified variant,
. In the reported optimisation runs of this paper, however, parameter selection is performed using the raw error criterion
, as described in
Appendix B.
3.4. CBM Multi-Stage Decision Logic
We next introduce an algorithmic–structural model of a monitoring channel as a formal scaffold for synthesising the multi-stage decision logic, analogous to the prototype in [
58]. The model makes explicit (i) repeated measurements and averaging within a measurement path (
), (ii) within-path confirmation over repeated sub-cycles (
out of
), (iii) aggregation across active multi-fidelity measurement paths within a channel (
out of
), and (iv) temporal filtering over a decision window (
out of
). The output is the binary channel decision
(no alarm) or
(alarm).
3.4.1. Structural Channel Parameters and Their Relation to CRIs
Consider a single monitoring channel i that may include measurement paths (in general, multi-fidelity), . Within each path, repeated measurements and repeated decision cycles are allowed, together with multi-level aggregation of decisions (Stages 0–5).
The monitoring channel is described by the following set of design parameters (structural and threshold parameters):
where
is the number of repeated measurements within one decision cycle in measurement path j (for averaging and suppression of random variability);
is the number of repeated decision cycles in measurement path j (for confirmation and temporal stability within the path);
is the number of measurement paths in the channel (multi-fidelity and/or redundancy);
is the number of upper-level decision cycles of the channel (filtering of single triggers);
, , and are k-out-of-n thresholds (the minimum required number of no-alarm decisions ), respectively: (a) for confirmation within one measurement path over cycles; (b) for aggregating decisions across measurement paths; (c) for aggregating across upper-level channel cycles.
Functionally, the channel model links (i) the distribution of
and the error/uncertainty distributions of the measurement paths, (ii) the regions
and
, and (iii) the parameters
to unconditional CRIs and engineering constraints:
where
and
are the total cost and time characteristics of the channel,
and
are the characteristics of the elements composing measurement path
j, and
and
are factors accounting for the contribution of switching and logical devices. Equation (
19) formalizes the key position of this paper: the synthesis target is the decision logic (the structure and parameters
), while quality is evaluated through unconditional CRIs computed using the same definitions of
and
as in
Section 3.3.
3.4.2. Multi-Stage Decision Logic in Monitoring Channel
Below, Stages 0–4 define the within-cycle structure of the proposed decision-logic framework, while Stage 5 aggregates the resulting outcomes over a decision window of upper-level cycles to enforce decision persistence and produce the final channel decision .
Within each upper-level cycle k, repeated measurements are indexed by , and repeated sub-cycles used for within-path confirmation are indexed by . The primary (sub-cycle) decision in path j is , the confirmed path decision for cycle k is , the channel decision in cycle k is , and the final channel decision over the window is .
Stage 0. Data quality gate. For each measurement path
j, we introduce a binary flag indicating whether the data are suitable for decision making in the current decision window:
where
means that the path is admitted for decision formation in cycle
k (no critical missing data, gross outliers, obvious physical infeasibility, detectable drift, time misalignment, etc.). The flag
is subsequently used to exclude unsuitable measurement paths from aggregation.
Stage 1. Threshold check (with within-cycle averaging). Within measurement path
j, repeated observations
are acquired in sub-cycle
q (
) according to the adopted observation model:
To suppress random variability, averaging (or another robust aggregator; the baseline implementation uses the mean) is applied:
A primary binary decision for measurement path
j in sub-cycle
q is then formed using the no-alarm decision region
:
Here, may be fixed, operating-condition dependent, or adaptive (e.g., defined via a residual or an anomaly score); however, in all cases the primary decision is based on versus .
Stage 2. Within-path confirmation. For measurement path
j,
sub-cycles are executed,
. The confirmed path decision for cycle
k is defined by an “
out of
” rule (the minimum required number of no-alarm decisions
):
where
is the event indicator. The parameter
controls confirmation strictness, ranging from a single admissible decision (
) to a strict requirement of persistent normality (
). In the reported experiment, the confirmation rule is evaluated at one reference setting,
.
Stage 3. AD–FD separation (after confirmation). Within the adopted two-stage “AD–FD” formulation:
The AD stage corresponds to abnormality detection according to Equations (
23) and (
24), where
is typically constructed from an AI/ML model of normal behaviour (e.g., a residual, reconstruction error, or anomaly score) and calibrated as a function of operating conditions;
The FD stage is triggered only after a confirmed alarm (e.g., when
holds for one or more measurement paths after Equation (
24)) and is implemented by a separate module (classifier/rule-based/hybrid) that refines the likely cause or fault class. The FD module does not substitute for CRI-based risk control: it operates on top of a confirmed event and does not remove the need to control the risks
at the decision-logic level.
Stage 4. Aggregation of multi-fidelity measurement paths within a channel. Let
denote the set of active measurement paths in cycle
k, and let
. If
, the channel does not form a per-cycle decision in that cycle. When aggregating over the decision window, summation is performed over the set of valid cycles
The channel output over the window is considered defined if , where is an operational policy parameter.
Decision-level aggregation. The channel decision in a single upper-level cycle
k is formed by a “
out of
” rule:
Here, is the minimum required number of measurement paths that return a no-alarm decision in cycle k. For example, corresponds to a strict “all paths agree on normality” rule, while allows partial redundancy and multi-fidelity operation.
Measurement-level aggregation. In addition to decision-level fusion, measurement-level fusion is also allowed. In this case, we first form one cycle-level averaged estimate per path in cycle
k:
In the special case
(used in the corresponding experiment series),
. The fused estimate and the resulting per-cycle channel decision are then:
Both approaches (decision-level and measurement-level) produce a binary output
and are evaluated comparably via unconditional
,
,
, and
using the method in
Section 3.3.
In the reported experiment, measurement-level fusion is evaluated with fixed weights.
Stage 5. Temporal filtering at the channel level (
K-cycle logic). To suppress single triggers and outliers, the channel decision is repeated over the decision window and the final channel decision is defined by a “
out of
” rule, where
:
Equivalently, an alarm is declared only after accumulating at least alarm outcomes over the valid cycles in the decision window, which provides a formal mechanism for escalation and decision persistence.
In summary, the multi-stage logic (Stages 0–5) defines an unambiguous mapping from observations to the channel decision for given , , and . This full mapping serves as the architectural and notational basis of the framework used in the subsequent computational experiment.
3.5. Computational Experiment Plan
The computational experiment is organised as four Monte Carlo series. The reported series address baseline thresholding, repeated-measurement averaging, within-path confirmation, and measurement-level multi-fidelity fusion. In all reported series, solution quality is evaluated via unconditional CRIs
,
,
, and
(
Section 3.3), together with the corresponding conditional risks
and
(Equations (
6) and (
7)) and the selected comparison criterion (e.g.,
).
Series 1 (baseline threshold logic): (i) calibrate the baseline threshold formulation for the prescribed abnormal-state prior (normal-region bound set via ); (ii) select the optimal decision bound for fixed measurement-path accuracy; and (iii) assess sensitivity to s and the – trade-off as functions of .
Series 2 (within-cycle averaging): quantify the effect of averaging repeated measurements (representative point ) using the fixed thresholds obtained in Series 1.
Series 3 (within-path confirmation): assess the effect of “a out of m” confirmation over repeated cycles. Here, the setting is used as an illustrative reference case for comparison rather than as a universally optimal confirmation policy. For shipboard application, confirmation parameters should be linked to the sampling interval, the confirmation window, and the expected time scale of fault evolution.
Series 4 (multi-fidelity fusion): assess the effect of two-path multi-fidelity fusion using a measurement-level combination with a fixed fusion weight. In this series, the weight w is optimised within a controlled setting and treated as a constant across the operating profile in order to isolate the contribution of multi-fidelity fusion. This formulation is intended for comparative analysis rather than as a final weighting policy for practical shipboard use.
The series parameters, fixed settings of
(Equation (
18)), and the controlled outputs are summarized in
Table 3.
The reported series should be interpreted as control-point comparisons within a unified verification framework rather than as a full optimisation over the decision-logic design space. Accordingly, the reported gains quantify the behaviour of selected logic elements at representative points, not universal optimal settings for shipboard deployment.