Next Article in Journal
Progressive Damage Failure Criterion Establishment and Collapse Period Prediction for Coalbed Methane Wellbore: A Numerical Simulation Study
Next Article in Special Issue
Qualitative Analysis of Signaling Networks Using Petri Nets and Invariant Computation
Previous Article in Journal
Modular Artificial Neural Network to Classify Materials and Determine Water Content in Soil Samples
Previous Article in Special Issue
Real-Time Experimental Benchmarking of Control Strategies for a Coupled 2-DOF Helicopter
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Synthesis of Decision Logic for Predictive Maintenance of a Marine Diesel Engine Based on Unconditional Control-Reliability Indicators

Department of System Analysis and Control, Empress Catherine II Saint Petersburg Mining University, 199106 Saint Petersburg, Russia
*
Author to whom correspondence should be addressed.
Eng 2026, 7(5), 190; https://doi.org/10.3390/eng7050190
Submission received: 5 March 2026 / Revised: 16 April 2026 / Accepted: 22 April 2026 / Published: 23 April 2026
(This article belongs to the Special Issue Interdisciplinary Insights in Engineering Research 2026)

Abstract

This paper proposes a formal framework for synthesizing multi-stage condition-based maintenance (CBM) decision logic for marine diesel monitoring systems. The design object is treated not as a single threshold or classifier output, but as an implementable decision logic with explicit stages of data-quality gating, thresholding, confirmation, fusion, and temporal filtering. Decision quality is evaluated using unconditional control-reliability indicators (CRIs) under a prescribed prior probability of rare abnormal events within a unified Monte Carlo verification protocol. Within a simplified Gaussian surrogate model, we compare baseline thresholding, repeated-measurement averaging, within-path confirmation, and measurement-level fusion. For the reported reference configuration, averaging five repeated measurements yields the largest reduction in the raw error criterion, “2 out of 3” confirmation provides a smaller but consistent improvement, and two-path multi-fidelity fusion is beneficial only after calibration toward the more informative path. The results show that, under rare abnormal events and limited measurement accuracy, decision quality is determined primarily by calibration of the multi-stage channel-level logic rather than by thresholding alone.

1. Introduction

Marine diesel engines and their air-handling systems operate over long voyages under variable operating conditions and strict, safety-driven constraints on intervention [1,2,3]. In this setting, degradation and incipient faults often develop amid normal parameter fluctuations, while failure events and labeled fault cases are rare and thus underrepresented in operational datasets [4,5,6]. Therefore, the practical task of early abnormality detection (anomaly detection, AD) and subsequent fault diagnosis (FD) is addressed within a condition-based maintenance (CBM) framework (often referred to in industry as predictive maintenance, PdM), where transparent and reproducible decision logic is as important as model accuracy [7,8,9].
In contemporary PdM practice, the observable quantity that triggers decisions is frequently not a raw measurement but a diagnostic statistic computed by a normal-behaviour or state-estimation AI/ML (artificial intelligence and machine learning) model (e.g., a residual, an autoencoder reconstruction error, or an anomaly score) [10,11]. Such model-derived statistics are attractive under limited instrumentation, yet their outputs are uncertain, operating-condition dependent, and prone to spurious excursions; therefore, they require explicitly specified thresholding, confirmation, and fusion rules with controlled error risks. This paper addresses the engineering interface between AI/ML-based AD/FD modules and the shipboard cyber–physical monitoring loop by synthesising decision logic that converts model outputs into reproducible maintenance triggers.
A practical feature of shipboard CBM is limited data quality and availability. Using operational data from an in-service vessel, Upadrashta and Wijaya showed that datasets can be sparse, noisy, and incomplete [10]. They also presented an end-to-end CBM architecture spanning sensors, onboard edge computing, cloud services, and a monitoring dashboard, including a pragmatic workaround for limited parameter access via image-based acquisition from display screens when direct interfaces are unavailable [10]. The set of feasible monitoring signals is therefore often constrained by what can be accessed onboard (see Table 1 for an example). Related CBM architectures are described in [12,13].
A second persistent source of uncertainty is the rarity of failure cases and the resulting class imbalance [14,15,16]: normal observations far exceed failure-related observations, so conventional accuracy metrics can overestimate performance and understate operational risk [17]. For marine main-engine diagnostics based on operational data, the shortage of failure examples and the need to account for it during model development and validation are emphasized in [11]. Studies focused on adapting to imbalanced datasets and updating models during operation show that imbalance can be mitigated algorithmically (e.g., via synthetic data augmentation and feature dimensionality reduction), yet this alone does not guarantee controlled risks of false alarms and missed detections in service [18].
Marine diagnostics is almost always multichannel and multi-fidelity [19,20,21]. Some parameters are measured directly, whereas others are available only through proxy indicators or with limited accuracy; therefore, the set of monitoring channels is largely determined by engineering feasibility and measurement cost [22]. Applied studies emphasize that simple, operationally feasible proxy indicators (e.g., combustion-process and fuel-consumption parameters) can support condition monitoring without invasive measurements [23,24,25]. In parallel, non-traditional sources (e.g., generator electrical signals and their spectral descriptors) have been shown to improve observability without increasing physical access requirements to engine components [26,27,28]. Given this diversity of measurement paths, a central question is how to fuse heterogeneous evidence and make reliable decisions under conflicting, incomplete, or noisy observations [29,30,31].
Related work also considers formalised decision making under deployment constraints, data uncertainty, and asymmetric error costs. In marine applications, optimisation models and simulation-based analyses have been used to select decisions under time, cost, and technical constraints [32,33,34,35,36]. Decision criteria under uncertainty are supported by scenario analysis and economic performance indicators [37,38], while the interpretation of losses and managerial consequences is developed in methods for economic assessment of resource losses [39]. From a safety perspective, multi-stage strategies are used when adverse outcomes are costly, which is methodologically consistent with confirmation and escalation in CBM decision logic [40]. Related work on proxy modelling and computationally efficient models further supports treating diagnostic quantities as computed indicators that must be embedded into a multi-stage decision logic [25,41,42]. Practical deployability is illustrated by implemented monitoring systems [43,44] and by engineering solutions aimed at reducing intervention frequency and improving process efficiency [45,46,47]. Finally, the need for reproducible tuning is reinforced by changes in operating profiles caused by alternative diesel-fuel components (biodiesel/green diesel) [48,49] and the associated shifts in diagnostic-indicator statistics [50,51,52].
CBM decision making for a marine power plant is considered here under three practical constraints: limited deployability and data accessibility onboard, severe class imbalance due to failure rarity, and heterogeneous (multi-fidelity) measurement quality. The task is to define a channel-level logic that maps the observed statistic Z to an alarm/no-alarm decision and to compare how thresholding, averaging, confirmation, and fusion change α , β , and J err , raw under the same uncertainty model.
As a reference context, we use the practical two-stage “AD–FD” setup and the end-to-end (E2E) CBM architecture demonstrated on an in-service vessel in [10]. Within this context, Section 3 formulates the proposed decision logic and evaluates selected mechanisms through unconditional control-reliability indicators (CRIs) in a Monte Carlo setting.
The goal of this paper is to synthesise a CRI-based CBM decision-logic framework for marine diesel engines and to verify selected logic components via Monte Carlo simulation on a simplified Gaussian surrogate model. The main contributions are as follows:
  • A formal design framework is proposed in which the object of synthesis is an implementable multi-stage CBM decision logic rather than a single threshold or classifier output;
  • Decision quality is quantified by unconditional control-reliability indicators (CRIs), allowing channel-level logic variants to be compared under an explicit rare-event prior and controlled operating uncertainty;
  • A structured channel description is introduced that makes the logic parameters explicit, including repeated measurements, within-path confirmation, multi-fidelity aggregation, and temporal filtering;
  • Selected logic elements are compared within a unified Monte Carlo verification protocol, which enables direct assessment of how pre-binarisation and post-binarisation mechanisms change false-alarm and missed-detection risks;
  • The framework is positioned as a downstream layer between AI/ML-based AD/FD modules and maintenance action, providing a basis for later validation on operational marine data.

2. Related Work

Studies on condition-based maintenance (CBM) for marine diesel engines and their subsystems, often discussed in practice under the broader label of predictive maintenance (PdM), develop along several applied lines. These include shipboard deployment under limited data accessibility, two-stage “anomaly detection–fault diagnosis” workflows under failure rarity, rule- and threshold-based maintenance decisions, and model-centric diagnostic systems that report strong classification metrics but leave the operational decision layer only partly specified.
Because the present paper focuses on decision logic evaluated through unconditional control-reliability indicators (CRIs), the review below concentrates on the parts of prior work that are methodologically most relevant for this task: practical deployment constraints, threshold-based triggering, confirmation procedures, state zoning, uncertainty of measurements and proxy indicators, and the possibility of combining channels with unequal fidelity within an implementable maintenance-triggering logic.

2.1. Deployment Constraints and Data Availability in Shipboard CBM/PdM

One of the most practice-oriented end-to-end CBM/PdM pipelines based on operational data is presented by Upadrashta and Wijaya [10], who study an in-service marine diesel engine under explicitly stated limitations in data quality and completeness. The importance of this work lies not only in the use of anomaly-detection and classification models as such, but also in the direct demonstration of actual shipboard constraints: parameter acquisition, local processing, data transmission, diagnostic interpretation, and the construction of labelled fault scenarios with severity grading. In methodological terms, this is important because it places classification and maintenance decisions inside a real monitoring chain affected by incomplete observations, heterogeneous signals, and class imbalance.
The set of monitored engine and turbocharger parameters reported in [10] is summarised in Table 1. This list is useful not as a descriptive inventory alone, but as an illustration of the practical channel set available in shipboard monitoring: a mixture of directly measured variables and proxy indicators with different noise levels, sampling properties, and diagnostic informativeness. Such heterogeneity motivates treating marine monitoring as a multi-channel and, in a broader sense, multi-fidelity decision problem rather than as a single-signal thresholding task.
From the viewpoint of the present study, this type of deployment context is fundamental because it shows that the decision problem in marine CBM/PdM is not limited to selecting a classifier. It also requires an explicit logic for handling channel heterogeneity, imperfect measurements, and operational triggering under rare adverse conditions.

2.2. From Anomaly Scores to Maintenance Decisions: Thresholds, Confirmation, Zones, and Fusion

A central methodological line directly related to decision-logic synthesis is the two-stage workflow in which anomaly detection is trained on data representing healthy operation, while fault diagnosis is performed only after failure data become available or are progressively accumulated [53,54,55]. In [10], this idea is implemented explicitly: abnormality detection is based on the reconstruction error of an autoencoder with a threshold of the form T = μ E + k σ E , after which supervised diagnosis is performed across fault classes and severity levels. Even in this relatively standard form, thresholding already acts as a first-level decision rule and therefore directly influences false-alarm and missed-detection behaviour near the boundary of normal fluctuations, especially under early-stage degradation and severe class imbalance.
A related line is represented by model-centric diagnostics that operate with a limited set of informative signals and achieve high test-set performance but leave the downstream decision structure comparatively implicit. Kim et al. [11], for example, diagnose a low-speed marine main engine from exhaust-gas temperature using a hybrid deep-learning model. Two aspects are especially relevant here. First, the authors explicitly consider controlled overlap between classes, which is consistent with increasing decision uncertainty under degradation, noise, and blurred state boundaries. Second, despite the claimed real-time applicability, the final output remains essentially a model-level class label rather than a maintenance-triggering logic with explicit re-check, confirmation, escalation, or persistence rules.
Another important line is formed by rule-based and zone-based approaches, in which the maintenance decision is represented not only through a score or class label but through interpretable operating regions associated with actions. In [56], Karatug et al. propose a hybrid strategy combining condition-based maintenance with reliability-centred maintenance by partitioning the time to functional failure into zones such as ideal state, warning, pre-failure, and failure. The importance of this approach for the present paper lies in the structure of the decision itself: each zone corresponds to an action class, which makes the logic operationally meaningful even when the boundaries between zones are not explicitly calibrated through probabilities of false alarms and missed detections.
Under scarce failure data, a closely related strategy is residual-based monitoring, where a model of normal behaviour is first constructed and alarms are then generated from deviations relative to thresholds or fault boundaries. This line is illustrated by Ceglie et al. [57], where an artificial neural network is trained on normal data and fault boundaries are formed from simulation results; an alarm is issued when several residuals jointly enter a region interpreted as fault-related. Methodologically, such formulations are important because they naturally motivate confirmation over observation sequences and the explicit use of alarm history when noise, transients, or operating-condition variability can produce spurious single-point decisions.
Additional motivation for an explicit multi-fidelity decision layer comes from studies based on proxy indicators and indirect measurements, which recognise that ideal diagnostic channels are often unavailable in practice. Duc Nghia et al. [58] relate fuel-injection parameters to smoke and soot as practical diagnostic indicators, while Markelova and Ivanovskaya [23] discuss diagnostics based on indirect indicators with explicit threshold rules under shipboard measurement constraints. From another perspective, Gasparjans et al. [26] use generator electrical quantities and spectral descriptors as an alternative diagnostic path. By contrast, works primarily focused on class imbalance and model adaptation address algorithmic robustness, for example through deep Q-network (DQN)-based diagnosis on simulated data [59] or through schemes based on the Borderline Synthetic Minority Over-sampling Technique (BorderlineSMOTE), kernel principal component analysis (KPCA), and online sequential extreme learning machine (OSELM) with online updating [18], but usually leave heterogeneous channel trust, confirmation, and action logic only implicit. Taken together, these studies support the need for a separate decision-logic layer capable of comparing, calibrating, and combining channels of unequal fidelity under incomplete and noisy observations.

2.3. Methodological Gap and Positioning of This Paper

The reviewed literature indicates a recurring gap between model-centric diagnostics and operational maintenance decision making. In many studies, the effective decision layer is reduced to a single threshold on an anomaly score, residual, or classifier output, whereas confirmation, escalation, temporal persistence, and multi-channel fusion remain only partly specified or are absorbed implicitly into the model architecture.
In the present paper, we therefore treat the decision logic itself, together with its structure and parameters θ i , as the design object. The proposed formulation is methodologically close to classical tools such as Type I/Type II error analysis and ROC-style reasoning, but it differs in one essential respect: for shipboard application, the object to be calibrated and compared is not a single threshold or classifier score but an implementable decision logic composed of channel-level operations.
More specifically, the contribution is formulated as a decision-layer synthesis framework in which thresholding, repeated-measurement averaging, within-path confirmation, measurement-level fusion, and temporal filtering can be structured, tuned, and compared under a common probabilistic verification protocol using unconditional CRIs. In this sense, the paper is positioned not as another classifier-evaluation study, but as a formal and reproducible scheme for maintenance-trigger design under rare abnormal events, limited measurement quality, and heterogeneous diagnostic channels.

3. Materials and Methods

3.1. Formalising CBM as a Condition Monitoring and Decision-Making Problem in a Multichannel Multi-Fidelity System

CBM for a marine diesel engine is formulated as a condition monitoring and decision-making problem under (i) failure rarity, (ii) operating-condition variability, and (iii) heterogeneous measurement quality. CBM is treated not as “a single classifier model”, but as a reproducible decision logic that, based on available observations, produces a binary decision—no alarm ( A ¯ ) or alarm (A)—and, when required, triggers subsequent fault diagnosis (FD) to identify the likely cause. In the present paper, this logic is treated as a formal design language for maintenance triggering, in which the structure of decision stages and the associated parameters are explicit design variables subject to verification through CRIs.
Terminology and scope. Unless stated otherwise, terminology follows ISO 13372 and ISO 2041, while the general CM/diagnostics workflow is interpreted in line with ISO 17359 [60,61,62]. In ISO 13372, predictive maintenance is deprecated in favour of condition-based maintenance; throughout the paper, “PdM” is used only as an industry synonym. We use fault for a state and failure for an event (loss of function), consistent with ISO usage. Finally, the proposed control-reliability indicators quantify correctness of monitoring decisions (TN/TP/FA/MD) and should not be confused with equipment reliability in the ISO sense.
The following definitions and notations are used.
Diagnostic variable X i . The diagnostic variable X i is the quantity used to infer the state of a component or subsystem and detect abnormalities. The variable X i can be (i) a physical parameter (e.g., charged-air pressure, exhaust gas temperature (EGT), turbocharger speed) or (ii) a computed diagnostic indicator (descriptor), such as a residual, reconstruction error, or anomaly score, obtained from measurement data and a model. In ML-enabled CBM/PdM, such computed indicators are typically the direct outputs of AI/ML models trained to represent normal behaviour or to estimate state from limited signals; in this case, the monitoring channel operates on a model-derived statistic rather than on a single physical measurement.
Measurement path j. A measurement path j is a specific route by which an estimate of X i is obtained, including the sensor, signal conversion, transmission, computation, and preprocessing. The path is characterized by the resulting error Y i j and, in the general case, by delays, missing data, and noise. The measurement outcome in path j is described by the model:
Z i j = X i + Y i j ,
For computed diagnostic statistics, Equation (1) is treated as an estimation-error model: X i denotes the latent (true) value of the diagnostic indicator, whereas Y i j is an aggregate random term that represents the combined uncertainty introduced along measurement path j. This term subsumes instrumentation and acquisition effects (measurement noise, calibration changes, time misalignment, missing-data handling) as well as processing and modelling effects (preprocessing, feature extraction, operating-condition normalisation, and model mismatch).
Accordingly, when X i is a directly measured physical parameter, Y i j is dominated by instrumentation and acquisition uncertainty; when X i is a computed diagnostic statistic (e.g., a residual of a normal-behaviour model, an autoencoder reconstruction error, or an anomaly score), Y i j also includes the approximation error of the underlying model and the variability introduced by preprocessing. In all cases, Z i j is the observable diagnostic quantity used for decision making in a monitoring channel, and a decision region ω is defined in the space of Z i j .
Regardless of whether Z i j is a physical measurement or an AI/ML-derived diagnostic statistic, the operational task is to map Z i j to a reproducible alarm/no-alarm decision with controlled risks. In this formulation, the CRI-based layer is placed downstream of the AI/ML module and upstream of maintenance actions, which is typical of cyber–physical PdM systems.
Monitoring channel i. A monitoring channel i aggregates one or more measurement paths j used to monitor a single diagnostic variable X i . A channel may include multi-fidelity measurement paths, for example, a “high-accuracy” sensor and a “coarse” or indirect source, different sensor types, or a combination of a measurement and a model-based estimator.
Proxy indicator. A proxy indicator is a quantity that does not coincide directly with X i but is informative about the system state. Proxies may be direct measurements (e.g., generator electrical quantities) or derived descriptors (e.g., RMS, kurtosis, spectral peaks of a vibration signal, or time-series statistics). Proxies are introduced either as separate diagnostic variables X i or as an additional measurement path within a channel when the proxy is used for confirmation or refinement.
Fusion. Fusion denotes a rule for combining evidence across measurement paths within a channel and/or across channels of the overall system. Fusion is used broadly, ranging from simple logical combination (e.g., “a out of m”, majority voting, averaging, or weighted fusion) to multi-level aggregation of decisions (e.g., confirmation over observation sequences, alarm escalation, and consistency checks between independent channels).
The complete list of symbols and notation is summarized in Appendix A.
From a decision-making perspective, each monitoring channel produces a binary decision based on the observation(s) Z i j :
  • A ¯ : no alarm;
  • A: alarm.
At the CBM level, the decision is constructed in a multi-stage manner. First, a data quality gate is applied. Next, threshold checks are performed at the measurement-path level. This is followed by confirmation using repeated measurements, repeated decision cycles, or an alternative measurement path. The workflow may then proceed to the two-stage “AD–FD” separation (abnormality detection followed by fault diagnosis). Finally, the logic aggregates evidence across channels via fusion and issues an action (observe/alert/intervene).
In the subsequent methodology, solution quality is not defined solely through model metrics; it is quantified using probabilistic indicators of decision correctness—the risks of false alarms and missed detections—in an unconditional setting, i.e., accounting for typical operating conditions and failure rarity. This provides a direct link between decision-logic design and verifiable operational requirements.

3.2. Object and Data/Signal Loop

The application context of this study is a marine diesel engine and its air-handling system, including subsystems that affect cylinder charging, thermal conditions, and turbocharging operation. The data loop comprises measurements available from the ship’s standard automation system and/or additional monitoring devices; the available channels, sampling rates, and data completeness depend on the specific vessel and interface implementation.
To ensure comparability under variable operating conditions, the diagnostic variables X i are interpreted in the context of operating conditions (load, speed, and thermal conditions). Consequently, within the CBM logic, some thresholds and confirmation criteria are defined not in absolute terms, but as functions of operating conditions or as deviations from expected normal behaviour at a given operating condition (e.g., via residuals of a model of normal behaviour).
Measurement quality and availability in marine operation are heterogeneous (see Table 1 for a representative mix of channels reported in [10]). For a practical CBM formulation, we consider the following typical constraints:
  • Sparse measurements (low logging rate, irregular recording);
  • Missing values, data gaps, and incomplete intervals;
  • Noise and outliers (spikes), sensor drift, and calibration shifts;
  • Delays and time misalignment between channels;
  • Unavailability of some parameters due to missing direct interfaces or constraints on sensor installation;
  • Unequal accuracy and noise immunity across measurement paths (multi-fidelity).
The computational experiment is used as a verification protocol for decision-logic elements rather than as a data-driven identification of an engine model. To isolate the contribution of thresholding, averaging, confirmation, and multi-fidelity fusion, we use a dimensionless Gaussian baseline with symmetric regions and a prescribed abnormal-state prior. This controlled baseline is intentionally simplified and serves as a reference for comparing CRI changes induced by individual logic stages.

3.3. Monitoring Model and CRIs

This section introduces unconditional control-reliability indicators (CRIs) in a self-contained probabilistic setting for a single monitoring channel and a single effective measurement path. The formulation is conceptually related to earlier work on decision-correctness indicators in multiparameter control, including the formulation reported by Filatov et al. [63], which is cited as an early related source rather than as the sole basis of the method.
The true state with respect to the diagnostic variable X is defined by a normal region Ω and an abnormal region Ω ¯ . In the simplest threshold formulation, the normal region is an interval:
Ω = [ X L , X U ] , X Ω X L X X U ,
where X L and X U are the lower and upper normal bounds.
The channel decision is formed from the observable value Z using the no-alarm decision region ω :
ω = [ Z L , Z U ] , A ¯ : Z ω , A : Z ω ,
where Z L and Z U are the lower and upper decision bounds; A ¯ denotes the no-alarm decision and A denotes the alarm decision.
Here, Ω / Ω ¯ describe the true state with respect to X, whereas A ¯ / A denote the decision made from the observed Z based on ω . Using Equation (1), the no-alarm condition Z ω can be written as a constraint on the measurement error Y (for a given X):
Z L X Y Z U X .
The corresponding decision outcomes (TN/TP/FA/MD) are summarized in Table 2.
In the CBM interpretation, α corresponds to the conditional false-alarm risk under a true normal state, i.e., α = P ( A Ω ) (Equation (6)), whereas β corresponds to the conditional missed-detection risk under a true abnormal state, i.e., β = P ( A ¯ Ω ¯ ) (Equation (7)).
Unconditional CRIs are estimated using Monte Carlo simulation. We simulate pairs of random variables ( X , Y ) according to the prescribed distributions F ( X ) and F ( Y ) . The distribution F ( X ) is specified as a mixture over operating conditions (e.g., by load and speed), and F ( Y ) may also depend on operating conditions. Therefore, the unconditional CRIs are integral with respect to the actual operating profile.
For each realization, we compute Z = X + Y , determine the true state (via whether X Ω ), and determine the decision (via whether Z ω ). For each of the four situations, we count the number of hits n hit and estimate the corresponding indicator as:
CRI ^ = n hit N ,
where N is the number of Monte Carlo trials. The hat notation denotes a Monte Carlo estimate of the corresponding population CRI, obtained from the relative frequency of the event in N simulated trials.
The use of unconditional CRIs is motivated by the practical difficulty of ensuring reference monitoring with negligible error (i.e., “ideal” measurement paths). As a result, directly defining Type I and Type II error probabilities is nontrivial. However, the conditional error characteristics α and β can be obtained from the unconditional CRIs as follows:
α = P ( A Ω ) = P ( Ω A ) P ( Ω ) = P FA P TN + P FA ,
β = P ( A ¯ Ω ¯ ) = P ( Ω ¯ A ¯ ) P ( Ω ¯ ) = P MD P TP + P MD .
To compare decision-logic variants (different thresholds, confirmation rules, numbers of repeats, etc.), it is convenient to use an integral performance index defined as an additive aggregation of normalised CRIs:
J int = γ 1 P ˜ TN + γ 2 P ˜ TP + γ 3 P ˜ MD + γ 4 P ˜ FA ,
where P ˜ TN , P ˜ TP , P ˜ MD , and P ˜ FA are normalised values of the corresponding indicators, and γ 1 , , γ 4 are expert-defined importance weights reflecting the relative cost of false alarms and missed detections in the considered CBM loop.
For P TP and P TN , normalisation is performed as for indicators whose values should be increased, whereas for P MD and P FA it is performed as for indicators whose values should be decreased. Linear normalisation is defined over an admissible (engineering-relevant) value range:
  • if an increase of the indicator is desired,
    k ˜ = k k min k max k min ,
  • if a decrease of the indicator is desired,
    k ˜ = k max k k max k min .
Here, k is the indicator being normalised, and k min and k max define not the theoretical probability bounds [ 0 , 1 ] , but an admissible range specified relative to a baseline decision-logic variant or to target (acceptable) indicator levels for the considered CBM loop. Because these normalisation bounds are defined relative to a baseline variant or to an admissible design set, the ranking of alternatives may depend on the chosen practical range. Accordingly, the normalisation in Equations (9) and (10) is used here as a tool for engineering comparison within a specified set of logic variants rather than as a universal optimality statement.
This choice follows from the relationships between “correct” and “erroneous” outcomes implied by the definitions in Table 2. For a true normal state Ω , the decision outcomes form a partition into two mutually exclusive events:
P TN + P FA = P ( Ω ) ,
and for a true abnormal state Ω ¯ :
P TP + P MD = P ( Ω ¯ ) .
It follows that
P TN = P ( Ω ) P FA , P TP = P ( Ω ¯ ) P MD ,
and, consequently, when the decision logic is modified (thresholds, confirmation rules, numbers of repeats),
Δ P TN = Δ P FA , Δ P TP = Δ P MD .
That is, reducing the false-alarm risk P FA by an absolute amount δ automatically increases P TN by the same amount δ , while reducing the missed-detection risk P MD by δ increases P TP by δ . At the same time, a multiple-fold reduction of errors in relative terms corresponds to a moderate increase in P TN and P TP in percentage points. For example, if P FA is reduced by a factor of r, then
P FA ( 1 ) = P FA ( 0 ) r , Δ P TN = P TN ( 1 ) P TN ( 0 ) = P FA ( 0 ) P FA ( 1 ) = P FA ( 0 ) 1 1 r .
For r = 2 –3 and, for example, P FA ( 0 ) = 0.06 , we obtain Δ P TN 0.03 0.04 , i.e., an increase of P TN by 3–4 percentage points while P FA is reduced by a factor of 2–3. The reasoning is fully analogous for the pair P TP and P MD :
P MD ( 1 ) = P MD ( 0 ) r , Δ P TP = P MD ( 0 ) 1 1 r .
For this reason, when applying the normalisation in Equations (9) and (10), it is advisable to set k min and k max as a practical range of attainable or admissible values relative to a baseline variant or to CBM requirements, rather than using the formal bounds [ 0 , 1 ] . Otherwise, the contributions of metrics that change “by a factor” (error probabilities) and metrics that change by a few percentage points (correct-decision probabilities) can be distorted disproportionately.
In a special case, a simplified index focusing only on errors can be used:
J err = γ 3 P ˜ MD + γ 4 P ˜ FA , γ 3 γ 4 .
Within the normalised-criterion formulation, the decision-logic parameters (thresholds, confirmation rules, and numbers of repeats) can be selected by maximising J int or, in the simplified variant, J err . In the reported optimisation runs of this paper, however, parameter selection is performed using the raw error criterion J err , raw , as described in Appendix B.

3.4. CBM Multi-Stage Decision Logic

We next introduce an algorithmic–structural model of a monitoring channel as a formal scaffold for synthesising the multi-stage decision logic, analogous to the prototype in [58]. The model makes explicit (i) repeated measurements and averaging within a measurement path ( n i j ), (ii) within-path confirmation over repeated sub-cycles ( a i out of m i j ), (iii) aggregation across active multi-fidelity measurement paths within a channel ( b i out of M i act ), and (iv) temporal filtering over a decision window ( c i out of K i act ). The output is the binary channel decision A ¯ i (no alarm) or A i (alarm).

3.4.1. Structural Channel Parameters and Their Relation to CRIs

Consider a single monitoring channel i that may include M i measurement paths (in general, multi-fidelity), j = 1 , , M i . Within each path, repeated measurements and repeated decision cycles are allowed, together with multi-level aggregation of decisions (Stages 0–5).
The monitoring channel is described by the following set of design parameters (structural and threshold parameters):
θ i = { n i j , m i j } j = 1 M i , M i , a i , b i , c i , K i ,
where
  • n i j is the number of repeated measurements within one decision cycle in measurement path j (for averaging and suppression of random variability);
  • m i j is the number of repeated decision cycles in measurement path j (for confirmation and temporal stability within the path);
  • M i is the number of measurement paths in the channel (multi-fidelity and/or redundancy);
  • K i is the number of upper-level decision cycles of the channel (filtering of single triggers);
  • a i , b i , and c i are k-out-of-n thresholds (the minimum required number of no-alarm decisions A ¯ ), respectively: (a) for confirmation within one measurement path over m i j cycles; (b) for aggregating decisions across M i measurement paths; (c) for aggregating across K i upper-level channel cycles.
Functionally, the channel model links (i) the distribution of X i and the error/uncertainty distributions of the measurement paths, (ii) the regions Ω i and ω i j , and (iii) the parameters θ i to unconditional CRIs and engineering constraints:
P TN , i , P TP , i , P FA , i , P MD , i = Φ F i ( X ) , { F i j ( Y ) } j = 1 M i , Ω i , { ω i j } j = 1 M i , θ i , C i = Ψ C θ i , { C i j } , k C , T i = Ψ T θ i , { T i j } , k T ,
where C i and T i are the total cost and time characteristics of the channel, { C i j } and { T i j } are the characteristics of the elements composing measurement path j, and k C and k T are factors accounting for the contribution of switching and logical devices. Equation (19) formalizes the key position of this paper: the synthesis target is the decision logic (the structure and parameters θ i ), while quality is evaluated through unconditional CRIs computed using the same definitions of Ω / Ω ¯ and A ¯ / A as in Section 3.3.

3.4.2. Multi-Stage Decision Logic in Monitoring Channel

Below, Stages 0–4 define the within-cycle structure of the proposed decision-logic framework, while Stage 5 aggregates the resulting outcomes over a decision window of K i upper-level cycles to enforce decision persistence and produce the final channel decision A ¯ i / A i .
Within each upper-level cycle k, repeated measurements are indexed by l = 1 , , n i j , and repeated sub-cycles used for within-path confirmation are indexed by q = 1 , , m i j . The primary (sub-cycle) decision in path j is A ¯ i j ( k , q ) / A i j ( k , q ) , the confirmed path decision for cycle k is A ¯ i j ( k ) / A i j ( k ) , the channel decision in cycle k is A ¯ i ( k ) / A i ( k ) , and the final channel decision over the window is A ¯ i / A i .
Stage 0. Data quality gate. For each measurement path j, we introduce a binary flag indicating whether the data are suitable for decision making in the current decision window:
g i j ( k ) { 0 , 1 } ,
where g i j ( k ) = 1 means that the path is admitted for decision formation in cycle k (no critical missing data, gross outliers, obvious physical infeasibility, detectable drift, time misalignment, etc.). The flag g i j ( k ) is subsequently used to exclude unsuitable measurement paths from aggregation.
Stage 1. Threshold check (with within-cycle averaging). Within measurement path j, repeated observations Z i j ( k , q , l ) are acquired in sub-cycle q ( l = 1 , , n i j ) according to the adopted observation model:
Z i j ( k , q , l ) = X i + Y i j ( k , q , l ) .
To suppress random variability, averaging (or another robust aggregator; the baseline implementation uses the mean) is applied:
Z ¯ i j ( k , q ) = 1 n i j l = 1 n i j Z i j ( k , q , l ) .
A primary binary decision for measurement path j in sub-cycle q is then formed using the no-alarm decision region ω i j :
A ¯ i j ( k , q ) : Z ¯ i j ( k , q ) ω i j , A i j ( k , q ) : Z ¯ i j ( k , q ) ω i j .
Here, ω i j may be fixed, operating-condition dependent, or adaptive (e.g., defined via a residual or an anomaly score); however, in all cases the primary decision is based on Z ¯ ω versus Z ¯ ω .
Stage 2. Within-path confirmation. For measurement path j, m i j sub-cycles are executed, q = 1 , , m i j . The confirmed path decision for cycle k is defined by an “ a i out of m i j ” rule (the minimum required number of no-alarm decisions A ¯ ):
A ¯ i j ( k ) : q = 1 m i j I A ¯ i j ( k , q ) a i , A i j ( k ) : q = 1 m i j I A ¯ i j ( k , q ) < a i ,
where I ( · ) is the event indicator. The parameter a i controls confirmation strictness, ranging from a single admissible decision ( a i = 1 ) to a strict requirement of persistent normality ( a i = m i j ). In the reported experiment, the confirmation rule is evaluated at one reference setting, ( a i , m i j ) = ( 2 , 3 ) .
Stage 3. AD–FD separation (after confirmation). Within the adopted two-stage “AD–FD” formulation:
  • The AD stage corresponds to abnormality detection according to Equations (23) and (24), where ω i j is typically constructed from an AI/ML model of normal behaviour (e.g., a residual, reconstruction error, or anomaly score) and calibrated as a function of operating conditions;
  • The FD stage is triggered only after a confirmed alarm (e.g., when A i j ( k ) holds for one or more measurement paths after Equation (24)) and is implemented by a separate module (classifier/rule-based/hybrid) that refines the likely cause or fault class. The FD module does not substitute for CRI-based risk control: it operates on top of a confirmed event and does not remove the need to control the risks ( P FA , P MD ) at the decision-logic level.
Stage 4. Aggregation of multi-fidelity measurement paths within a channel. Let
J i ( k ) = j : g i j ( k ) = 1
denote the set of active measurement paths in cycle k, and let M i act = J i ( k ) . If J i ( k ) = 0 , the channel does not form a per-cycle decision in that cycle. When aggregating over the decision window, summation is performed over the set of valid cycles
K i = k : J i ( k ) > 0 .
The channel output over the window is considered defined if K i K i , min , where K i , min is an operational policy parameter.
Decision-level aggregation. The channel decision in a single upper-level cycle k is formed by a “ b i out of M i act ” rule:
A ¯ i ( k ) : j J i ( k ) I A ¯ i j ( k ) b i , A i ( k ) : j J i ( k ) I A ¯ i j ( k ) < b i , J i ( k ) > 0 .
Here, b i is the minimum required number of measurement paths that return a no-alarm decision in cycle k. For example, b i = M i act corresponds to a strict “all paths agree on normality” rule, while b i < M i act allows partial redundancy and multi-fidelity operation.
Measurement-level aggregation. In addition to decision-level fusion, measurement-level fusion is also allowed. In this case, we first form one cycle-level averaged estimate per path in cycle k:
Z ¯ i j ( k ) 1 m i j q = 1 m i j Z ¯ i j ( k , q ) .
In the special case m i j = 1 (used in the corresponding experiment series), Z ¯ i j ( k ) = Z ¯ i j ( k , 1 ) . The fused estimate and the resulting per-cycle channel decision are then:
Z ¯ i ( k ) = j J i ( k ) w i j Z ¯ i j ( k ) , j J i ( k ) w i j = 1 , w i j 0 , A ¯ i ( k ) : Z ¯ i ( k ) ω i , A i ( k ) : Z ¯ i ( k ) ω i .
Both approaches (decision-level and measurement-level) produce a binary output A ¯ / A and are evaluated comparably via unconditional P TN , P TP , P FA , and P MD using the method in Section 3.3.
In the reported experiment, measurement-level fusion is evaluated with fixed weights.
Stage 5. Temporal filtering at the channel level (K-cycle logic). To suppress single triggers and outliers, the channel decision is repeated over the decision window and the final channel decision is defined by a “ c i out of K i act ” rule, where K i act = K i :
A ¯ i : k K i I A ¯ i ( k ) c i , A i : k K i I A ¯ i ( k ) < c i .
Equivalently, an alarm A i is declared only after accumulating at least K i act c i + 1 alarm outcomes over the valid cycles in the decision window, which provides a formal mechanism for escalation and decision persistence.
In summary, the multi-stage logic (Stages 0–5) defines an unambiguous mapping from observations { Z i j } to the channel decision A ¯ i / A i for given Ω i , ω i j , and θ i . This full mapping serves as the architectural and notational basis of the framework used in the subsequent computational experiment.

3.5. Computational Experiment Plan

The computational experiment is organised as four Monte Carlo series. The reported series address baseline thresholding, repeated-measurement averaging, within-path confirmation, and measurement-level multi-fidelity fusion. In all reported series, solution quality is evaluated via unconditional CRIs P ^ TN , i , P ^ TP , i , P ^ FA , i , and P ^ MD , i (Section 3.3), together with the corresponding conditional risks α and β (Equations (6) and (7)) and the selected comparison criterion (e.g., J err , raw ).
Series 1 (baseline threshold logic): (i) calibrate the baseline threshold formulation for the prescribed abnormal-state prior π dev (normal-region bound set via π dev z 1 ); (ii) select the optimal decision bound z 2 * for fixed measurement-path accuracy; and (iii) assess sensitivity to s and the P FA P MD trade-off as functions of z 2 .
Series 2 (within-cycle averaging): quantify the effect of averaging repeated measurements (representative point n = 5 ) using the fixed thresholds obtained in Series 1.
Series 3 (within-path confirmation): assess the effect of “a out of m” confirmation over repeated cycles. Here, the setting ( a , m ) = ( 2 , 3 ) is used as an illustrative reference case for comparison rather than as a universally optimal confirmation policy. For shipboard application, confirmation parameters should be linked to the sampling interval, the confirmation window, and the expected time scale of fault evolution.
Series 4 (multi-fidelity fusion): assess the effect of two-path multi-fidelity fusion using a measurement-level combination with a fixed fusion weight. In this series, the weight w is optimised within a controlled setting and treated as a constant across the operating profile in order to isolate the contribution of multi-fidelity fusion. This formulation is intended for comparative analysis rather than as a final weighting policy for practical shipboard use.
The series parameters, fixed settings of θ i (Equation (18)), and the controlled outputs are summarized in Table 3.
The reported series should be interpreted as control-point comparisons within a unified verification framework rather than as a full optimisation over the decision-logic design space. Accordingly, the reported gains quantify the behaviour of selected logic elements at representative points, not universal optimal settings for shipboard deployment.

4. Results

The results reported below correspond to the four Monte Carlo series introduced in Section 3.5. In all reported series, we used the same baseline setup (Section 3): X and Y were normally distributed; the regions were symmetric, Ω = [ z 1 σ X , z 1 σ X ] and ω = [ Z L , Z U ] with Z L = z 2 σ X and Z U = z 2 σ X ; and σ X = 1 with a fixed rarity parameter π dev = 0.001 . This stylized Z = X + Y uncertainty model is used as a surrogate for ML-derived diagnostic statistics in PdM: X represents the latent health-related component of the statistic, whereas Y aggregates sensor and model-induced uncertainty that drives spurious threshold crossings and missed detections in practice. All results reported below are based on N = 2 , 000 , 000 Monte Carlo trials, which yields small estimation error even for probabilities on the order of 10 4 (for p 10 4 , the standard error is on the order of 10 5 ). Therefore, between-series differences at the level of 10 4 10 3 are interpretable.

4.1. Series 1: Threshold Calibration and Decision-Region Selection (Baseline Single-Cycle Channel)

Calibration using a prescribed abnormal-state prior was performed by setting the scenario parameter π dev = 0.001 , which in this simulation corresponds to z 1 = 3.290527 and an empirical fraction of normal realizations close to the target 1 π dev = 0.999 . This parameter is used to represent an extreme imbalance regime for controlled comparison of decision-logic variants, rather than as a direct engineering estimate of the actual failure rate of marine diesel engines.
Next, for a fixed noise ratio s = σ Y / σ X = 0.3 , we selected the decision-region width parameter z 2 by minimizing the raw error criterion J err , raw . In this simulation, we used
J err , raw = P FA + 10 P MD ,
which prioritizes missed detections relative to false alarms.
A coarse grid search followed by local refinement yielded z 2 ( coarse ) = 3.20 and z 2 * = 3.23 as the minimizer of J err , raw .
The key interpretation of z 2 * is as follows. When z 2 < z 1 , the no-alarm region ω is narrower than the true normal region Ω , so the channel becomes more conservative (it raises an alarm earlier). This reflects the adopted weighting in J err , raw : narrowing ω reduces P MD and β but increases P FA and α . The optimum z 2 * fixes the resulting trade-off for the chosen weights. The coarse search and local refinement of z 2 * are shown in Figure 1 and Figure 2, and the P FA P MD trade-off as a function of z 2 is shown in Figure 3.
For the selected baseline point (Series 1, s = 0.3 , z 1 = 3.290527 , z 2 * = 3.23 ), we obtained P FA = 0.0011845 , P MD = 0.0001850 , α = 0.001186 , β = 0.181997 , and J err , raw = 0.0030345 . Under extremely rare abnormalities ( 0.1 % ), the unconditional missed-detection probability P MD remains small (on the order of 10 4 ); however, the conditional missed-detection risk given an abnormal state, β , is substantial (about 0.182 ) in the baseline single-cycle scheme.

4.2. Series 1 (Continued): Sensitivity of Conditional Risks to the Noise Ratio

With fixed z 1 and z 2 * , we evaluated a grid of s { 0.1 , , 0.8 } . Both conditional risks increase as relative measurement accuracy deteriorates, with a markedly stronger increase in β than in α . At s = 0.1 , we obtained α 3.72 × 10 4 and β 0.0526 , whereas at s = 0.8 we obtained α 0.0109 and β 0.3424 . Thus, as measurement-error variance increases relative to the variability of X, the channel rapidly loses sensitivity to rare abnormalities (i.e., β increases). Methodologically, this supports the need for multi-stage decision logic (repeats, confirmation, and multi-fidelity), rather than reliance on a single threshold trigger. The dependencies α ( s ) and β ( s ) are shown in Figure 4.

4.3. Series 2: Effect of Averaging Repeated Measurements

In Series 2, for the reference configuration n = 5 , we obtained P FA = 0.0004450 , P MD = 7.75 × 10 5 , α = 0.000445 , β = 0.076242 , and J err , raw = 0.0012200 .
Relative to the baseline Series 1, this corresponds to:
  • A 62.43 % reduction in α (ratio 0.376 );
  • A 58.11 % reduction in β (ratio 0.419 );
  • A 59.80 % reduction in J err , raw (ratio 0.402 ).
From an engineering perspective, averaging five independent measurements substantially increases the robustness of the threshold decision against the random component Y and provides the largest gain among the methods evaluated at this point. In particular, it reduces both the false-alarm rate and the probability of missing a rare abnormality.

4.4. Series 3: Effect of “2 out of 3” Confirmation

Series 3 evaluates within-path confirmation at the illustrative reference case ( a , m ) = ( 2 , 3 ) under the same baseline conditions used for Series 1. This setting is introduced here as a representative comparison case rather than as a universally optimal confirmation rule because a principled choice of a and m should depend on the sampling interval, the length of the confirmation window, and the expected time scale of fault evolution in the monitored object. At this point, we obtained P FA = 0.0006885 , P MD = 1.435 × 10 4 , α = 0.000689 , β = 0.141171 , and J err , raw = 0.0021235 .
Relative to the baseline Series 1:
  • α decreased by 41.89% (ratio 0.581);
  • β decreased by 22.43% (ratio 0.776);
  • J err , raw decreased by 30.02% (ratio 0.700).
Confirmation based on binary outcomes reduces both types of errors (false alarms and missed detections), but in this setup it is less effective than averaging (Series 2). This is expected because Series 3 operates after information loss due to binarisation (decisions A / A ¯ rather than a continuous Z).

4.5. Series 4: Multi-Fidelity Two-Path Fusion and Weight Optimization

Series 4 considers a controlled two-path multi-fidelity setting in which measurement-level fusion is implemented with a fixed weight across the operating profile. This fixed-weight formulation is used here as a deliberate simplification for comparing fusion against other decision-logic mechanisms, rather than as a general deployment policy. Within this setting, the optimisation yielded w * = 0.83 . At this point, we obtained P FA = 0.0010615 , P MD = 0.0001840 , α = 0.001063 , β = 0.181013 , and J err , raw = 0.0029015 .
The engineering reference weight w 0 = 0.70 yielded J err , raw = 0.0031825 , which is slightly worse than the baseline Series 1 ( J err , raw = 0.0030345 ). Moving from w 0 to w * improves J err , raw by approximately 8.83 % within Series 4. The objective landscape and the selected optimum w * are shown in Figure 5.

4.6. Summary Comparison of Series (Control Point) and Implications for Multi-Stage Synthesis

To ensure a fair comparison, all variants are evaluated under the same baseline configuration with identical z 1 and z 2 * ; therefore, differences are driven only by the logic component (averaging, confirmation, or multi-fidelity fusion). The results are summarized in Table 4 and Figure 6. Here, n norm and n dev denote the Monte Carlo counts of realizations with X Ω and X Ω ¯ , respectively (identical across Series 1–4 because F ( X ) is fixed).
From the summary plot and Table 4, the following points can be concluded:
1.
The largest gain under the shared baseline configuration is achieved by averaging (Series 2): α decreases by a factor of ≈2.66, β by a factor of ≈2.39, and J err , raw by a factor of ≈2.49 relative to the baseline threshold decision.
2.
“2 out of 3” confirmation (Series 3) yields a consistent improvement, but weaker than averaging: J err , raw decreases by about 30%, and β by about 22% relative to Series 1.
3.
Multi-fidelity fusion (Series 4) with s i 1 = 0.3 and s i 2 = 0.7 provides only a small improvement over Series 1 when the weight is optimised ( w * = 0.83 ): J err , raw decreases by ≈4.4%, while β remains nearly unchanged. Without optimisation (the reference point w 0 = 0.70 ), performance degrades slightly relative to Series 1.
Under the shared baseline configuration, averaging reduced both α and β more strongly than “2 out of 3” confirmation, whereas two-path fusion provided only a limited gain after calibration of w. In other words, the compared mechanisms shift the P FA P MD trade-off in quantitatively different ways even under the same baseline threshold z 2 * .

5. Discussion

5.1. Interpretation of the Results

We evaluate decision-logic quality not by “classifier accuracy” but by probabilistic risks of false alarms and missed detections, defined through the unconditional control-reliability indicators P TN , P TP , P FA , and P MD , and the corresponding conditional risks.
Under extremely rare abnormal states, π dev = 0.001 (i.e., P ( Ω ¯ ) 10 3 ), the unconditional missed-detection probability P MD is inevitably small in magnitude. Here, π dev = 0.001 is treated as a scenario parameter of the simulation. Within rare-abnormality settings of this kind, however, the conditional risk β becomes especially informative, because it quantifies the probability of failing to detect an abnormality given that it is truly present. This is observed at the Series 1 reference setting: for s = σ Y / σ X = 0.3 and the error-optimal threshold z 2 * = 3.23 , we obtained β 0.182 . Thus, in the single-cycle threshold scheme, roughly one out of five to six abnormal cases does not trigger an alarm, indicating limited sensitivity to rare events at the given measurement-path accuracy.
Series 2 (within-path averaging of repeated measurements, n = 5 ) provided the largest gain at the reference setting among the evaluated mechanisms: both α and β decreased simultaneously (in particular, α decreased by approximately a factor of 2.66 and β by approximately a factor of 2.39 compared to the baseline Series 1). At the common comparison point, the n = 5 averaging scheme reduced α from 0.001186 to 0.000445 and β from 0.182 to 0.076 . This indicates that aggregation of the continuous quantity Z prior to binarisation suppresses the random component of the error Y more effectively than the other mechanisms compared at this setting.
Series 3 (confirmation over repeated cycles using a “2 out of 3” rule) also improves the indicators consistently (both α and β decrease relative to Series 1), but it is less effective than averaging. A plausible explanation is information loss due to binarisation: Series 3 operates on discrete outcomes A / A ¯ , whereas Series 2 aggregates continuous values Z before applying the threshold decision. At the reported control point, mechanisms operating before binarisation appear more effective than those applied after binarisation because they retain information carried by the continuous statistic Z. This observation should be interpreted as a result of the adopted surrogate setting and comparison protocol; confirming its generality requires broader parameter sweeps and validation on operational data.
Series 4 (multi-fidelity, two paths with s i 1 = 0.3 and s i 2 = 0.7 , fused at the measurement level) reveals an important effect: adding a coarse measurement path does not guarantee improvement unless fusion parameters are tuned appropriately. Weight optimisation in the measurement-level fusion Z ¯ i ( k ) = w Z ¯ i 1 ( k ) + ( 1 w ) Z ¯ i 2 ( k ) (Equation (26); in Series 4, n i j = m i j = 1 so Z ¯ i j ( k ) Z i j ( k , 1 , 1 ) ) yielded w * = 0.83 , implying dominance of the more accurate path; with the engineering reference weight w 0 = 0.70 , performance slightly degraded relative to the baseline Series 1, whereas the optimised weight yielded a small but consistent improvement in J err , raw . Thus, multi-fidelity design should be treated not as an automatic performance gain, but as a mechanism whose weights or voting rules require calibration.

5.2. Comparison with the Literature

Compared with the literature, the distinctive feature of the present study is not the use of thresholds, confirmation, or fusion as such, but their treatment as explicit, calibratable components of a channel-level decision logic evaluated under a common CRI-based protocol. In many applied marine studies, these elements are embedded implicitly in a model pipeline or heuristic warning scheme; here, they are isolated as design variables and compared on the same probabilistic basis.
Many applied studies on marine CBM/PdM specify the AD–FD structure and discuss deployment constraints, but the decision layer is often reduced to a single threshold on an anomaly score or residual without an explicit link to controlled risks such as α and β . In the reported baseline configuration ( z 2 * = 3.23 , s = 0.3 ), this distinction is visible: although unconditional P MD is small, the conditional missed-detection risk remains β 0.182 . Rule-based and zone-based approaches are practically interpretable because they connect diagnostics to maintenance actions; however, zone boundaries are most often defined heuristically and are rarely quantified via error probabilities [56]. Residual-based approaches also motivate confirmation over observation sequences and explicit use of alarm history [57]. In the present formulation, the decision region ω , confirmation parameters, and fusion rule are treated as explicit components of channel logic and compared through the same set of CRIs.

5.3. Study Limitations and Threats to Validity

1.
Distributional assumptions. The computational experiment uses normal distributions for X and Y and symmetric regions Ω and ω . Real operational data may exhibit asymmetry, heavy tails, outliers, and nonstationarity. Therefore, the reported quantitative gains (in percent) should be interpreted within the adopted uncertainty model; for real data, the robustness of the conclusions should be verified under more realistic distributions for F ( X ) and F ( Y ) .
2.
Interpretation of the rarity parameter. The value π dev = 0.001 is introduced in this paper as a scenario parameter for studying the decision logic under extreme class imbalance. It is not calibrated here against fleet-level maintenance data or failure statistics for marine diesel engines. Therefore, conclusions related to rarity should be interpreted within the adopted simulation setting, whereas deployment-oriented calibration would require maintenance logs, fleet statistics, or domain-specific abnormality priors.
3.
Independence of repeats. The averaging effect (Series 2) assumes that the noise component across repeated measurements is largely uncorrelated. Under strong autocorrelation (sensor drift, slow operating-condition fluctuations, or process inertia), the effective variance reduction will be smaller. In such cases, it may be necessary to increase the interval between measurements, use robust aggregation, or combine averaging with confirmation over cycles.
4.
Operating-condition heterogeneity. In real operation, the distribution of X is a mixture over operating conditions (load, speed, and thermal conditions), and the error Y may also depend on operating conditions. Consequently, transferring thresholds and confirmation parameters requires operating-condition normalisation (or a residual-based formulation relative to expected normal behaviour under a given operating condition) and calibration of ω as a function of operating conditions.
5.
Confirmation-parameter selection. In this study, the setting ( a , m ) = ( 2 ,   3 ) is used as an illustrative reference case for comparing confirmation against other decision-logic mechanisms. This choice is not sufficient to support a general design recommendation. For shipboard application, confirmation parameters should be linked to the sampling interval, the observation window, and the characteristic time scale of abnormality development. Generalisable recommendations therefore require either a parameter sweep over ( a , m ) or calibration against operational data and domain-specific dynamics.
6.
Multi-fidelity monitoring and correlated path errors. Series 4 considers a specific scenario with two measurement paths and linear measurement-level fusion using a fixed weight across the operating profile. In practice, however, path reliability may vary with operating conditions, and errors across measurement paths may be correlated because of shared disturbance sources, common calibration drift, or synchronization effects. Under such conditions, a single global weight may be insufficient for deployment-oriented decision logic. More general recommendations therefore require parameterization over path accuracies, correlation structure, and regime-dependent weighting policies.
7.
Monte Carlo accuracy. With N = 2 × 10 6 , the estimation accuracy is sufficient for probabilities on the order of 10 4 10 3 and for comparing series under the shared baseline configuration. For more stringent target levels (e.g., P MD 10 6 ), either larger N or variance-reduction techniques would be required, together with checks of reproducibility and stability of random number generation.

5.4. Practical Integration of CBM Logic into CMMS/EAM and Requirements for Measurement and Computation

For shipboard implementation, the Stages 0–5 logic may be mapped onto a CMMS/EAM workflow by logging a minimal, traceable set of decision attributes for each monitored channel (e.g., the admitted-path flags g i j ( k ) , the effective window size | K i | , and the outcomes of within-path confirmation and K i -cycle filtering). This makes each alarm/no-alarm outcome A ¯ i / A i auditable in terms of which paths participated and which rule confirmed (or rejected) the event.
A maintenance trigger can then be defined in terms of a confirmed event (e.g., after “a out of m” confirmation and/or “c out of K” filtering) rather than a single threshold crossing. In this study, mechanisms that act before binarisation (e.g., averaging in Series 2) reduced both α and β more strongly than confirmation alone. Whether the same prioritisation holds in deployment depends on the sampling interval, operating-profile variability, and data quality, and should be verified on operational data.
One plausible architecture is to run data-quality gating, computation of diagnostic statistics, and confirmation/filtering onboard (edge), while performing periodic recalibration of thresholds and fusion weights offline using accumulated data and maintenance logs. The feasibility of this split depends on the available computing budget and connectivity constraints.
Computational budget and scalability should also be considered when the proposed CRI-based calibration is used in practice. In the supplementary implementation, a full Monte Carlo run for Series 1–4 with N = 2 × 10 6 required about 5.8 s of wall-clock time on a 12th Gen Intel Core i5-12500H system with 16 GB RAM (Python 3.13.13, NumPy 2.3.3). Peak memory usage remained below 1 GB (about 0.34 GiB working set and 0.82 GiB private memory). Because the threshold and weight sweeps use fixed grids, the computational cost is governed mainly by the Monte Carlo sample size and scales approximately linearly with N. This supports periodic recalibration as an offline engineering task after operating-profile updates; onboard operation can then use fixed pre-calibrated parameters in real time.

5.5. Directions for Future Research

Further work should extend the control-point comparisons reported here (e.g., n = 5 , ( a , m ) = ( 2 ,   3 ) , and two-path fusion with w * = 0.83 ) to broader parameter grids over ( s , π dev ) and ( n i j , m i j , a i , b i , K i , c i ) and, where relevant, over the number of active paths M i . Such computations would enable linking design choices to sampling intervals, confirmation-window duration, expected fault-evolution time scales, and time/cost constraints. A related step is to replace the scenario value of π dev used here with priors derived from maintenance logs, fleet statistics, or domain-specific estimates of abnormality occurrence.
The uncertainty model should also be extended beyond the Gaussian stationary baseline to include outliers, drift, autocorrelation, mixtures of operating conditions in both F ( X ) and F ( Y ) , missing data, correlated path errors, asynchronous sampling, and dynamically changing sets of active paths.
In addition, the framework should be validated on long-term operational data from physical marine diesel engines. This includes formation of the observable diagnostic quantity Z, operating-condition-dependent specification of the decision region ω , calibration of confirmation parameters through CRIs using maintenance logs and expert-verified events, and evaluation of adaptive fusion policies instead of the fixed-weight scheme used in Series 4.

6. Conclusions

This study establishes a CRI-based framework for synthesising and comparing multi-stage CBM decision logic for marine diesel monitoring under rare abnormalities, measurement uncertainty, and heterogeneous channel quality. The key result is methodological: the paper treats implementable decision logic—not a single threshold or classifier output—as the primary design object and evaluates its variants through a unified Monte Carlo verification protocol.
Using a Monte Carlo computational experiment ( N = 2 × 10 6 ) with π dev = 0.001 , we compared a baseline single-cycle threshold scheme with three reliability-enhancement mechanisms: averaging of repeated measurements, “a out of m” confirmation over repeated cycles, and two-path multi-fidelity measurement-level fusion. The baseline risks were α = 0.001186 , β 0.182 , and J err , raw = 0.00303 . Averaging (Series 2, n = 5 ) reduced them to α = 0.000445 , β = 0.0762 , and J err , raw = 0.00122 , whereas confirmation (Series 3, ( a , m ) = ( 2 ,   3 ) ) yielded α = 0.000689 , β = 0.141 , and J err , raw = 0.00212 . Multi-fidelity monitoring (Series 4) improved performance only when fusion parameters were tuned appropriately: the optimal weight w * shifted toward the more accurate path, whereas non-optimal engineering weights yielded little or no gain under the adopted error criterion.
The practical significance of the framework lies in its compatibility with AD–FD architectures in which AI/ML modules generate the diagnostic statistic Z, while a downstream decision layer transforms this statistic into traceable maintenance triggers. In the present paper, this integration is demonstrated at the level of a Gaussian surrogate verification protocol; extending the framework to implementation-specific, operating-condition-dependent shipboard data remains the next necessary step.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/eng7050190/s1.

Author Contributions

Conceptualization, D.T. and O.A.; methodology, D.T., O.A. and A.K.; software, A.K.; validation, D.T., O.A. and A.K.; formal analysis, A.K.; investigation, D.T., O.A. and A.K.; resources, A.K.; data curation, D.T. and A.K.; writing—original draft preparation, D.T. and A.K.; writing—review and editing, A.K.; visualization, O.A. and A.K.; supervision, O.A.; project administration, D.T.; funding acquisition, D.T. and O.A. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The code and data supporting the findings of this study are provided in the Supplementary Materials.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
ADanomaly detection
ANNartificial neural network
CBMcondition-based maintenance
CMMScomputerized maintenance management system
CRIcontrol-reliability indicator
DQNdeep Q-network
E2Eend-to-end
EAMenterprise asset management
EGTexhaust gas temperature
FDfault diagnosis
KPCAkernel principal component analysis
MDmissed detection
OSELMonline sequence extreme learning machine
PdMpredictive maintenance
RCMreliability-centered maintenance
RMSroot mean square

Appendix A. Symbols and Notation

Appendix A.1. Indices, Cycles, and Sets

Table A1. Indices, cycles, and sets used in the multi-stage condition-based maintenance (CBM) logic.
Table A1. Indices, cycles, and sets used in the multi-stage condition-based maintenance (CBM) logic.
SymbolMeaning
iMonitoring channel (diagnostic variable) index; i = 1 , , n .
jMeasurement-path index in channel i; j = 1 , , M i .
kUpper-level channel cycle index within a decision window (Stages 0–4 yield A ¯ i ( k ) / A i ( k ) ).
qSub-cycle index used for within-path confirmation; q = 1 , , m i j .
lRepeated measurement index within sub-cycle q; l = 1 , , n i j .
tMonte Carlo trial index; t = 1 , , N .
J i ( k ) Active paths in cycle k: J i ( k ) = { j : g i j ( k ) = 1 } .
M i act Number of active paths in cycle k: M i act = | J i ( k ) | .
K i Valid cycles in the window: K i = { k : | J i ( k ) | > 0 } .
K i act Number of valid cycles in the window: K i act = | K i | .
K i , min Minimum admissible number of valid cycles required to issue a channel decision (policy parameter).

Appendix A.2. Random Variables, Observation Model, and Regions

Table A2. Observation model, uncertainty, and state/decision regions.
Table A2. Observation model, uncertainty, and state/decision regions.
SymbolMeaning
X i True diagnostic variable for channel i (physical parameter or computed diagnostic indicator).
Y i j Aggregate error/uncertainty of measurement path j in channel i.
Z i j Observable diagnostic quantity; base model Z i j = X i + Y i j .
Z i j ( k , q , l ) Repeated observation: Z i j ( k , q , l ) = X i + Y i j ( k , q , l ) .
Z ¯ i j ( k , q ) Within-sub-cycle average over n i j repeats: Z ¯ i j ( k , q ) = 1 n i j l = 1 n i j Z i j ( k , q , l ) .
Z ¯ i j ( k ) Cycle-level average over m i j sub-cycles (when used): Z ¯ i j ( k ) = 1 m i j q = 1 m i j Z ¯ i j ( k , q ) .
F i ( X ) Unconditional distribution of X i (mixture over operating conditions).
F i j ( Y ) Distribution of Y i j (may depend on operating conditions).
σ X Standard deviation of X in the adopted uncertainty model.
σ Y , i j Standard deviation of Y i j .
sNoise ratio (single-path case): s = σ Y / σ X .
s i j Path-specific noise ratio: s i j = σ Y , i j / σ X .
ε Standard normal variate used in Monte Carlo schemes: ε N ( 0 , 1 ) .
Ω i , Ω ¯ i True normal and abnormal regions for X i (complementary).
X L , X U Lower/upper bounds of the true normal region (threshold form).
ω i j , ω i Path-level and fused no-alarm decision regions in Z.
Z L , Z U Lower/upper bounds of a decision region in Z.
π dev Abnormality rarity: π dev = P ( Ω ¯ ) .
z 1 , z 2 Standardized parameters defining symmetric Ω and ω in the computational experiment.

Appendix A.3. Decisions and Multi-Stage Logic Parameters

Table A3. Binary decisions and structural design parameters of the multi-stage logic.
Table A3. Binary decisions and structural design parameters of the multi-stage logic.
SymbolMeaning
A ¯ , A Binary decisions: no alarm ( A ¯ ) and alarm (A).
A ¯ i j ( k , q ) / A i j ( k , q ) Primary path decision in cycle k, sub-cycle q (based on Z ¯ i j ( k , q ) ω i j ).
A ¯ i j ( k ) / A i j ( k ) Confirmed path decision for cycle k after within-path confirmation.
A ¯ i ( k ) / A i ( k ) Per-cycle channel decision in cycle k after aggregation across active paths.
A ¯ i / A i Final channel decision after temporal filtering over the decision window.
g i j ( k ) Data-quality flag for path j in cycle k; g i j ( k ) = 1 admits the path for voting/fusion.
I ( · ) Event indicator function used in counting rules.
n i j Number of repeated measurements within one sub-cycle (averaging, noise suppression).
m i j Number of sub-cycles per upper-level cycle for within-path confirmation.
M i Number of measurement paths in channel i (multi-fidelity/redundancy).
a i Within-path confirmation threshold: minimum number of no-alarm outcomes out of m i j .
b i Path-aggregation threshold: minimum number of no-alarm paths out of M i act .
K i Number of upper-level cycles in the decision window (temporal filtering).
c i Temporal-filter threshold: minimum number of no-alarm outcomes out of K i act .
θ i Channel design vector: θ i = ( { n i j , m i j } j = 1 M i , M i , a i , b i , c i , K i ) .
w i j Nonnegative fusion weights for measurement-level aggregation; j J i ( k ) w i j = 1 .
w , w * , w 0 Two-path fusion weight, its optimised value, and the engineering reference value.

Appendix A.4. CRIs, Conditional Risks, Monte Carlo Counters, and Criteria

Table A4. Quality indicators (control-reliability indicators (CRIs)), conditional risks, Monte Carlo counters, and comparison criteria.
Table A4. Quality indicators (control-reliability indicators (CRIs)), conditional risks, Monte Carlo counters, and comparison criteria.
SymbolMeaning
P TN , i True negative probability: P TN , i = P ( Ω i A ¯ i ) .
P TP , i True positive probability: P TP , i = P ( Ω ¯ i A i ) .
P FA , i False alarm probability: P FA , i = P ( Ω i A i ) .
P MD , i Missed detection probability: P MD , i = P ( Ω ¯ i A ¯ i ) .
α Conditional false-alarm risk: α = P ( A Ω ) = P FA P TN + P FA .
β Conditional missed-detection risk: β = P ( A ¯ Ω ¯ ) = P MD P TP + P MD .
NNumber of Monte Carlo trials.
n hit Number of hits of a given event in N trials (frequency estimator).
P ^ · Monte Carlo estimate of the corresponding unconditional CRI (frequency-based estimator of the population indicator).
n TN , n TP , n FA , n MD Monte Carlo counters of TN/TP/FA/MD outcomes.
n norm , n dev Counts of true normal/abnormal realizations: n norm = n TN + n FA , n dev = n TP + n MD .
J int Integral index (weighted sum of normalised CRIs).
J err Simplified index based on normalised errors only.
J err , raw Raw error criterion used in tuning (reported run): J err , raw = P FA + 10 P MD .
γ 1 , , γ 4 Expert-defined weights in J int and/or J err .
k ˜ Linearly normalised indicator value defined over a practical engineering range.
k min , k max Practical normalisation bounds (set relative to baseline/requirements).

Appendix B. Computational Experiment: Description and Implementation

Appendix B.1. General Settings and Adopted Parameter Values (Common to All Series)

1.
Distributions and dimensionless parametrization. The current implementation uses a dimensionless formulation with X 0 = 0 , σ X = 1.0 , and s = σ Y / σ X . Therefore, σ Y = s σ X = s . Observations follow Equation (1), Z = X + Y , where Y N ( 0 , σ Y 2 ) .
2.
Calibration of the normal region Ω via a prescribed rarity parameter. In the reported simulation, π dev = 0.001 is used as a scenario parameter representing an extreme class-imbalance setting in the computational experiment:
P | X | > z 1 σ X = π dev .
Hence,
z 1 = Φ 1 1 π dev 2 ,
where Φ 1 ( · ) is the quantile function of the standard normal distribution. For the scenario value π dev = 0.001 , we obtained z 1 = 3.290527 . This value is used only to define a controlled rare-abnormality setting for the Monte Carlo analysis. This quantity should not be interpreted here as a direct estimate of the fleet-level failure rate of marine diesel engines; rather, it is introduced to study the behaviour of the decision logic under rare-abnormality conditions.
3.
Symmetric region bounds (as implemented in the code). Given X 0 = 0 and symmetry, we used
Ω = [ z 1 σ X , + z 1 σ X ] , ω = [ z 2 σ X , + z 2 σ X ] .
In the reported simulation, σ X = 1 , so the bounds are ± z 1 and ± z 2 , respectively.
4.
Number of trials and reproducibility. All series use N = 2 , 000 , 000 Monte Carlo trials. For reproducibility, the pseudo-random number generator seed is fixed (see Appendix B.8).
5.
Parameter-selection criterion (for the optimisation steps in Series 1 and 4). To select z 2 * and w * , we use the raw (non-normalised) error criterion:
J err , raw = γ MD P MD + γ FA P FA ,
where, in the reported simulation, γ MD = 10 and γ FA = 1 . Conceptually, this corresponds to the simplified index in Equation (17), but without the normalisation step in Equations (9) and (10), because optimisation is performed directly over the error probabilities.

Appendix B.2. General Monte Carlo Run Scheme (Common to All Series)

In all series, estimation follows the definitions in Section 3.3: the true state is determined by whether X Ω , and the decision is determined by whether Z ω . Outcomes are classified into TN/FA/TP/MD according to Table 2; unconditional probabilities are estimated by Equation (5), and conditional risks α and β are computed using Equations (6) and (7).

Appendix B.3. Universal Scheme for a Monte Carlo Run at Fixed Series Parameters

For fixed series parameters, the simulation procedure is as follows:
1.
Specify N, σ X , and π dev z 1 ; set measurement-accuracy parameters (a single value s or a set of values); specify the series logic parameters (e.g., n, ( a , m ) , or ( s 1 , s 2 , w ) ); and set the threshold z 2 (or a procedure for selecting z 2 * ).
2.
Generate a sample X ( t ) , t = 1 , , N , from N ( 0 , σ X 2 ) and mark the true state via X ( t ) Ω .
3.
Form observations Z ( t ) (or Z ¯ ( t ) ) according to the series rules (see Appendix B.4, Appendix B.5, Appendix B.6 and Appendix B.7).
4.
Make the no-alarm/alarm decision based on membership in ω (in Series 2–4, the decision is made using Z ¯ or a confirmed rule).
5.
Accumulate counters n TN , n FA , n TP , and n MD , then compute P ^ TN , P ^ FA , P ^ TP , P ^ MD , and the conditional risks α and β . For optimisation steps, also compute J err , raw .
A key point for comparability is that all single-point calculations in Series 1–4 use the same sample of X. This guarantees identical n norm and n dev and enables direct comparison of α , β , and J err , raw across series without differences caused by a randomly generated operating profile of X.

Appendix B.4. Series 1: Baseline Threshold Decision and Parameter Selection

The logic parameters are M i = 1 , n i 1 = 1 , m i 1 = 1 , and K i = 1 . In each trial, one observation is generated, Z i 1 ( t ) = X ( t ) + Y i 1 ( t ) , and the no-alarm decision is made when Z i 1 ( t ) ω i 1 , i.e., | Z i 1 ( t ) | z 2 σ X (Equation (3)).
In the reported simulation, Series 1 consists of three technically distinct steps.
Step 1: Selecting z 2 * at fixed z 1 and s = s fixed . We fix σ X = 1 and π dev = 0.001 z 1 = 3.290527 ; s fixed = 0.3 σ Y = 0.3 ; and N = 2 × 10 6 . To reduce Monte Carlo noise in threshold comparison, the same realizations of X and Y are used (i.e., one array Z for the entire sweep over z 2 ); only the check | Z i 1 | z 2 σ X changes (with σ X = 1 in the reported run, this is equivalent to | Z i 1 | z 2 ). A coarse sweep uses z 2 [ 2.8 , 4.6 ] with step 0.05 . For each z 2 , we compute P ^ FA , P ^ MD , and J err , raw and select the minimizer z 2 ( coarse ) . Local refinement then builds a grid of width ± 0.20 around z 2 ( coarse ) with step 0.01 ; the minimizer of J err , raw on this grid is selected as z 2 * . In the reported simulation, the coarse sweep yielded z 2 ( coarse ) = 3.20 , and the local refinement yielded z 2 * = 3.23 .
Step 2: Sensitivity to measurement-path accuracy s at fixed z 1 and z 2 = z 2 * . We fix z 1 and z 2 = z 2 * and consider s { 0.1 , 0.2 , 0.3 , 0.4 , 0.5 , 0.6 , 0.7 , 0.8 } . To improve smoothness across s, we use common random numbers: a single array of standard normal noise ε ( t ) N ( 0 , 1 ) is generated, and the measurement-path error is set as Y i 1 ( t ) = σ Y , i 1 ε ( t ) = ( s σ X ) ε ( t ) . This eliminates curve “jitter” caused by different random seeds across s and isolates the effect of changing the variance. For each s, we estimate α , β , and J err , raw .
Step 3: Control illustration of the z 2 trade-off at fixed z 1 and s = s fixed . For interpretability, we additionally evaluate z 2 { 3.0 , 3.5 , 4.0 , 4.5 } at s fixed = 0.3 (using the same X and Y as in Step 1). The output comprises the values P FA ( z 2 ) , P MD ( z 2 ) , and J err , raw ( z 2 ) to illustrate the trade-off.

Appendix B.5. Series 2: Repeated Measurements with Averaging

The logic parameters are M i = 1 , m i 1 = 1 , K i = 1 , and n i 1 = n . In the reported simulation, we use n i 1 = 5 , while z 1 and z 2 are fixed as ( z 1 , z 2 * ) from Series 1, and the measurement-path accuracy is fixed as s = s fixed = 0.3 .
Conceptually, this series corresponds to Equations (21)–(23): within one trial, n independent measurements are performed, followed by averaging and a threshold check against ω .
In implementation, an equivalent accelerated form is used. For independent Gaussian noise Y ( l ) N ( 0 , σ Y 2 ) , the sample mean Y ¯ = 1 n l = 1 n Y ( l ) follows N ( 0 , σ Y 2 / n ) . Therefore, Z ¯ is simulated directly as
Z ¯ i 1 ( t ) = X ( t ) + σ Y , i 1 n i 1 ε ( t ) , σ Y , i 1 = s σ X , ε ( t ) N ( 0 , 1 ) .
which is fully equivalent to explicitly generating n noise samples and averaging, but is substantially faster for large N.
After Z ¯ is formed, the decision is made using the same region ω (Equation (3)): no alarm if Z ¯ i 1 ( t ) ω i 1 , i.e., | Z ¯ i 1 ( t ) | z 2 * σ X . We then estimate P ^ , α , β , and J err , raw .

Appendix B.6. Series 3: “a out of m” Confirmation (Control Point 2 out of 3)

The logic parameters are M i = 1 , n i 1 = 1 , m i 1 = m , a i = a , and K i = 1 . In the reported simulation, m i 1 = 3 and a i = 2 are used; the thresholds and accuracy are fixed as ( z 1 , z 2 * , s fixed = 0.3 ) .
Within one trial t:
1.
Generate m = 3 independent noise samples Y i 1 ( t , q ) and observations Z i 1 ( t , q ) = X ( t ) + Y i 1 ( t , q ) (a special case of Equation (21) with n = 1 );
2.
For each cycle q, form the primary decision no alarm if | Z i 1 ( t , q ) | z 2 * σ X (Equation (23) with n = 1 );
3.
Form the confirmed decision using the “a out of m” rule (Equation (24)): no alarm if the number of cycles that produced no alarm is at least a = 2 .
We then estimate P ^ , α , β , and J err , raw and compare with the baseline Series 1 using the same sample of X.

Appendix B.7. Series 4: Multi-Fidelity Two-Path Measurement-Level Fusion

The logic parameters are M i = 2 , and fusion is implemented as measurement-level fusion (Equation (26)) in the special case of two paths and a single cycle ( n i 1 = n i 2 = 1 , m i 1 = m i 2 = 1 , K i = 1 ). In the reported simulation, s i 1 = 0.3 , s i 2 = 0.7 , σ Y 1 = s i 1 σ X , and σ Y 2 = s i 2 σ X . The thresholds ( z 1 , z 2 ) are fixed as ( z 1 , z 2 * ) from Series 1.
Within one trial t:
1.
Form two independent observations of the same X ( t ) :
Z i 1 ( t ) = X ( t ) + Y i 1 ( t ) , Z i 2 ( t ) = X ( t ) + Y i 2 ( t ) ,
where Y i 1 ( t ) N ( 0 , σ Y 1 2 ) and Y i 2 ( t ) N ( 0 , σ Y 2 2 ) are independent;
2.
Perform measurement fusion:
Z ¯ i ( t ) = w Z i 1 ( t ) + ( 1 w ) Z i 2 ( t ) , w [ 0 , 1 ] ;
3.
Apply the same threshold decision (Equation (3)): no alarm if | Z ¯ i ( t ) | z 2 * σ X .
The optimal weight w * is selected by a grid search:
1.
w [ 0 , 1 ] with step 0.01 ;
2.
For each w, compute J err , raw ( w ) = P FA ( w ) + 10 P MD ( w ) (cf. Equation (28));
3.
Select the minimizer as w * .
In addition, we evaluate the engineering reference point w 0 = 0.7 (without optimisation) to demonstrate the sensitivity of solution quality to weight selection.

Appendix B.8. Reproducibility, Monte Carlo Noise Control, and Statistical Error Assessment

1.
Fixed seeds and consistency of comparisons. In the reported simulation, the base seed was set to seed _ base = 1 . For reproducibility and comparability, fixed seeds were used, and for the single-point Series 1–4 runs the same sample of X was reused. This guarantees identical n norm and n dev across all series and eliminates the impact of a random skew in the realized operating profile of X on between-series comparisons. In Series 1, during the search for z 2 * , the same array Z was reused for all values of z 2 (only the threshold changed), which substantially reduces noise when comparing thresholds.
2.
Common random numbers for α ( s ) and β ( s ) . To compute α ( s ) and β ( s ) in Series 1, a single shared array ε ( t ) N ( 0 , 1 ) was used, while the noise level was varied by scaling σ Y = s σ X . This ensures that differences between points in s are driven by the change in s itself rather than by changes in the particular noise realization.
3.
Statistical error assessment. For unconditional probabilities p ^ { P ^ TN , i , P ^ FA , i , P ^ TP , i , P ^ MD , i } , a standard approximation for the standard error of the frequency estimator is
se ( p ^ ) p ^ ( 1 p ^ ) N .
For conditional risks α and β , a first-order guideline is to use the effective sample size of the corresponding group: N norm = n TN + n FA for α and N dev = n TP + n MD for β . This matters because β is typically estimated from a much smaller number of abnormal realizations.

References

  1. Ceylan, B.O. Marine diesel engine turbocharger fouling phenomenon risk assessment application by using fuzzy FMEA method. Proc. Inst. Mech. Eng. M J. Eng. Marit. Environ. 2024, 238, 514–530. [Google Scholar] [CrossRef]
  2. Wang, T. Research on fault diagnosis of exhaust gas turbocharger of Marine diesel engine. Agro Food Ind. Hi Tech 2017, 28, 2988–2991. [Google Scholar]
  3. Gao, Z.; Huo, B.; Jinjie, J.; Jiang, Z. Failure investigation of gear teeth fracture of seawater pump in a diesel engine. Eng. Fail. Anal. 2019, 105, 1079–1092. [Google Scholar] [CrossRef]
  4. Peng, C. Wear test of cylinder liner and piston ring of marine diesel engine based on computer simulation technology. J. Phys. Conf. Ser. 2021, 2074, 012033. [Google Scholar] [CrossRef]
  5. Gao, B.; Xu, J.; Zhang, Z.; Liu, Y.; Chang, X. Marine diesel engine piston ring fault diagnosis based on LSTM and improved beluga whale optimization. Alex. Eng. J. 2024, 109, 213–228. [Google Scholar] [CrossRef]
  6. Afanaseva, O.; Pervukhin, D.; Afanasyev, M.; Khatrusov, A. Assessment of the State and Development Trends of Centrifugal Compressors for Marine Power Plants. Energies 2026, 19, 991. [Google Scholar] [CrossRef]
  7. Lazakis, I.; Raptodimos, Y.; Varelas, T. Predicting ship machinery system condition through analytical reliability tools and artificial neural networks. Ocean Eng. 2018, 152, 404–415. [Google Scholar] [CrossRef]
  8. Wang, M.; Qin, G.; Chen, J.; Liao, Y. Design of vibration monitoring and fault diagnosis system for marine diesel engine. In Proceedings of the 11th International Conference on Prognostics and System Health Management, PHM-Jinan 2020, Jinan, China, 23–25 October 2020; pp. 481–485. [Google Scholar] [CrossRef]
  9. Wu, H.; Yan, Q.; Ling, X. Survey on fault diagnosis of diesel engine based on vibration signal. In Proceedings of the 2017 3rd International Conference on Information Management, ICIM 2017, Chengdu, China, 21–23 April 2017; pp. 494–495. [Google Scholar] [CrossRef]
  10. Upadrashta, D.; Wijaya, T. AI/ML Based Anomaly Detection and Fault Diagnosis of Turbocharged Marine Diesel Engines: Experimental Study on Engine of an Operational Vessel. Information 2026, 17, 16. [Google Scholar] [CrossRef]
  11. Kim, S.H.; Kim, T.G.; Lee, J.; Song, H.K.; Moon, H.; Chun, C.J. Deep Hybrid Model for Fault Diagnosis of Ship’s Main Engine. J. Mar. Sci. Eng. 2025, 13, 1398. [Google Scholar] [CrossRef]
  12. Liu, H.; Sun, Y.; Chen, D.; Huang, T.; Hou, X.; Ren, Y.; Ding, L.; Liu, X. An Optimized Vibration Signal Compressed Sensing Based on Phase Blocking K-SVD Algorithm for Marine Diesel Engine Cylinders. IEEE Trans. Instrum. Meas. 2025, 74, 7501322. [Google Scholar] [CrossRef]
  13. He, X.; Wei, M.; Qiu, B.; Wang, J.; Yang, Y. Thermodynamic mechanism and data hybrid driven model based marine diesel engine turbocharger anomaly detection with performance analysis. In Proceedings of the 2019 11th CAA Symposium on Fault Detection, Supervision, and Safety for Technical Processes, SAFEPROCESS 2019, Xiamen, China, 5–7 July 2019; pp. 477–482. [Google Scholar] [CrossRef]
  14. Ouyang, S.; Chen, H. Research on Unbalanced Data Fault Diagnosis of Diesel Engine Based on WACGAN-GP. In Proceedings of the 2025 5th International Conference on Mechanical, Electronics and Electrical and Automation Control, METMS 2025, Chongqing, China, 9–11 May 2025; pp. 324–332. [Google Scholar]
  15. Li, B.; Yu, Y.; Wang, W.; Cao, B.; Xu, D.; Yao, Y. Sample generation method for marine diesel engines based on FEM simulation and DCGAN. J. Mech. Sci. Technol. 2024, 38, 2335–2345. [Google Scholar] [CrossRef]
  16. Zheng, H.; Zhou, H.; Kang, C.; Liu, Z.; Dou, Z.; Liu, J.; Li, B.; Chen, Y. Modeling and prediction for diesel performance based on deep neural network combined with virtual sample. Sci. Rep. 2021, 11, 16709. [Google Scholar] [CrossRef]
  17. Luo, C.; Zhao, M.; Fu, X.; Zhong, S.; Fu, S.; Zhang, K.; Yu, X. Thermodynamic simulation-assisted random forest: Towards explainable fault diagnosis of combustion chamber components of marine diesel engines. Meas. J. Int. Meas. Confed. 2025, 251, 117252. [Google Scholar] [CrossRef]
  18. Wang, L.; Cao, H.; Wei, L. Research on Online Diagnosis of Ship Main Engine with Unbalanced Sample. J. Phys. Conf. Ser. 2023, 2508, 012008. [Google Scholar] [CrossRef]
  19. Jiang, R.; Ou, S.; Li, B.; Liu, W.; Cao, B.; Yu, Y. A Fault Diagnosis Method for Typical Failures of Marine Diesel Engines Based on Multisource Information Fusion. Shock Vib. 2025, 2025, 1904885. [Google Scholar] [CrossRef]
  20. Asif, M.B.; Ali, H.; Aslam, H.; Safdar, R.; Abd Manan, T.S.B.; Basit, A.; Pajak, M. A multi-domain machine learning framework for intelligent condition monitoring of marine diesel engines. Eng. Res. Express 2026, 8, 025201. [Google Scholar] [CrossRef]
  21. Xi, W.; Li, Z.; Tian, Z.; Duan, Z. A feature extraction and visualization method for fault detection of marine diesel engines. Meas. J. Int. Meas. Confed. 2018, 116, 429–437. [Google Scholar] [CrossRef]
  22. Hu, L.; Wang, W.; Yu, X.; Yu, Y.; Hu, J.; Ma, B.; Yang, J. Research on sensitivity of measurement points and diagnostic method of combustion chamber health status for marine diesel engine. J. Mech. Sci. Technol. 2025, 39, 2549–2562. [Google Scholar] [CrossRef]
  23. Markelova, O.S.; Ivanovskaya, A.V. Diagnostics of a Ship Diesel Engine by Proxy Indicators of Fuel Combustion. In Proceedings of the 2022 Conference of Russian Young Researchers in Electrical and Electronic Engineering, ElConRus 2022, St. Petersburg, Russia, 25–28 January 2022; pp. 1226–1229. [Google Scholar] [CrossRef]
  24. Patil, C.; Theotokatos, G.; Wu, Y.; Lyons, T. Investigation of logarithmic signatures for feature extraction and application to marine engine fault diagnosis. Eng. Appl. Artif. Intell. 2024, 138, 109299. [Google Scholar] [CrossRef]
  25. Li, Y.; Guo, Z.; Li, Z.; Deng, Z.; Noman, K. Instantaneous Angular Speed-Based Fault Diagnosis of Multicylinder Marine Diesel Engine Using Intrinsic Multiscale Dispersion Entropy. IEEE Sens. J. 2023, 23, 9523–9535. [Google Scholar] [CrossRef]
  26. Gasparjans, A.; Terebkov, A.; Priednieks, V.; Klaucans, R. Spectral analysis of output voltages and currents as a criterion for technical diagnostics of synchronous generators of ship diesel engine. Latv. J. Phys. Tech. Sci. 2020, 57, 12–23. [Google Scholar] [CrossRef]
  27. Bui, T.M.; Nguyen, Q.D.; Nguyen, D.T.; Le Thi, H.; Thu, T.N.T.; Nguyen, T.P. Failure Analysis for Diesel Generator on Marine Ship Based Convolutional Neural Network. In Proceedings of the 9th International Artificial Intelligence and Data Processing Symposium, IDAP 2025, Malatya, Turkiye, 6–7 September 2025. [Google Scholar] [CrossRef]
  28. Drewing, S.; Witkowski, K. Spectral analysis of torsional vibrations measured by optical sensors, as a method for diagnosing injector nozzle coking in marine diesel engines. Sensors 2021, 21, 775. [Google Scholar] [CrossRef]
  29. Cui, D.; Hu, Y. Fault Diagnosis for Marine Two-Stroke Diesel Engine Based on CEEMDAN-Swin Transformer Algorithm. J. Fail. Anal. Prev. 2023, 23, 988–1000. [Google Scholar] [CrossRef]
  30. Li, C.; Cui, D. Fault Diagnosis of Marine Diesel Engine Based on Multi-scale Time Domain Decomposition and Convolutional Neural Network. Pol. Marit. Res. 2024, 31, 85–93. [Google Scholar] [CrossRef]
  31. Li, K.; Wang, H.; Gong, Q.; Hu, X.; Sun, Y.; Ou, S. Iterative generation differential decomposition with multi-component vibration signals for marine diesel engine. J. Navig. 2024, 77, 526–542. [Google Scholar] [CrossRef]
  32. Pervukhin, D.A.; Neyrus, S.K. Improving the Efficiency of the Bunkering Enterprise on the Basis of Simulation Modeling. In Proceedings of the 8th International Conference on Computational Methods in Systems and Software, CoMeSySo 2024, Online 23–26 October 2024; Silhavy, R., Silhavy, P., Eds.; Springer Science and Business Media Deutschland GmbH: Berlin/Heidelberg, Germany, 2025; Volume 1491 LNNS, pp. 481–490. [Google Scholar] [CrossRef]
  33. Pervukhin, D.; Neyrus, S. Optimization of Bunkering Logistics at Sea, Taking into Account Cost, Time and Technical Constraints. Eng 2025, 6, 364. [Google Scholar] [CrossRef]
  34. Golovina, E.; Shchelkonogova, O. Possibilities of Using the Unitization Model in the Development of Transboundary Groundwater Deposits. Water 2023, 15, 298. [Google Scholar] [CrossRef]
  35. Martirosyan, A.V.; Romashin, D.V. Investigation of the Control Strategies for Enhancing the Efficiency of Natural Gas Separation and Purification Processes. Processes 2026, 14, 700. [Google Scholar] [CrossRef]
  36. Ilyushin, Y.; Martirosyan, A.V.; Asadulagi, M.A.; Kukharova, T. Modeling and Optimization of an Automatic Temperature Control System for the Catalytic Cracking Process. Modelling 2026, 7, 68. [Google Scholar] [CrossRef]
  37. Matrokhina, K.V.; Trofimets, V.Y.; Mazakov, E.B.; Makhovikov, A.B.; Khaykin, M.M. Development of methodology for scenario analysis of investment projects of enterprises of the mineral resource complex. J. Min. Inst. 2023, 259, 112–124. [Google Scholar] [CrossRef]
  38. Galevskiy, S.G.; Qian, H. Developing and validating comprehensive indicators to evaluate the economic efficiency of hydrogen energy investments. Oper. Res. Eng. Sci. Theor. Appl. 2024, 7, 188–207. [Google Scholar] [CrossRef]
  39. Marinina, O.A. Methodological approach to economic assessment of losses of balance coal reserves. Min. Inf. Anal. Bull. 2025, 2025, 183–197. [Google Scholar]
  40. Sidorenko, A.A.; Sidorenko, S.A. A Comprehensive Strategy for Safe and Efficient Mining of Thick, Spontaneous Combustion-prone Coal Seams under Geodynamic Hazard Conditions. Int. J. Eng. Trans. B Appl. 2026, 39, 818–827. [Google Scholar] [CrossRef]
  41. Perepelkin, A.; Sharifov, A.; Titov, D.; Shandrygolov, Z.; Derkach, D.; Islamov, S. Approaches to Proxy Modeling of Gas Reservoirs. Energies 2025, 18, 3881. [Google Scholar] [CrossRef]
  42. Kukharova, T.; Maltsev, P.; Abramkin, S.; Novozhilov, I. Analysis of Modern Challenges and Technological Solutions in Natural Gas Production at Fields with Complex Geological Structure: A Review. Resources 2026, 15, 32. [Google Scholar] [CrossRef]
  43. Kukharova, T.; Martirosyan, A.; Asadulagi, M.A.; Ilyushin, Y. Development of the Separation Column’s Temperature Field Monitoring System. Energies 2024, 17, 5175. [Google Scholar] [CrossRef]
  44. Ilyushin, Y.V.; Boronko, E.A. Analysis of Energy Sustainability and Problems of Technological Process of Primary Aluminum Production. Energies 2025, 18, 2194. [Google Scholar] [CrossRef]
  45. Linh, N.K.; Tien, N.T.; Luan, D.C.; Dinh, D.V.; Thang, N.V. Enhancing Efficiency of Steel Prop Recovery Processes in Unused Mining Excavation. Int. J. Eng. Trans. B Appl. 2025, 38, 400–407. [Google Scholar] [CrossRef]
  46. Korobov, G.Y.; Parfenov, D.V.; Nguyen, V.T. Long-term Inhibition of Paraffin Deposits Using Porous Ceramic Proppant Containing Solid Ethylene-vinyl Acetate. Int. J. Eng. Trans. B Appl. 2025, 38, 1887–1897. [Google Scholar] [CrossRef]
  47. Martynenko, Y.V.; Bolobov, V.I. Effect of Motive Pressure on Ejector Entrainment Capacity. Int. J. Eng. 2026, 39, 841–848. [Google Scholar] [CrossRef]
  48. Eremeeva, A.M.; Marinets, A.R.; Oleynik, I.L.; Povarov, V.G. Recycling of Waste Cooking Oils into a Biodiesel Fuel: Kinetics and Analysis. Recycling 2026, 11, 41. [Google Scholar] [CrossRef]
  49. Eremeeva, A.M.; Chumachenko, Y.A.; Khasanov, A.F.; Oleynik, I.L. Advanced hydroprocessing technology for sustainable diesel: Hydrotreatment of renewable and fossil feedstocks. Bioresour. Technol. Rep. 2026, 33, 102499. [Google Scholar] [CrossRef]
  50. Gang, W.; Yuan, Y.; Jiang, G.; Guo, H.; Zhang, Z. Experimental study on combustion and vibration characteristics of low-speed marine diesel engine fuelled with biodiesel. Pol. Marit. Res. 2025, 32, 154–162. [Google Scholar] [CrossRef]
  51. Geng, P.; Hu, X.; Chang, X. Research on Combustion, Emissions, and Fault Diagnosis of Ternary Mixed Fuel Marine Diesel Engine. J. Mar. Sci. Eng. 2025, 13, 1561. [Google Scholar] [CrossRef]
  52. Wu, G.; Jiang, G.; Chen, C.; Jiang, G.; Pu, X.; Chen, B. An Experimental Study of the Effects of Cylinder Lubricating Oils on the Vibration Characteristics of a Two-Stroke Low-Speed Marine Diesel Engine. Pol. Marit. Res. 2023, 30, 92–101. [Google Scholar] [CrossRef]
  53. Ji, Z.; Gan, H.; Liu, B. A Deep Learning-Based Fault Warning Model for Exhaust Temperature Prediction and Fault Warning of Marine Diesel Engine. J. Mar. Sci. Eng. 2023, 11, 1509. [Google Scholar] [CrossRef]
  54. Zeng, H.; Sun, J.; Chen, C.; Jiang, K.; Wu, Z. The marine diesel engine exhaust gas temperature baseline model based on particle swarm optimised generalised regression neural network. Ships Offshore Struct. 2025, 20, 429–440. [Google Scholar] [CrossRef]
  55. Kuang, W.; Tang, Y.; Cao, L.; Liang, B.; Liu, S. The marine diesel engine exhaust gas temperature baseline model based on Improved Slime Mould Algorithm optimized generalized regression neural network. Eng. Res. Express 2025, 7, 035515. [Google Scholar] [CrossRef]
  56. Karatug, C.; Ceylan, B.O.; Arslanoglu, Y. A hybrid predictive maintenance approach for ship machinery systems: A case of main engine bearings. J. Mar. Eng. Technol. 2025, 24, 12–21. [Google Scholar] [CrossRef]
  57. Ceglie, M.; Ferrante, F.; Giannino, G. Employing Artificial Neural Network for Process Signal Estimation in the Monitoring of Smart Shipboard Diesel Engine Systems. In Progress in Marine Science and Technology; IOS Press BV: Amsterdam, The Netherlands, 2023; Volume 7, pp. 93–100. [Google Scholar] [CrossRef]
  58. Duc Nghia, M.D.; Duc Tuan, H.D. Study to Establish the Relationship Between Fuel Injection Parameters and Exhaust Emission Content of Fishing Vessels’ Diesel Engines to Diagnose the Technical State. J. Adv. Res. Fluid Mech. Therm. Sci. 2024, 116, 158–169. [Google Scholar] [CrossRef]
  59. Wang, M.; Cao, H.; Li, G. Study on the fault diagnosis method of ship main engine unbalanced data based on improved DQN. In Proceedings of the ACM International Conference Proceeding Series; Association for Computing Machinery: New York, NY, USA, 2023; pp. 15–23. [Google Scholar] [CrossRef]
  60. ISO 13372:2012; Condition Monitoring and Diagnostics of Machines—Vocabulary. International Organization for Standardization: Geneva, Switzerland, 2012.
  61. ISO 2041:2018; Mechanical Vibration, Shock and Condition Monitoring—Vocabulary. International Organization for Standardization: Geneva, Switzerland, 2018.
  62. ISO 17359:2018; Condition Monitoring and Diagnostics of Machines—General Guidelines. International Organization for Standardization: Geneva, Switzerland, 2018.
  63. Filatov, I.N.; Litvinov, S.V.; Filyustin, A.E.; Rossoshansky, P.V.; Sharapov, V.P.; Tukeev, D.L.; Skazkin, V.P.; Golovlev, D.S.; Bazhin, D.A.; Tkachenko, V.P.; et al. Device for Solving the Problem of Evaluating the Performance Indicators of Multi-Parameter Control of Missile and Artillery Weapons. Utility Model Patent RU30205U1, 20 June 2003. Available online: https://patents.google.com/patent/RU30205U1/en (accessed on 15 April 2026).
Figure 1. Series 1: selection of z 2 * by minimizing J err , raw (coarse search).
Figure 1. Series 1: selection of z 2 * by minimizing J err , raw (coarse search).
Eng 07 00190 g001
Figure 2. Series 1: local refinement of z 2 * by minimizing J err , raw .
Figure 2. Series 1: local refinement of z 2 * by minimizing J err , raw .
Eng 07 00190 g002
Figure 3. Series 1: error trade-off as z 2 varies (dependencies of P FA , P MD , and J err , raw on z 2 at fixed z 1 and s = 0.3 ).
Figure 3. Series 1: error trade-off as z 2 varies (dependencies of P FA , P MD , and J err , raw on z 2 at fixed z 1 and s = 0.3 ).
Eng 07 00190 g003
Figure 4. Series 1: α and β versus s = σ Y / σ X at fixed z 1 and z 2 * .
Figure 4. Series 1: α and β versus s = σ Y / σ X at fixed z 1 and z 2 * .
Eng 07 00190 g004
Figure 5. Series 4: selection of w * by minimizing J err , raw for two-path measurement-level fusion.
Figure 5. Series 4: selection of w * by minimizing J err , raw for two-path measurement-level fusion.
Eng 07 00190 g005
Figure 6. Comparison of Series 1–4 in terms of α / α 0 , β / β 0 , and J / J 0 relative to the baseline Series 1 (smaller is better).
Figure 6. Comparison of Series 1–4 in terms of α / α 0 , β / β 0 , and J / J 0 relative to the baseline Series 1 (smaller is better).
Eng 07 00190 g006
Table 1. List of available monitoring parameters [10].
Table 1. List of available monitoring parameters [10].
Engine ParametersTurbocharger Parameters
Engine generator statusTurbine inlet gas temperature
Engine RPMTurbine outlet gas temperature
Engine loadTurbocharger speed
Voltage and AmperageBearing and lubrication system parameters
Charged air pressureLubrication oil pressure
Charged air temperatureLubrication oil inlet temperature
Engine start pressureBearing temperature
Engine cylinder (1–6) temperaturesCooling system parameters
Fuel injection system parametersCooling water pressure
Fuel oil pressureCooling water temperature (inlet and outlet)
Fuel oil temperatureVibration
Fuel oil mass flow rateTime series and frequency domain features
Table 2. Possible decision outcomes for a single monitoring channel (TN, TP, FA, MD).
Table 2. Possible decision outcomes for a single monitoring channel (TN, TP, FA, MD).
SituationNotation
Correct alarm decision (true positive, TP) P TP = P ( Ω ¯ A )
Correct no-alarm decision (true negative, TN) P TN = P ( Ω A ¯ )
False alarm (FA) P FA = P ( Ω A )
Missed detection (MD) P MD = P ( Ω ¯ A ¯ )
Table 3. Computational experiment series plan (Monte Carlo verification protocol for decision logic using unconditional control-reliability indicators (CRIs)).
Table 3. Computational experiment series plan (Monte Carlo verification protocol for decision logic using unconditional control-reliability indicators (CRIs)).
SeriesVerified Decision-Logic ElementFixed θ i (Equation (18))Varied Factors (Reported Simulation)Model Relations UsedControlled Outputs
1Baseline single-path threshold decision (no repeats) M i = 1 , n i 1 = 1 , m i 1 = 1 , K i = 1 ; a i = b i = c i = 1 . Normal bound set via π dev z 1 .1.1: z 2 (search for z 2 * ) at fixed s = s fixed = 0.3 and fixed z 1 .
1.2: s { 0.1 , , 0.8 } at fixed ( z 1 , z 2 * ) .
1.3: control grid z 2 { 3.0 , 3.5 , 4.0 , 4.5 } at fixed s fixed .
Observation model (Equation (1)); regions Ω and ω (Equations (2) and (3)); CRI estimation (Equation (5)); risks α , β (Equations (6) and (7)); J err , raw (Appendix B.2, Appendix B.4). P ^ TN , i , P ^ TP , i , P ^ FA , i , P ^ MD , i ; α , β ; J err , raw ; dependencies on z 2 and s.
2Repeated measurements and averaging Z ¯ within a path M i = 1 , m i 1 = 1 , K i = 1 ; n i 1 = n ; a i = b i = c i = 1 . Reported point: n = 5 .No sweep: one point n = 5 at s fixed = 0.3 and fixed ( z 1 , z 2 * ) .Equations (21)–(23); implementation uses the equivalent σ Y / n representation for acceleration (Appendix B.5).Same P ^ · , i , α , β , J err , raw ; comparison with Series 1.
3Within-path confirmation (“a out of m”) over sub-cycles M i = 1 , n i 1 = 1 , m i 1 = m , K i = 1 ; a i = a ; b i = c i = 1 . Reported point: ( a , m ) = ( 2 , 3 ) .No sweep: one point ( a , m ) = ( 2 , 3 ) at s fixed = 0.3 and fixed ( z 1 , z 2 * ) .Primary decision (Equation (23)) with n = 1 ; confirmation (Equation (24)); CRI estimation (Equation (5)); risks α , β (Equations (6) and (7)).Same P ^ · , i , α , β , J err , raw ; comparison with Series 1.
4Two-path multi-fidelity measurement-level fusion ( M i = 2 ) M i = 2 ; n i 1 = n i 2 = 1 , m i 1 = m i 2 = 1 , K i = 1 ; a i = b i = c i = 1 . Reported point: ( s i 1 , s i 2 ) = ( 0.3 , 0.7 ) . w [ 0 , 1 ] with step 0.01 for selecting w * by minimizing J err , raw ; the reference point w 0 = 0.7 is also evaluated.Model (Equation (1)); measurement-level fusion (Equation (26)) in the special case M i = 2 , m i j = 1 ; CRI estimation (Equation (5)); risks α , β (Equations (6) and (7)); J err , raw (Appendix B.7).Same P ^ · , i , α , β , J err , raw ; J err , raw ( w ) ; comparison of w * and w 0 .
Note: the reported series quantify selected decision-logic elements within the proposed multi-stage framework; Stage 0 (data-quality gating) and Stage 5 (temporal filtering) are included in the architecture but are not separately isolated as standalone Monte Carlo series here. Detailed verbal descriptions of the series algorithms are provided in Appendix B; the Python implementation is included in the Supplementary Materials.
Table 4. Summary results at the shared baseline configuration for Series 1–4 (fixed z 1 = 3.290527 , z 2 * = 3.23 ; for Series 4 additionally s i 1 = 0.3 and s i 2 = 0.7 ).
Table 4. Summary results at the shared baseline configuration for Series 1–4 (fixed z 1 = 3.290527 , z 2 * = 3.23 ; for Series 4 additionally s i 1 = 0.3 and s i 2 = 0.7 ).
Series P TN P FA P TP P MD α β n norm n dev J err , raw
Series 10.997800.001180.000830.000190.001190.182001,997,96720330.00303
Series 20.998540.000450.000940.000080.000450.076241,997,96720330.00122
Series 30.998300.000690.000870.000140.000690.141171,997,96720330.00212
Series 4 ( w * )0.997920.001060.000830.000180.001060.181011,997,96720330.00290
Series 4 ( w 0 )0.997760.001220.000820.000200.001220.192821,997,96720330.00318
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Tukeev, D.; Afanaseva, O.; Khatrusov, A. Synthesis of Decision Logic for Predictive Maintenance of a Marine Diesel Engine Based on Unconditional Control-Reliability Indicators. Eng 2026, 7, 190. https://doi.org/10.3390/eng7050190

AMA Style

Tukeev D, Afanaseva O, Khatrusov A. Synthesis of Decision Logic for Predictive Maintenance of a Marine Diesel Engine Based on Unconditional Control-Reliability Indicators. Eng. 2026; 7(5):190. https://doi.org/10.3390/eng7050190

Chicago/Turabian Style

Tukeev, Dmitry, Olga Afanaseva, and Aleksandr Khatrusov. 2026. "Synthesis of Decision Logic for Predictive Maintenance of a Marine Diesel Engine Based on Unconditional Control-Reliability Indicators" Eng 7, no. 5: 190. https://doi.org/10.3390/eng7050190

APA Style

Tukeev, D., Afanaseva, O., & Khatrusov, A. (2026). Synthesis of Decision Logic for Predictive Maintenance of a Marine Diesel Engine Based on Unconditional Control-Reliability Indicators. Eng, 7(5), 190. https://doi.org/10.3390/eng7050190

Article Metrics

Back to TopTop