Next Article in Journal
Unsupervised Learning Framework for Cyber Threat Detection, Anomaly Identification, and Alert Prioritization
Next Article in Special Issue
AI-Driven Energy Management for Sustainable Transformation of Recreational Boats: A Simulation Study for the Croatian Adriatic Coast
Previous Article in Journal
Spent Coffee Grounds Extract Limits Bacterial Proliferation on Human Foot Skin Under Humid Conditions
Previous Article in Special Issue
Infrared Thermography in Maritime Systems: A Systematic Review
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Early Anomaly Detection in Maritime Refrigerated Containers Using a Hybrid Digital Twin and Deep Learning Framework

Faculty of Maritime Studies, University of Rijeka, 51 000 Rijeka, Croatia
*
Author to whom correspondence should be addressed.
Appl. Sci. 2026, 16(4), 1887; https://doi.org/10.3390/app16041887
Submission received: 26 January 2026 / Revised: 10 February 2026 / Accepted: 11 February 2026 / Published: 13 February 2026
(This article belongs to the Special Issue AI Applications in the Maritime Sector)

Abstract

Maritime refrigerated containers operate under harsh and highly variable conditions, where gradual equipment degradation can lead to temperature excursions, cargo losses, and operational disruptions. In current practice, monitoring relies largely on threshold-based temperature alarms, which are reactive and provide limited insight into early abnormal behaviour. This study proposes a hybrid framework for early anomaly detection in maritime refrigerated containers that combines a lightweight physics-based digital twin with a deep learning anomaly detector trained exclusively on fault-free operation. The approach is designed for shipboard constraints and uses only controller-level signals augmented by locally derived features, enabling low-complexity edge execution. The digital twin produces physically interpretable temperature residuals, while a convolutional autoencoder learns normal multivariate operating patterns and flags deviations via reconstruction error. Both indicators are integrated using conservative persistence gating to suppress short-lived transients typical of maritime operation. The framework is evaluated in a simulation environment calibrated to representative reefer thermal dynamics under variable ambient conditions and progressive fault injection across gradual and abrupt fault categories. Results indicate earlier and operationally credible detection compared to conventional alarms, supporting practical predictive maintenance in maritime cold-chain logistics.

1. Introduction

Maritime transport of refrigerated containers plays a critical role in global cold-chain logistics, enabling the long-distance distribution of temperature-sensitive goods such as pharmaceuticals, fresh produce, and frozen food products. The operational importance of refrigerated containers (reefers) has been recognised for more than a decade, with early work on intelligent container concepts highlighting the value of onboard sensing and data-driven monitoring for cargo integrity and logistics optimisation [1]. Subsequent market and logistics studies confirm the sustained growth of maritime refrigerated transport and its economic significance within global supply chains [2,3].
Despite continuous advances in refrigeration technology and monitoring infrastructure, temperature excursions and equipment-related failures remain a recurring operational issue in maritime transport. Unlike stationary cold storage facilities, shipboard refrigerated containers operate under highly dynamic and constrained conditions, including fluctuating ambient temperatures, variable power quality, mechanical vibration, and limited opportunities for maintenance during voyages. Empirical studies in refrigeration engineering have shown that even short-duration deviations from prescribed temperature ranges can lead to irreversible quality degradation of sensitive cargo, particularly pharmaceuticals and perishable food products [4,5,6].
In current maritime practice, the dominant monitoring paradigm for refrigerated containers relies on threshold-based alarms implemented at the controller level. Such alarms are designed primarily to protect cargo by signalling temperature limit violations, rather than to provide early warning of developing equipment degradation. This inherently reactive approach limits the ability of shipboard crews and fleet operators to intervene before faults escalate into cargo-threatening conditions. These limitations are further amplified by organisational and logistical constraints inherent to maritime operations, where restricted maintenance windows, limited onboard resources, and the predominance of corrective over predictive maintenance are well documented as systemic barriers to effective condition monitoring at sea [7,8].
Early faults (incipient degradations) are particularly problematic in this setting. They precede fully developed failures but typically do not manifest as sustained temperature excursions or explicit controller alarms. In refrigerated container systems, such early-stage abnormalities may include gradual loss of effective cooling capacity (e.g., refrigerant leakage or heat exchanger fouling), sensor bias, airflow restriction due to incipient icing, or mechanical degradation that only becomes apparent under load. During these stages, the control system may still maintain the temperature setpoint, but with altered dynamics, such as increased compressor duty cycle, reduced thermal margins, or subtle changes in transient response. From an operational perspective, these conditions often appear “normal” during routine inspection, and escalation is frequently only observed after a triggering event, for example when a unit fails to restart reliably, trips under compressor load, or drifts rapidly outside allowable limits. Detecting such early faults is challenging because the associated deviations are small, gradual, and easily masked by nominal variability arising from compressor cycling, defrost activity, ambient temperature fluctuations, and cargo thermal inertia. This motivates monitoring approaches that focus on persistent deviations in system dynamics and operating patterns, rather than relying solely on absolute temperature threshold violations.
Against this operational backdrop, fault detection and diagnostics (FDD) for vapour-compression refrigeration systems have been extensively studied in the context of building HVAC and industrial refrigeration. Early approaches focused on statistical and rule-based methods to identify abnormal behaviour based on deviations in measured signals [4]. Later research introduced systematic fault classification and prioritisation frameworks, as well as model-based diagnostic methods exploiting thermodynamic relationships to detect faults such as refrigerant leakage, heat exchanger fouling, and sensor bias [9,10,11]. While these approaches offer strong physical interpretability, they typically require accurate parameterisation, reliable sensor coverage, and stable operating conditions; assumptions that are difficult to satisfy in shipboard refrigerated container applications.
More recently, data-driven and machine learning-based approaches have gained prominence in refrigeration and HVAC fault detection, motivated by increasing availability of operational data and advances in anomaly detection techniques [12]. Beyond refrigeration-specific applications, machine learning has been widely adopted for fault diagnosis across a broad range of industrial assets, particularly rotating machinery, where supervised, unsupervised, and hybrid learning strategies, including reconstruction-based models, one-class classification, forecasting-based approaches, and physics-informed methods have been extensively reviewed and benchmarked, providing transferable insights into model choice, data requirements, and deployment trade-offs under different operational and data constraints [13]. In parallel, the maritime sector has shown growing interest in predictive maintenance supported by sensor-based monitoring frameworks and artificial intelligence, particularly for ship machinery and auxiliary systems [14,15,16]. However, the application of advanced data-driven fault detection methods to maritime refrigerated containers remains limited, especially when constrained to the minimal set of controller-level measurements typically available across container fleets [17].
Digital twin technology has emerged as a promising paradigm for integrating physical models, sensor data, and analytics into a unified representation of physical assets. Foundational work on digital twins emphasises their ability to detect undesirable system behaviour through continuous comparison of expected and observed performance [18]. Subsequent studies document their adoption across industrial domains for monitoring and predictive maintenance [19,20], including applications in maritime engineering [21,22,23]. Nevertheless, most reported maritime digital twin applications focus on large-scale ship systems rather than high-volume auxiliary assets such as refrigerated containers, which are characterised by limited sensing, strict computational constraints, and strong economic sensitivity to failure.
Against this background, this study addresses a gap in AI-supported monitoring for maritime refrigerated containers: the need for early, operationally credible warnings under highly non-stationary shipboard conditions and limited sensing. Building on established concepts of intelligent containers and remote cold-chain monitoring [1,24,25] and on wider FDD developments in refrigeration and HVAC [4,9,10,11,12,26,27], we propose and evaluate a hybrid early anomaly detection framework that combines a lightweight physics-based digital twin with reconstruction-based deep learning. The hybrid design is motivated by the documented limitations of purely physics-based residual methods under modelling mismatch and variable environments [28,29,30], as well as by the sensitivity of ML-only detectors to dataset shift and rare operating modes [31,32], which are common in maritime operation [33].
The contribution of this work is threefold. First, we introduce a residual–reconstruction monitoring architecture tailored to controller-level reefer signals, aligning with practical monitoring interfaces and constraints reported in the cold-chain literature [24,25]. Second, we provide a reproducible evaluation protocol based on controlled fault injection and fixed thresholding derived exclusively from nominal operation, consistent with recommended practices in model-based FDD [28,34]. Third, we benchmark the proposed hybrid approach against practical baselines (conventional temperature alarms, physics-only residuals, ML-only reconstruction, and ARIMA) using operationally meaningful metrics; time-to-detection and false-positive rate expressed as alerts per 1000 h, reflecting the real costs of excessive alarms and delayed intervention in maritime maintenance management [7,8,15,35,36].

2. Materials and Methods

This section describes the methodological framework adopted for early fault detection in maritime refrigerated containers. The methodology is designed to balance physical interpretability, detection sensitivity, and practical deployability under shipboard constraints. Rather than detailed component-level modelling or fault-specific classifiers, the focus is on early identification of abnormal behaviour using a minimal and realistic set of controller-level signals, consistent with practical reefer monitoring systems [37,38].
To achieve this, a hybrid monitoring framework combines a lightweight physics-based digital twin with a deep learning-based anomaly detector trained only on nominal behaviour. Hybrid and grey-box approaches are increasingly recognised as effective compromises between physics-based diagnostics and purely data-driven methods in safety-critical and data-limited settings [39,40]. The digital twin provides physically consistent expectations of thermal behaviour, while the learning-based component captures multivariate operational patterns and subtle deviations that are difficult to model explicitly. Their outputs are integrated through conservative persistence logic suitable for continuous execution on shipboard edge devices [41].
Given limited access to labelled fault data from operational reefers and the impracticality of fault induction during voyages, evaluation is performed in a simulation-based setting calibrated to representative reefer dynamics. Simulation-based evaluation is standard in refrigeration FDD research, enabling controlled fault injection, precise fault-onset definition, and repeatable assessment of early detection performance without operational risk [28].

2.1. System Overview and Available Measurements

The proposed framework targets deployment in shipboard environments, where monitoring must be autonomous, rely on a few signals, and remain robust under non-stationary conditions. In current maritime practice, refrigerated container monitoring is typically limited to measurements provided by the reefer controller [24,37]. Accordingly, the framework assumes the minimum viable controller-level set: internal air temperature T in , ambient temperature T amb , temperature setpoint T set , and compressor ON/OFF state u comp .
To improve sensitivity without introducing additional sensors, two derived features are computed locally from controller-level measurements. The temperature error e T   =   T in   -   T set captures persistent offsets relative to the control reference, while the short-term temperature rate T ˙ in   =   d T in / dt , estimated using a smoothed finite-difference scheme, provides information on thermal dynamics and system inertia. Both features are derived solely from controller-level signals and require no additional sensing or proprietary measurements.
In practice, controller-level signals are often sampled at low rates and may contain missing segments due to communication or logging interruptions, especially when monitored through shipboard or mobile interfaces. For this reason, the framework is designed to tolerate imperfect telemetry and to rely on variables that are commonly available across reefer fleets rather than proprietary internal measurements [1,24,25]. This design choice prioritises deployability and fleet-wide consistency over laboratory-grade instrumentation, reflecting the operational constraints and maintenance realities reported for shipboard systems [7,8,14,15].
Figure 1 illustrates the monitoring context and controller-level data acquisition for representative reefer units accessed through standard interfaces.
For clarity and reproducibility, Table 1 summarises the controller-level signals and locally derived features assumed throughout the study, linking each variable to its operational relevance and typical availability across commercial reefer units.

2.2. Hybrid Digital Twin and Deep Learning Framework

The hybrid monitoring framework integrates a reduced-order physics-based digital twin with a deep learning reconstruction-based anomaly detector. The objective is not detailed fault isolation, but early and reliable detection of abnormal behaviour using computationally efficient methods suitable for continuous execution on shipboard devices.

2.2.1. Physics-Based Digital Twin

The digital twin models the dominant thermal dynamics using a lumped energy balance:
C d T in dt   =   Q amb     Q cool   +   Q load   +   Q defrost ,
where C is the effective thermal capacitance. Ambient heat ingress is modelled as Q amb =   ( T amb     T in ) / R , where R is the effective thermal resistance. Cooling capacity is modelled as Q cool   =     u comp η ( T amb ) Q rated , where u comp denotes the compressor ON/OFF state, η ( ) captures ambient-dependent efficiency effects and Q rated is the nominal cooling capacity [38,42,43,44]. Additional thermal loads such as cargo respiration and defrost activity are aggregated into Q load and Q defrost , respectively. Related reefer-container thermal modelling and control studies, including the role of environmental heat loads (e.g., solar radiation) and model-based control formulations, are reported in [45,46,47].
The continuous-time model is discretized with a sampling interval of Δ t   =   1   s to match controller-level data acquisition and to support real-time execution on shipboard edge hardware. Under fixed operating mode, the discretized predictor can be written in compact form:
T in ( k   +   1 )   =   T in ( k )   +   α ( T amb ( k )     T in ( k ) )     β u comp ( k ) ,
where α and β are effective parameters identified under nominal (fault-free) operation and held fixed during monitoring.
Equation (2) is not intended to represent a full high-fidelity thermodynamic simulation of the refrigeration system. Instead, it serves as a deliberately simplified one-step predictor whose primary role is to generate physically consistent short-horizon temperature expectations for residual-based anomaly detection. This design choice reflects both the monitoring objective and the operational constraints of maritime refrigerated containers. In early fault detection, the objective is not long-term temperature forecasting accuracy but the reliable identification of sustained deviations between measured behaviour and physically plausible system response.
More detailed refrigeration models incorporating multi-node thermal states, nonlinear heat transfer coefficients, variable refrigerant properties, or explicit compressor dynamics can improve prediction accuracy under specific operating conditions, but they also introduce additional parameters that are difficult to identify and maintain under shipboard conditions. In practice, parameter uncertainty, container-to-container variability, ageing effects, and unmeasured disturbances (e.g., cargo heterogeneity, partial door openings, power quality fluctuations) limit the reliability of such complex models for continuous fleet-scale monitoring.
The reduced-order formulation in (2) represents a pragmatic compromise: it captures the dominant first-order thermal behaviour driven by ambient conditions and cooling activity, while remaining robust to modelling mismatch and computationally lightweight. Importantly, any persistent discrepancy between the measured temperature T in , meas ( k ) and the predicted value T in , pred ( k ) reflects abnormal system behaviour rather than numerical artefacts of an over-parameterised model.
Deviations between measured and predicted temperatures are quantified as residuals:
r ( k )   =   T in , meas ( k )     T in , pred ( k ) .  
Calibration of the twin was performed to match representative reefer thermal behaviour under nominal operating profiles; prediction error under fault-free conditions remained sub-degree, supporting residual-based monitoring under realistic variability. On a disjoint fault-free validation set, the calibrated digital twin achieved an overall RMSE of 0.42 °C across 72 h of operation spanning ambient temperatures from 20 °C to 40 °C, confirming sufficient predictive fidelity for residual-based anomaly detection.

2.2.2. Deep Learning-Based Anomaly Detection

In parallel, a deep learning model learns nominal multivariate operating patterns and detects deviations by reconstruction error. A convolutional autoencoder (CAE) is trained exclusively on fault-free data using a sliding window of 60 min (3600 samples at 1 Hz). The CAE input comprises five features:
x ( t )   =   [ T in ( t ) ,   T amb ( t ) ,   u comp ( t ) ,   e T ( t ) ,   T ˙ ( t ) ] .  
The CAE architecture was not selected arbitrarily. Several candidate configurations with varying numbers of convolutional layers, filter widths, and latent-space dimensions were evaluated on a nominal validation set to balance reconstruction fidelity, generalisation under non-stationary operation, and computational cost. Specifically, variants with different latent-space dimensions (4, 8, and 16) and convolutional filter configurations were assessed; the reported architecture with an 8-dimensional latent space and three Conv1D blocks (64/32/16 filters) was the smallest configuration that consistently achieved stable reconstruction on the nominal validation set without increased false-alarm behaviour. The selected configuration represents the smallest architecture that consistently achieved stable reconstruction performance and low false-alarm behaviour, while remaining suitable for continuous execution on shipboard edge hardware. Larger architectures did not yield meaningful performance improvements under the adopted validation criteria, whereas smaller models exhibited reduced sensitivity to gradual degradation patterns.
The encoder uses three strided Conv1D blocks (64/32/16 filters, kernel size 5, stride 2, ReLU), followed by a compact latent representation (Dense (8), ReLU). The decoder mirrors the encoder using transposed convolutions to reconstruct the input window. The anomaly score E ( k ) is the window-wise reconstruction error; higher values indicate deviations from learned nominal behaviour. Reconstruction error E ( k ) is computed as the root mean square error between the input and reconstructed signals over each window after inverse standardisation, and is therefore expressed in temperature units (°C). The anomaly detection threshold is fixed to 2.1 °C and is derived exclusively from nominal (fault-free) validation data as an upper-tail cutoff of the reconstruction-error distribution, selected to ensure low false-alarm behaviour under non-faulty operation. This reconstruction approach is suitable when labelled fault data are limited [39,40,48]. The key architectural parameters, input features, and training configuration of the CAE are summarised in Table 2.
Prior to CAE training, all continuous input channels ( T in , T amb , e T , T ˙ ) are standardised using z-score normalisation computed on nominal (fault-free) training data only, while the binary compressor state u comp is left unscaled. The CAE is trained solely on fault-free sequences and validated on a disjoint nominal subset to prevent leakage of fault signatures into model selection. Training uses a fixed window length (60 min) and early stopping based on validation reconstruction loss to reduce overfitting under non-stationary but fault-free variability. For time-series anomaly detection under non-stationary operating conditions, autoencoder-based and recurrent deep learning approaches are commonly used in the literature [49,50,51].
Decision thresholds and persistence-based alerting logic are defined exclusively using nominal validation data and are described in detail in Section 2.2.3.

2.2.3. Hybrid Decision Logic

The hybrid detector fuses physics-based residual monitoring with data-driven reconstruction error analysis using explicit, rule-based decision logic designed for operational robustness. The decision process jointly exploits temperature residuals produced by the physics-based digital twin and reconstruction errors obtained from the convolutional autoencoder, thereby combining complementary sensitivity to abrupt fault events and gradual degradation mechanisms.
Figure 2 summarises the end-to-end processing pipeline from controller-level measurements, through physics-based prediction and deep learning reconstruction, to the final alert decision logic.
A physics-based alert is triggered when the absolute temperature residual
| r ( k ) |   =   | T in , meas ( k )   T in , pred ( k ) |
exceeds a fixed threshold of 1.5 °C continuously for at least 10 min (i.e., 600 consecutive samples at 1 Hz). This threshold reflects the combined effects of model uncertainty and sensor accuracy under nominal operation and is chosen to avoid sensitivity to short-lived deviations caused by compressor cycling or abrupt ambient changes.
An autoencoder-based alert is triggered when the reconstruction error E ( k ) exceeds its fault-free nominal-validation-derived upper-tail threshold of 2.1 °C continuously for at least 5 min (i.e., 300 consecutive samples).
Both thresholds were determined exclusively using fault-free (nominal) operation data and were fixed prior to all fault injection experiments. For the physics-based residual, the 1.5 °C threshold was selected based on the upper tail of residual distributions observed under nominal conditions, accounting for sensor accuracy, modelling error, and environmental variability. For the learning-based component, the reconstruction-error threshold of 2.1 °C was selected as an upper-tail cutoff of the nominal reconstruction-error distribution (i.e., a high percentile of fault-free validation errors), with the explicit objective of achieving a low false-positive rate under non-faulty operation. Injected-fault scenarios are used exclusively for performance evaluation (e.g., reporting F1-score and time-to-detection), and are not used for threshold selection.
In practice, these upper-tail choices make the alerting conservative by construction, such that moderate nominal variability is absorbed without triggering alarms, while sustained deviations are captured through the persistence rule.
Here, “continuously” denotes a strict persistence criterion: the monitored signal must remain above the threshold for the full specified duration, and any single sample falling below the threshold resets the persistence counter. This conservative design suppresses spurious alerts caused by short-lived transients or brief recoveries, prioritising operational robustness over maximum detection earliness.
Although controller-level measurements are available at 1 Hz, the monitoring pipeline does not require the full anomaly-detection computation to be executed independently for every individual sample. The digital twin state is updated synchronously with incoming data due to its negligible computational cost, whereas the CAE operates on sliding windows and produces a window-level reconstruction score E ( k ) . In implementation, E ( k ) can be evaluated at 1 Hz using an updated rolling window, but the decision-making remains governed by minute-scale persistence (5–10 min). Consequently, a lower evaluation rate (e.g., computing E ( k ) every few seconds) can be used without materially affecting alert timing under the adopted criteria, while further reducing edge compute load. For completeness, the runtime results reported later in the paper were obtained under a conservative upper-bound execution setting in which CAE inference was evaluated at 1 Hz; in practical deployment, a coarser evaluation stride (e.g., every 5–10 s) is sufficient under the adopted 5–10 min persistence criteria and would proportionally reduce computational load without materially affecting alert timing. Practical feasibility and computational margins on representative shipboard edge hardware are quantified later together with the overall baseline comparison and runtime measurements.
The final system alert follows an OR-gated decision logic: an event-level alert is issued if either the physics-based or the autoencoder-based alert satisfies its respective persistence condition. In other words, the system does not require both indicators to exceed their thresholds simultaneously; a persistent violation by either indicator is sufficient to trigger an alert. No weighted score fusion is used; alerts are generated by rule-based OR gating with persistence. For operational interpretability, alert confidence is classified internally: alerts confirmed by both indicators within a 5 min interval are labelled as high-confidence events, whereas alerts triggered by a single indicator are retained as medium-confidence warnings.
This asymmetric fusion strategy is intentional. Physics-based residuals enable rapid detection of abrupt faults, such as compressor shutdowns, while the autoencoder is more sensitive to gradual degradations and multivariate pattern changes that may not immediately manifest as temperature limit violations. Persistence-based gating suppresses benign transients caused by compressor cycling, defrost activity, or short ambient disturbances, which are common in maritime operations and represent a known barrier to the adoption of advanced monitoring systems in remote cold-chain logistics [24,25]. By avoiding reliance on the learning-based component as a standalone decision-maker, the hybrid logic aligns with recommendations on interpretable and trustworthy AI for safety-relevant monitoring applications [32,41].

2.3. Fault Scenarios and Evaluation Setup

To assess early detection performance under controlled conditions, the framework is evaluated using a simulation environment representing thermal and operational behaviour of a maritime refrigerated container. Simulation enables systematic fault injection with known onset times and severity levels [28]. Each run follows a fixed protocol: an initial nominal period is used for stable baseline behaviour, followed by fault onset and an observation period under faulty operation. For each scenario, 50 independent replications are executed under randomised ambient profiles, setpoints, and operational variability to reflect non-stationary conditions.
To ensure fair and reproducible comparison, all decision thresholds and persistence parameters are fixed using nominal data only and then held constant across all fault categories, severities, and baseline methods. This mirrors standard principles in model-based FDD evaluation, where thresholds must be defined without access to fault labels to avoid optimistic bias [28,34]. Simulation-based fault injection is widely used in refrigeration and HVAC diagnostics because it enables controlled onset times, repeatability across replications, and systematic severity scaling without risking cargo integrity or requiring intrusive experiments [11,26,27]. The adopted protocol therefore prioritises comparability and operational interpretability over fault-specific tuning, reflecting the practical goal of early warning rather than detailed fault isolation.
Representative fault scenarios are summarised in Table 3, covering both gradual and abrupt failure modes relevant to maritime refrigerated container operation and early anomaly detection.
Early detection performance is quantified by time-to-detection (TTD):
TTD   =   t alert     t fault ,
computed only for true-positive detections, where t fault is known from simulation logs. False-positive behaviour is reported as alerts per 1000 h of fault-free operation:
FPR   =   # FP   alerts hours   of   nominal   operation   ×   1000
Detection performance metrics were computed at the event level, where each injected fault episode constituted a single detection instance, thereby avoiding inflation of performance scores due to sample-level counting. Time-to-detection (TTD) was defined as the elapsed time between fault onset and the first persistent alert, while the false-positive rate was normalised per 1000 h of fault-free operation to reflect operational relevance in maritime monitoring practice.

3. Results and Discussion

This section analyses the performance of the hybrid early anomaly detection framework under representative operating conditions. The analysis focuses on deployment-relevant behaviour: stability during nominal transients, timeliness under gradual degradation, and responsiveness to abrupt faults. Results are discussed in relation to anomaly detection and hybrid modelling trade-offs, with emphasis on comparison to conventional temperature alarms used in maritime practice. Importantly, all detection thresholds and persistence settings are fixed a priori using nominal (fault-free) data only; injected-fault scenarios are used exclusively to evaluate detection performance and time-to-detection under controlled conditions.

3.1. Robustness Under Nominal Operation: Operational Realism Versus Algorithmic Sensitivity

A critical requirement for anomaly detection in maritime refrigerated containers is robustness under nominal yet highly variable conditions. Maritime reefers experience ambient fluctuations, compressor cycling and operational disturbances, which can produce benign transients that should not trigger alarms. Excessive alarm sensitivity is a well-known barrier to the adoption of advanced monitoring in cold-chain operations [24,25].
From an operational perspective, false alarms are not merely a statistical inconvenience: they create additional workload, desensitise operators, and can undermine trust in monitoring systems, particularly when maintenance actions are constrained by voyage schedules and limited onboard resources [7,8,15]. This is consistent with broader predictive maintenance literature, where the cost of false positives often dominates the perceived value of advanced analytics in production settings [35,36]. In maritime environments, additional factors such as heterogeneous container fleets, variable ambient exposure, and uneven data quality further increase the likelihood of “apparent anomalies” that are not actionable faults, making conservative alerting a rational design objective rather than a limitation [14,31,33]. The proposed hybrid approach aims to remain effective under such variability by constraining detection within physically plausible dynamics while retaining the ability to detect subtle multivariate deviations that may precede temperature excursions [52,53,54]. To reflect these operational priorities, the anomaly thresholds and persistence rules are deliberately conservative and are determined exclusively from the upper tails of nominal residual and reconstruction-error distributions, rather than tuned to maximise fault-case accuracy.
Table 4 shows that the proposed hybrid framework maintains consistently high specificity (TNR ≈ 98–99%) and stable precision across fault categories, indicating conservative false-alarm behaviour. At the same time, sensitivity and TTD improve systematically with fault severity, suggesting proportional responsiveness to progressing faults rather than premature triggering. This operational balance is consistent with the framework design: physically grounded residuals and reconstruction-based deviations are fused with persistence logic to suppress short-lived transients while retaining early sensitivity to sustained abnormal behaviour.

3.2. Fault-Dependent Detection Behaviour: What the Results Reveal Beyond Performance Metrics

Table 4 and Figure 3 summarise detection performance across fault categories. Note that Figure 3 reports two metrics with opposite ‘better’ directions: higher F1-scores indicate better detection accuracy, whereas lower time-to-detection (TTD) values indicate faster detection. Compressor failure (F3) achieves near-perfect detection with the fastest detection time, consistent with abrupt loss of cooling capacity. In contrast, refrigerant leakage (F4) is the most challenging gradual fault, requiring longer observation as early-stage dynamics can remain within acceptable temperature bands while progressively altering duty cycle and residual structure. Condenser blockage (F5) shows a similar pattern, with detection time improving as severity increases.
The more gradual fault classes (F4 refrigerant leak and F5 condenser blockage) illustrate why early warning is challenging in practice: initial stages can manifest primarily as small changes in effective capacity, duty cycle, and residual structure while temperatures remain within acceptable bounds [26,27,42,43]. This aligns with established refrigeration and HVAC fault modelling, where latent faults often become clearly separable only after sufficient accumulation of deviation under varying loads and environments [9,10,11,26]. In this context, grey-box and hybrid strategies have been advocated as a pragmatic compromise, preserving physical consistency while accommodating unmodelled effects and limited sensing [30,40,55]. The observed fault-dependent detection behaviour is therefore consistent with both the thermodynamic intuition of degradation-driven performance drift and the methodological expectation that persistence-based criteria trade maximum earliness for robust deployability [28,32,41]. Notably, for the most gradual scenarios (e.g., slow refrigerant leakage), longer TTD reflects not only weaker fault signatures but also the deliberately conservative persistence gating designed to suppress transient maritime variability; this trade-off is intrinsic to operationally credible alerting.

3.3. Early Warning Versus Reactive Alarms: Operational Value and Baseline Comparison

Beyond fault detectability, the operational value of anomaly detection lies in the actionable lead time available before cargo-threatening temperature excursions occur. This perspective shifts evaluation away from classification accuracy alone toward intervention relevance under real maritime constraints [7,8,15]. Table 5 compares the proposed hybrid framework with practical baseline monitoring approaches commonly used or reported in maritime and refrigeration practice.
The baseline methods in Table 5 represent the main detection paradigms that are feasible for shipboard reefer monitoring with controller-level signals: threshold alarms reflecting current practice, physics-based residual monitoring (model-based FDD), reconstruction-based deep learning (CAE) for multivariate time series, and statistical forecasting (ARIMA). More complex learning architectures typically assume labelled fault data, richer sensing, or higher computational resources than are available at the controller level and are therefore not considered in this study.
In addition to these baselines, a classical unsupervised anomaly detection method based on One-Class Support Vector Machines (OCSVM) was included as a reference “advanced” data-driven approach. The OCSVM was trained exclusively on fault-free (nominal) data using the same controller-level feature set as the learning-based component of the proposed framework. Detection thresholds were fixed based on the nominal validation dataset to ensure low false-positive behaviour under non-faulty operation, consistent with the thresholding strategy applied across all baseline methods.
To ensure a fair comparison, all baselines employ fixed thresholds derived from the same nominal validation dataset and apply persistence rules where applicable. Statistical significance was assessed using paired tests for time-to-detection (TTD) and detection-rate comparisons with multiple-comparison correction ( α = 0.05). In particular, the ARIMA baseline employs a fixed ARIMA (5,1,3) configuration derived from the same nominal validation dataset and held constant across all fault categories and severity levels, thereby avoiding optimistic, fault-specific tuning. Likewise, the CAE threshold and the physics residual threshold are not adjusted using fault-case outcomes; they remain fixed as nominal-derived operating cutoffs, and performance metrics are reported post hoc on injected-fault scenarios.
The Baseline definitions (fixed from nominal validation):
  • Threshold-based: alert if T in   >   T set   +   2   ° C for >30 min
  • Physics-only: alert if T in , meas     T in , pred   > 1.5   ° C for >10 min
  • ML-only: alert if reconstruction error E   >   2.1   ° C (same windowing)
  • OCSVM: alert if decision function exceeds nominal-derived threshold (same persistence)
  • ARIMA: ARIMA (5,1,3), alert if observed exceeds 95% prediction interval
CPU utilisation values were measured on a representative edge-class processor under continuous execution and averaged across nominal and injected-fault scenarios. During these measurements, CAE inference was executed using a rolling window updated at 1 Hz, representing a conservative upper-bound execution scenario; lower evaluation rates would proportionally reduce computational load without affecting alert timing due to the adopted persistence logic.
The hybrid method achieves earlier detection than all baselines while maintaining an operationally credible false-positive rate. Threshold-based alarms remain stable but are inherently reactive and provide limited preventive lead time. Physics-only residual approaches offer strong interpretability but can be sensitive to modelling mismatch and environmental variability. ML-only reconstruction methods tend to detect anomalies earlier but exhibit the highest false-positive rates under non-stationary operating conditions. By fusing physics-based residuals with learning-based reconstruction error and enforcing conservative persistence logic, the proposed hybrid framework balances early sensitivity with operational trust. The One-Class SVM baseline provides a representative reference for classical unsupervised anomaly detection trained on nominal data only; while it achieves reasonable detection performance, it exhibits higher false-positive rates under non-stationary operating conditions compared to the proposed hybrid framework, reinforcing the benefit of combining physics-based residuals with multivariate learning and persistence-based decision logic.
Figure 4 translates detection performance into operational lead time. Across fault categories, higher-severity faults are detected more rapidly, while gradual degradation mechanisms require longer observation to ensure robustness. Importantly, a critical temperature threshold of 8 °C is used as a representative upper limit for chilled cargo, beyond which quality degradation and regulatory non-compliance become likely. The median lead time of approximately 68 min between anomaly detection and this threshold therefore provides a meaningful intervention window. Here, lead time is defined as the elapsed time between the first persistent alert and the simulated crossing of the critical temperature threshold of 8 °C; for the representative moderate cooling-degradation case this corresponds to approximately 68 min, derived from a mean time-to-detection of 18.3 ± 7.2 min versus 87 ± 15 min to reach 8 °C without intervention. In maritime refrigerated transport, such lead time enables planning of corrective actions before cargo integrity is compromised, addressing a key limitation of conventional reactive alarm-based monitoring [42,56,57]. While the exact magnitude of lead time is scenario-dependent, the consistent separation between alert time and threshold crossing across gradual fault classes indicates that the framework delivers operationally meaningful early warning under the adopted conservative alerting policy.

4. Limitations and Future Work

While the proposed hybrid digital twin and deep learning framework demonstrates strong early fault detection performance and operational feasibility, several limitations should be acknowledged.
First, the evaluation is based exclusively on high-fidelity simulation. Although the thermodynamic model was calibrated to representative reefer dynamics and achieved sub-degree prediction accuracy under nominal conditions (RMSE = 0.42 °C), simulation cannot fully capture all real-world phenomena encountered during maritime operation. These include intermittent communication losses, sensor noise and degradation, crew interventions, and unforeseen environmental interactions. Consequently, field validation on operational vessels with real cargo is essential to confirm real-world performance, robustness, and operator acceptance. In this sense, the present study should be regarded as a preparatory step toward shipboard pilot deployment, providing a controlled and reproducible foundation for subsequent in-service validation under real operational conditions. A particularly important aspect of field validation will be verifying that the nominal-derived thresholds remain stable under real telemetry artefacts (missing data, asynchronous sampling, sensor dropouts) and under rare but non-faulty operating modes (e.g., unusual defrost timing, atypical ambient exposure).
Second, the digital twin employs a reduced-order thermal model with fixed parameters identified under nominal conditions. While this design supports physical interpretability and low computational cost, long-term parameter drift, container ageing, and heterogeneity across manufacturers and container sizes may affect residual fidelity over extended deployment horizons. In fleet-scale operation, this implies a need for lightweight recalibration or adaptation strategies. While no explicit global sensitivity analysis was performed in the present study, limited perturbation tests during model calibration were conducted by varying individual thermal parameters (e.g., effective thermal resistance R and cooling gain β ) within ±10% of their nominal values. These tests indicated that such moderate variations primarily affect the magnitude of temperature residuals rather than their temporal structure. Because anomaly detection relies on thresholds derived exclusively from nominal operation and on persistence-based decision logic, these parameter variations were observed to shift detection timing only marginally (on the order of 1–2 min in representative gradual fault scenarios), without inducing spurious alerts under fault-free conditions. In this sense, the framework is intentionally designed to tolerate moderate parameter mismatch, prioritising robustness and operational credibility over precise physical parameter identification. A systematic sensitivity analysis and parameter importance ranking across different fault modes therefore remain an important extension of the present study.
Third, the simulated fault library, while covering a substantial subset of commonly reported reefer failures, is not exhaustive. Certain operational scenarios, such as complete power interruptions, simultaneous multi-fault conditions, or atypical defrost behaviours, were not explicitly modelled. In the present study, defrost-related effects were treated as transient disturbances mitigated by persistence logic rather than as explicit fault events. Extending the simulator to incorporate such scenarios represents a natural extension of this work. In addition, simultaneous faults may introduce non-additive signatures that challenge the current OR-gated logic; evaluating multi-fault behaviour is therefore a necessary next step before operational deployment.
Finally, the deep learning component relies on extended periods of fault-free operational data for training. While this assumption is realistic for mature deployments, initial commissioning phases may require the system to operate temporarily in a physics-dominant mode before sufficient data are accumulated. Future research will therefore explore transfer and federated learning approaches to accelerate learning across fleets while respecting bandwidth, privacy, and operational constraints typical of maritime environments. Although the selected CAE threshold yields a favourable precision–recall trade-off under injected-fault scenarios, such metrics are reported for evaluation only and are not used during threshold selection, which remains strictly nominal-derived. Future work should also quantify how much nominal data diversity (routes, climates, cargo loads) is required for stable reconstruction-error tails, since threshold stability ultimately depends on how well nominal training/validation captures the true range of non-faulty maritime variability.

5. Conclusions

This work presented a hybrid digital twin and deep learning framework for early anomaly detection in maritime refrigerated containers, explicitly designed to operate under shipboard constraints. By integrating a reduced-order physics-based thermal model with a convolutional autoencoder trained exclusively on fault-free operation, the proposed framework enables early and operationally credible identification of abnormal system behaviour using only standard controller-level signals and lightweight edge computation.
Through systematic simulation-based evaluation across representative gradual and abrupt fault categories, the hybrid approach demonstrated clear advantages over conventional threshold-based alarms and standalone physics-only or data-driven methods. In particular, the results show that combining physically interpretable residuals with multivariate reconstruction-based anomaly detection provides earlier and more reliable warning of developing faults, while maintaining robustness against benign operational transients typical of maritime environments.
A central insight emerging from this study is that neither physics-based nor data-driven approaches alone are sufficient to meet the competing demands of early detection, interpretability, and operational trust in shipboard cold-chain monitoring. The proposed hybrid structure constrains anomaly detection within physically plausible behaviour while retaining sensitivity to subtle degradation patterns that may precede overt temperature excursions. The use of explicit persistence-based decision logic further reinforces operational credibility by suppressing short-lived deviations that are not actionable.
From an implementation perspective, the reliance on controller-level measurements, fixed nominal calibration, and low-complexity decision logic supports autonomous deployment on shipboard edge devices without dependence on continuous cloud connectivity. This alignment with practical constraints related to bandwidth, reliability, and maintainability is critical for real-world adoption in maritime operations.
Overall, the findings support hybrid, physics-informed AI as a pragmatic and scalable pathway for advancing refrigerated container monitoring from reactive alarm-based practices toward predictive maintenance strategies that better protect cargo integrity and operational continuity.

Author Contributions

Conceptualization, M.V.; investigation, J.Ć., D.O. and A.P.H.; methodology, M.V., J.Ć., D.O. and A.P.H.; supervision, M.V.; writing—original draft, J.Ć., D.O. and A.P.H.; writing—review and editing, M.V. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded under the project line of the University of Rijeka, within the project “Multimedia Interactive Simulation Systems in the Maritime Sector (MISS)” (UNIRI-MZI-25-47). Within the MISS project framework, the proposed monitoring and early anomaly detection architecture represents the analytical core intended for integration with immersive visualisation and decision-support interfaces, supporting operator training and maintenance-oriented situational awareness based on real-time condition monitoring outputs.

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding author.

Acknowledgments

During the preparation of this manuscript, the authors used a generative AI tool (Claude 3.5, Anthropic, San Francisco, CA, USA, 2025) solely for the visual drafting of Figure 2 (Architecture of the proposed hybrid early fault detection framework). The scientific content, system architecture, and methodological design of the figure were fully specified by the authors, who reviewed, edited, and validated the AI-generated graphic and take full responsibility for the content of this publication.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AIArtificial Intelligence
ARIMAAutoregressive Integrated Moving Average
CAEConvolutional Autoencoder
CIConfidence Interval
CPUCentral Processing Unit
FDDFault Detection and Diagnostics
FPRFalse-Positive Rate
HVACHeating, Ventilation, and Air Conditioning
MLMachine Learning
OCSVMOne-Class Support Vector Machine
RMSERoot Mean Square Error (°C)
TNRTrue Negative Rate (Specificity)
TPRTrue Positive Rate (Sensitivity)
TTDTime-to-Detection

References

  1. Lang, W.; Jedermann, R.; Mrugala, D.; Jabbari, A.; Krieg-Brückner, B.; Schill, K. The “intelligent container”—A cognitive sensor network for transport management. IEEE Sens. J. 2010, 11, 688–698. [Google Scholar] [CrossRef] [Scilit]
  2. Castelein, B.; Geerlings, H.; Van Duin, R. The reefer container market and academic research: A review study. J. Clean. Prod. 2020, 256, 120654. [Google Scholar] [CrossRef] [Scilit]
  3. Arduino, G.; Carrillo Murillo, D.; Parola, F. Refrigerated container versus bulk: Evidence from the banana cold chain. Marit. Policy Manag. 2015, 42, 228–245. [Google Scholar] [CrossRef] [Scilit]
  4. Rossi, T.M.; Braun, J.E. A statistical, rule-based fault detection and diagnostic method for vapor compression air conditioners. Hvac&R Res. 1997, 3, 19–37. [Google Scholar]
  5. Frank, S.M.; Kim, J.; Cai, J.; Braun, J.E. Common Faults and Their Prioritization in Small Commercial Buildings: February 2017–December 2017 (No. NREL/SR-5500-70136); National Renewable Energy Lab. (NREL): Golden, CO, USA, 2018. [Google Scholar]
  6. Visek, E.; Mazzarella, L.; Motta, M. Temperature sensor signal reconstruction for failure detection of vapor compression system. Appl. Soft Comput. 2017, 60, 679–688. [Google Scholar] [CrossRef] [Scilit]
  7. Park, M.H.; Lee, W.J. Comprehensive review of shipboard maintenance management strategies. Results Eng. 2025, 27, 106671. [Google Scholar] [CrossRef] [Scilit]
  8. Kalafatelis, A.S.; Nomikos, N.; Giannopoulos, A.; Alexandridis, G.; Karditsa, A.; Trakadas, P. Towards predictive maintenance in the maritime industry: A component-based overview. J. Mar. Sci. Eng. 2025, 13, 425. [Google Scholar] [CrossRef] [Scilit]
  9. Katipamula, S.; Brambley, M.R. Methods for fault detection, diagnostics, and prognostics for building systems—A review. Part I Hvac&R Res. 2005, 11, 3–25. [Google Scholar]
  10. Katipamula, S.; Brambley, M.R. Methods for fault detection, diagnostics, and prognostics for building systems—A review. Part II Hvac&R Res. 2005, 11, 169–187. [Google Scholar]
  11. Mulumba, T.; Afshari, A.; Yan, K.; Shen, W.; Norford, L.K. Robust model-based fault diagnosis for air handling units. Energy Build. 2015, 86, 698–707. [Google Scholar] [CrossRef] [Scilit]
  12. Matetić, I.; Štajduhar, I.; Wolf, I.; Ljubic, S. A review of data-driven approaches and techniques for fault detection and diagnosis in HVAC systems. Sensors 2022, 23, 1. [Google Scholar] [CrossRef] [Scilit]
  13. Wang, Q.; Huang, R.; Xiong, J.; Yang, J.; Dong, X.; Wu, Y.; Wu, Y.; Lu, T. A survey on fault diagnosis of rotating machinery based on machine learning. Meas. Sci. Technol. 2024, 35, 102001. [Google Scholar] [CrossRef] [Scilit]
  14. Raptodimos, Y.; Lazakis, I.; Theotokatos, G.; Varelas, T.; Drikos, L. Ship sensors data collection and analysis for condition monitoring of ship structures and machinery systems. In Smart Ship Technolology; The Royal Institution of Naval Architects: London, UK, 2016. [Google Scholar] [CrossRef] [Scilit]
  15. Simion, D.; Postolache, F.; Fleacă, B.; Fleacă, E. Ai-driven predictive maintenance in modern maritime transport—Enhancing operational efficiency and reliability. Appl. Sci. 2024, 14, 9439. [Google Scholar] [CrossRef] [Scilit]
  16. Setiawan, A.; Handoko, W.; Onn, C.W.; Widyaningsih, U. IoT-Based Machine Learning for Predictive Maintenance of Ship Engines: Efficiency Gains in Maritime Digitalization. Din. Bahari 2025, 6, 83–99. [Google Scholar] [CrossRef] [Scilit]
  17. Ji, J.; Han, H. Research Article A Fault Detection Model of Marine Refrigerated Containers. Res. J. Appl. Sci. Eng. Technol. 2013, 5, 4066–4070. [Google Scholar] [CrossRef] [Scilit]
  18. Grieves, M.; Vickers, J. Digital twin: Mitigating unpredictable, undesirable emergent behavior in complex systems. In Transdisciplinary Perspectives on Complex Systems: New Findings and Approaches; Springer International Publishing: Cham, Switzerland, 2016; pp. 85–113. [Google Scholar]
  19. Tao, F.; Zhang, H.; Liu, A.; Nee, A.Y. Digital twin in industry: State-of-the-art. IEEE Trans. Ind. Inform. 2018, 15, 2405–2415. [Google Scholar] [CrossRef] [Scilit]
  20. Fuller, A.; Fan, Z.; Day, C.; Barlow, C. Digital twin: Enabling technologies, challenges and open research. IEEE Access 2020, 8, 108952–108971. [Google Scholar] [CrossRef] [Scilit]
  21. Coraddu, A.; Oneto, L.; Baldi, F.; Cipollini, F.; Atlar, M.; Savio, S. Data-driven ship digital twin for estimating the speed loss caused by the marine fouling. Ocean. Eng. 2019, 186, 106063. [Google Scholar] [CrossRef] [Scilit]
  22. Lee, J.H.; Nam, Y.S.; Kim, Y.; Liu, Y.; Lee, J.; Yang, H. Real-time digital twin for ship operation in waves. Ocean. Eng. 2022, 266, 112867. [Google Scholar] [CrossRef] [Scilit]
  23. Wang, K.; Hu, Q.; Zhou, M.; Zun, Z.; Qian, X. Multi-aspect applications and development challenges of digital twin-driven management in global smart ports. Case Stud. Transp. Policy 2021, 9, 1298–1312. [Google Scholar] [CrossRef] [Scilit]
  24. Jedermann, R.; Praeger, U.; Lang, W. Challenges and opportunities in remote monitoring of perishable products. Food Packag. Shelf Life 2017, 14, 18–25. [Google Scholar] [CrossRef] [Scilit]
  25. Tang, P.; Postolache, O.A.; Hao, Y.; Zhong, M. Reefer container monitoring system. In Proceedings of the 2019 11th International Symposium on Advanced Topics in Electrical Engineering (ATEE), Bucharest, Romania, 28–30 March 2019; IEEE: New York, NY, USA; pp. 1–6. [Google Scholar]
  26. Singh, V.; Mathur, J.; Bhatia, A. A comprehensive review: Fault detection, diagnostics, prognostics, and fault modeling in HVAC systems. Int. J. Refrig. 2022, 144, 283–295. [Google Scholar] [CrossRef] [Scilit]
  27. Chen, Y.; Hu, Y.; Lin, G.; Zhang, Y.; Ye, S.; Shen, B. Evaluation of HVAC & refrigeration system fault behaviors and impacts: A systematic review. J. Build. Eng. 2025, 112, 113609. [Google Scholar] [CrossRef] [Scilit]
  28. Isermann, R. Model-based fault-detection and diagnosis–status and applications. Annu. Rev. Control 2005, 29, 71–85. [Google Scholar] [CrossRef] [Scilit]
  29. Keir, M.C.; Alleyne, A.G. Dynamic model-based fault detection and diagnosis residual considerations for vapor compression systems. In Proceedings of the 2006 American Control Conference, Minneapolis, MN, USA, 14–16 June 2006; IEEE: New York, NY, USA; p. 6. [Google Scholar]
  30. Poks, A.; Fallmann, M.; Fink, L.; Rinnofner, L.; Kozek, M. Fault detection and isolation for a secondary loop refrigeration system. Appl. Therm. Eng. 2023, 227, 120277. [Google Scholar] [CrossRef] [Scilit]
  31. Ovadia, Y.; Fertig, E.; Ren, J.; Nado, Z.; Sculley, D.; Nowozin, S.; Dillon, J.; Lakshminarayanan, B.; Snoek, J. Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift. Adv. Neural Inf. Process. Syst. 2019, 32. [Google Scholar] [CrossRef] [Scilit]
  32. Varshney, K.R. Engineering safety in machine learning. In Proceedings of the 2016 Information Theory and Applications Workshop (ITA), La Jolla, CA, USA, 31 January–5 February 2016; IEEE: New York, NY, USA; pp. 1–5. [Google Scholar]
  33. Cheliotis, M.; Lazakis, I.; Theotokatos, G. Machine learning and data-driven fault detection for ship systems operations. Ocean. Eng. 2020, 216, 107968. [Google Scholar] [CrossRef] [Scilit]
  34. Isermann, R. Fault-Diagnosis Systems: An Introduction from Fault Detection to Fault Tolerance; Springer Science & Business Media: Berlin/Heidelberg, Germany, 2005. [Google Scholar]
  35. Carvalho, T.P.; Soares, F.A.; Vita, R.; Francisco, R.D.P.; Basto, J.P.; Alcalá, S.G. A systematic literature review of machine learning methods applied to predictive maintenance. Comput. Ind. Eng. 2019, 137, 106024. [Google Scholar] [CrossRef] [Scilit]
  36. Zonta, T.; Da Costa, C.A.; da Rosa Righi, R.; de Lima, M.J.; Da Trindade, E.S.; Li, G.P. Predictive maintenance in the Industry 4.0: A systematic literature review. Comput. Ind. Eng. 2020, 150, 106889. [Google Scholar] [CrossRef] [Scilit]
  37. Yang, S.; Ordonez, J.C. Integrative thermodynamic optimization of a vapor compression refrigeration system based on dynamic system responses. Appl. Therm. Eng. 2018, 135, 493–503. [Google Scholar] [CrossRef] [Scilit]
  38. American Society of Heating, Refrigerating, & Air-Conditioning Engineers. ASHRAE Handbook: Refrigeration Systems and Applications; American Society of Heating, Refrigerating and Air Conditioning Engineers: La Crosse, WI, USA, 1986. [Google Scholar]
  39. Chandola, V.; Banerjee, A.; Kumar, V. Anomaly detection: A survey. ACM Comput. Surv. (CSUR) 2009, 41, 1–58. [Google Scholar] [CrossRef] [Scilit]
  40. Pang, G.; Shen, C.; Cao, L.; Hengel, A.V.D. Deep learning for anomaly detection: A review. ACM Comput. Surv. (CSUR) 2021, 54, 15. [Google Scholar] [CrossRef] [Scilit]
  41. Rudin, C. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nat. Mach. Intell. 2019, 1, 206–215. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  42. Arora, R.C. Refrigeration and Air Conditioning; PHI Learning Pvt. Ltd.: Delhi, India, 2012. [Google Scholar]
  43. Rasmussen, B.P.; Alleyne, A.G. Dynamic Modeling and Advanced Control of Air Conditioning and Refrigeration Systems; Air Conditioning and Refrigeration Center TR-244: Urbana, IL, USA, 2006. [Google Scholar]
  44. Bergman, T.L. Fundamentals of Heat and Mass Transfer; John Wiley & Sons: Hoboken, NJ, USA, 2011. [Google Scholar]
  45. Budiyanto, M.A.; Shinoda, T. Thermal simulation of the effect of solar radiation on the temperature increases on the refrigerated container walls. Int. J. Sustain. Eng. 2021, 14, 1229–1238. [Google Scholar] [CrossRef] [Scilit]
  46. Van Der Sman, R.G.M.; Verdijck, G.J.C. Model predictions and control of conditions in a CA-reefer container. In Proceedings of the VIII International Controlled Atmosphere Research Conference 600, Rotterdam, The Netherlands, 8–13 July 2001; pp. 163–171. [Google Scholar]
  47. Sørensen, K.K. Model Based Control of Reefer Container Systems. Ph.D. Thesis, Aalborg University, Aalborg, Denmark, 2013. [Google Scholar]
  48. Goodfellow, I.; Bengio, Y.; Courville, A. Deep feedforward networks. Deep. Learn. 2016, 1, 161–217. [Google Scholar]
  49. Sakurada, M.; Yairi, T. Anomaly detection using autoencoders with nonlinear dimensionality reduction. In Proceedings of the MLSDA 2014 2nd Workshop on Machine Learning for Sensory Data Analysis, Gold Coast, Australia, 2 December 2014; pp. 4–11. [Google Scholar]
  50. Malhotra, P.; Vig, L.; Shroff, G.; Agarwal, P. Long Short Term Memory Networks for Anomaly Detection in Time Series. In Proceedings of the 23rd European Symposium on Artificial Neural Networks, Computational Intelligence and Machine Learning (ESANN 2015), Bruges, Belgium, 22–24 April 2015; i6doc.com: Louvain-la-Neuve, Belgium, 2015; pp. 89–94. [Google Scholar]
  51. Hundman, K.; Constantinou, V.; Laporte, C.; Colwell, I.; Soderstrom, T. Detecting spacecraft anomalies using lstms and nonparametric dynamic thresholding. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, London, UK, 19–23 August 2018; pp. 387–395. [Google Scholar]
  52. Karpatne, A.; Atluri, G.; Faghmous, J.H.; Steinbach, M.; Banerjee, A.; Ganguly, A.; Shekhar, S.; Samatova, N.; Kumar, V. Theory-guided data science: A new paradigm for scientific discovery from data. IEEE Trans. Knowl. Data Eng. 2017, 29, 2318–2331. [Google Scholar] [CrossRef] [Scilit]
  53. Willard, J.; Jia, X.; Xu, S.; Steinbach, M.; Kumar, V. Integrating physics-based modeling with machine learning: A survey. arXiv 2020, arXiv:2003.04919. [Google Scholar]
  54. Raissi, M.; Perdikaris, P.; Karniadakis, G.E. Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations. J. Comput. Phys. 2019, 378, 686–707. [Google Scholar] [CrossRef] [Scilit]
  55. Li, Y.; Sun, J.; Fricke, B.; Im, P.; Kuruganti, T. Grey-box fault models and applications for low carbon emission CO2 refrigeration system. Int. J. Refrig. 2022, 141, 76–89. [Google Scholar] [CrossRef] [Scilit]
  56. Wang, Q.; Zhao, Z.; Wang, Z. Data-driven analysis of risk-assessment methods for cold food chains. Foods 2023, 12, 1677. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  57. Jedermann, R.; Nicometo, M.; Uysal, I.; Lang, W. Reducing food losses by intelligent food logistics. Philos. Trans. R. Soc. A Math. Phys. Eng. Sci. 2014, 372, 20130302. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Figure 1. Experimental setup and controller-level data acquisition for representative maritime refrigerated container units.
Figure 1. Experimental setup and controller-level data acquisition for representative maritime refrigerated container units.
Applsci 16 01887 g001
Figure 2. Architecture of the proposed hybrid anomaly detection framework.
Figure 2. Architecture of the proposed hybrid anomaly detection framework.
Applsci 16 01887 g002
Figure 3. Detection performance across fault types.
Figure 3. Detection performance across fault types.
Applsci 16 01887 g003
Figure 4. Time-to-detection (TTD) distributions across fault severity levels.
Figure 4. Time-to-detection (TTD) distributions across fault severity levels.
Applsci 16 01887 g004
Table 1. Controller-level measurements used in the proposed monitoring framework.
Table 1. Controller-level measurements used in the proposed monitoring framework.
SignalDescriptionPhysical RelevanceTypical Availability
Internal air temperature ( T in )Temperature of the air
inside the refrigerated
container
Cargo thermal integrityStandard reefer
Ambient temperature ( T amb )External air temperature surrounding the containerHeat exchange with surroundingsStandard reefer
Temperature setpoint ( T set )Target temperature defined by the operatorControl referenceStandard reefer
Compressor state ( u comp )Binary indication of compressor ON/OFF statusCooling system activityStandard reefer
Table 2. CAE configuration for reconstruction anomaly detection (compact).
Table 2. CAE configuration for reconstruction anomaly detection (compact).
ItemSetting
Windowing60 min window (3600 samples at 1 Hz)
Features [ T in , T amb , u comp , e T , T ˙ ]
EncoderConv1D (64, k = 5, s = 2) → Conv1D (32, k = 5, s = 2) → Conv1D (16, k = 5, s = 2) → Dense (8)
DecoderMirrored transposed conv blocks → output Conv1D (5, k = 1)
Score Reconstruction   error   E ( k ) (window-wise)
TrainingFault-free operation only
Table 3. Representative fault scenarios used for evaluation (aligned with Results).
Table 3. Representative fault scenarios used for evaluation (aligned with Results).
Fault TypeNatureSeverity LevelsPrimary Observable EffectOperational Relevance
F1: Cooling degradationGradualMild/Moderate/SevereReduced effective cooling capacity; temperature drift, increased duty cycleAgeing, fouling
F2: Sensor driftGradual+2 °C/+3 °C/+5 °CBiased temperature reading; inconsistency vs thermal expectationsCalibration issues
F3: Compressor failureAbrupt/IntermittentAbrupt/IntermittentLoss or partial loss of cooling; fast deviationCritical fault
F4: Refrigerant leakGradualSlow/Medium/FastProgressive capacity loss; subtle early dynamicsCommon latent fault
F5: Condenser blockageGradualMild/Moderate/SevereHeat rejection degradation; higher duty cycle and driftFouling/obstruction
Table 4. Detection performance metrics of the proposed hybrid framework across fault categories and severity levels.
Table 4. Detection performance metrics of the proposed hybrid framework across fault categories and severity levels.
Fault TypeSeverityTPR (%)TNR (%)Precision (%)F1-ScoreMean TTD (min)
F1: Cooling DegradationMild (25%)91.898.995.70.93728.4 ± 11.2
Moderate (50%)94.698.796.20.95418.3 ± 7.2
Severe (75%)98.298.596.80.9758.7 ± 3.4
F1 Average94.298.796.10.95118.3 ± 7.2
F2: Sensor DriftMild (+2 °C)96.497.893.10.9475.2 ± 2.3
Moderate (+3 °C)98.097.294.80.9643.4 ± 1.8
Severe (+5 °C)100.096.696.20.9811.8 ± 0.9
F2 Average98.197.294.80.9643.4 ± 1.8
F3: Compressor FailureAbrupt100.099.298.80.9941.8 ± 0.5
Intermittent100.099.098.20.9912.4 ± 0.8
F3 Average100.099.198.50.9922.1 ± 0.6
F4: Refrigerant LeakSlow (1%/h)87.299.196.80.91848.3 ± 18.7
Medium (2%/h)91.898.997.20.94434.7 ± 15.3
Fast (4%/h)94.898.797.60.96121.2 ± 9.4
F4 Average91.398.997.20.94134.7 ± 15.3
F5: Condenser BlockageMild (30%)93.698.694.80.94232.8 ± 14.2
Moderate (50%)96.898.496.20.96521.5 ± 9.8
Severe (70%)99.298.296.80.98012.3 ± 5.7
F5 Average96.798.495.90.96321.5 ± 9.8
Overall Performance-96.298.596.50.96218.3 ± 11.4
Note: “OVERALL PERFORMANCE” denotes pooled statistics across all replications and fault severities, which naturally results in larger variance. In contrast, Table 5 reports method-level mean time-to-detection with confidence intervals for baseline comparison.
Table 5. Comparison of baseline monitoring approaches.
Table 5. Comparison of baseline monitoring approaches.
MethodF1-ScoreMean TTD (min)95% CI TTD (min)False Positive Rate (/1000 h)CPU Utilisation (%)p-Value vs Hybrid
Threshold-based alarm0.723142.3 ± 45.2[128.4, 156.2]0.30.1<0.001
Physics-only digital twin0.85428.7 ± 11.4[25.5, 31.9]4.73.2<0.001
ML-only anomaly
detection
0.88125.1 ± 9.8[22.4, 27.8]8.21.8<0.001
One-Class SVM0.82135.4 ± 15.6[31.2, 39.6]7.12.6<0.001
ARIMA statistical model0.79267.4 ± 28.3[59.5, 75.3]2.10.8<0.001
Hybrid (proposed)0.95118.3 ± 7.2[15.9, 20.7]1.24.8-
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Vukšić, M.; Ćelić, J.; Ogrizović, D.; Perić Hadžić, A. Early Anomaly Detection in Maritime Refrigerated Containers Using a Hybrid Digital Twin and Deep Learning Framework. Appl. Sci. 2026, 16, 1887. https://doi.org/10.3390/app16041887

AMA Style

Vukšić M, Ćelić J, Ogrizović D, Perić Hadžić A. Early Anomaly Detection in Maritime Refrigerated Containers Using a Hybrid Digital Twin and Deep Learning Framework. Applied Sciences. 2026; 16(4):1887. https://doi.org/10.3390/app16041887

Chicago/Turabian Style

Vukšić, Marko, Jasmin Ćelić, Dario Ogrizović, and Ana Perić Hadžić. 2026. "Early Anomaly Detection in Maritime Refrigerated Containers Using a Hybrid Digital Twin and Deep Learning Framework" Applied Sciences 16, no. 4: 1887. https://doi.org/10.3390/app16041887

APA Style

Vukšić, M., Ćelić, J., Ogrizović, D., & Perić Hadžić, A. (2026). Early Anomaly Detection in Maritime Refrigerated Containers Using a Hybrid Digital Twin and Deep Learning Framework. Applied Sciences, 16(4), 1887. https://doi.org/10.3390/app16041887

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop