Next Article in Journal
Research on Rapid 3D Model Reconstruction Based on 3D Gaussian Splatting for Power Scenarios
Next Article in Special Issue
Seismic Disruption and Maritime Carbon Emissions for Sustainability in Maritime Transportation: A Natural Experiment from the 2023 Kahramanmaraş 7.6 Mwg Earthquake
Previous Article in Journal
Project-Based Learning in Geography and Its Impact on Developing Students’ Values, Attitudes and Pro-Environmental Behavior
Previous Article in Special Issue
Carbon Revenue Recycling: The Cornerstone of the Carbon Pricing Mechanism Within the Shipping Industry
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Decarbonizing Coastal Shipping: Voyage-Level CO2 Intensity, Fuel Switching and Carbon Pricing in a Distribution-Free Causal Framework

by
Murat Yildiz
1,
Abdurrahim Akgundogdu
2 and
Guldem Elmas
1,*
1
Department of Maritime Transportation Management Engineering, Faculty of Engineering, Istanbul University-Cerrahpasa, Avcılar, 34320 Istanbul, Turkey
2
Department of Electrical-Electronics Engineering, Faculty of Engineering, Istanbul University-Cerrahpasa, Avcılar, 34320 Istanbul, Turkey
*
Author to whom correspondence should be addressed.
Sustainability 2026, 18(2), 723; https://doi.org/10.3390/su18020723
Submission received: 1 December 2025 / Revised: 31 December 2025 / Accepted: 4 January 2026 / Published: 10 January 2026
(This article belongs to the Special Issue Sustainable Maritime Logistics and Low-Carbon Transportation)

Abstract

Coastal shipping plays a critical role in meeting maritime decarbonization targets under the International Maritime Organization’s (IMO) Carbon Intensity Indicator (CII) and the European Union Emissions Trading System (EU ETS); however, operators currently lack robust tools to forecast route-specific carbon intensity and evaluate the causal benefits of fuel switching. This study developed a distribution-free causal forecasting framework for voyage-level Carbon Dioxide (CO2) intensity using an enriched panel of 1440 real-world voyages across four Nigerian coastal routes (2022–2024). We employed a physics-informed monotonic Light Gradient Boosting Machine (LightGBM) model trained under a strict leave-one-route-out (LORO) protocol, integrated with split-conformal prediction for uncertainty quantification and Causal Forests for estimating heterogeneous treatment effects. The model predicted emission intensity on completely unseen corridors with a Mean Absolute Error (MAE) of 40.7 kg CO2/nm, while 90% conformal prediction intervals achieved 100% empirical coverage. While the global average effect of switching from heavy fuel oil to diesel was negligible (≈−0.07 kg CO2/nm), Causal Forests revealed significant heterogeneity, with effects ranging from −74 g to +29 g CO2/nm depending on route conditions. Economically, targeted diesel use becomes viable only when carbon prices exceed ~100 USD/tCO2. These findings demonstrate that effective coastal decarbonization requires moving beyond static baselines to uncertainty-aware planning and targeted, route-specific fuel strategies rather than uniform fleet-wide policies.

1. Introduction

The international maritime sector is undergoing a regulatory transformation. The International Maritime Organization (IMO) introduced the Carbon Intensity Indicator (CII) in 2023 as a mandatory operational rating scheme aimed at reducing greenhouse gas emissions from international shipping by at least 40% by 2030 compared to 2008 levels [1,2,3]. The CII grades vessels from A to E are based on annual CO2 per ton-nautical mile, with corrective action required for persistent D/E performance [1,2,4]. Non-compliance can lead to detention, fines, and loss of market access [5,6]. In parallel, the EU ETS extension to shipping and FuelEU Maritime are expected to impose carbon prices of several tens to about one hundred USD/tCO2 [2,7,8,9]. Together, these regimes require methods that forecast voyage-level emission intensity and credibly evaluate operational measures [10]. Even small efficiency gains reduce fuel and CO2 [11], and fuel choice—including cleaner alternatives to HFO—has become a key lever [12].
Maritime decarbonization supports the Sustainable Development Goals, notably Sustainable Development Goals (SDG) 13, SDG 9, and SDG 14 [13,14]. Because coastal and short-sea shipping connect hubs and communities, their emissions directly affect development and marine ecosystems [15], and coastal corridors can serve as practical testbeds where route-specific decision tools help align IMO measures with the Paris goals [16,17]. Achieving net-zero shipping by mid-century will require coupling operational improvements with high-level policy instruments (including coordinated carbon-pricing frameworks) and technology pathways such as alternative fuels and green-corridor deployment [18]. Consistent with the expert review-based agenda articulated by Govindan et al. [18], priority gaps include rigorous cost–benefit evaluation of port initiatives, the investment and techno-economic feasibility of onboard carbon capture and alternative fuels (with green shipping corridors as an adoption catalyst), and the design of carbon pricing under intertwined climate, economic, and socio-political constraints in ongoing IMO negotiations and in regional schemes such as the EU ETS [18]. Within this broader transition, our route- and season-specific risk and causal evidence provide an implementation-layer input for designing fair and effective corridor strategies. Near-coastal decarbonization is not a ship-only problem: emissions outcomes on short-sea corridors are shaped by port energy infrastructure, electrification of port-side services, and the availability of cleaner fuels and shore power; smart-port studies highlight ISO 50001-aligned energy management and digital monitoring as practical levers for reducing Greenhouse Gas (GHG) emissions and improving operational control [2], while complementary work emphasizes shore power (“cold ironing”) and port-side renewables as mitigation options for coastal networks [5]. Recent scoping evidence on renewable-powered seaport operations highlights the growing role of port Energy Management Systems and integrated renewables (e.g., solar, wind, and marine energy) in reducing seaport carbon footprints and enabling greener coastal logistics [19]. However, measure effectiveness is mixed, and switching from heavy fuel oil (HFO) to marine diesel oil (MDO) can be heavily confounded by operating conditions [20,21]; fuel use and carbon intensity respond nonlinearly to speed, draft, weather and hull condition, and coastal trades add shallow-water effects and maneuvering [22]. Naïve comparisons can therefore bias estimated switching benefits [21]. Causal methods are needed to isolate fuel effects, and findings should be interpreted alongside Energy Efficiency Design Index (EEDI) and Ship Energy Efficiency Management Plan (SEEMP) [23,24,25,26]; because fuel and routing choices have externalities, causal evidence can also support targeted governance interventions [27,28,29].
Machine learning (ML) adoption in maritime energy-efficiency modelling has accelerated [8]. Models such as neural networks and gradient boosting can predict fuel consumption and emission intensity within familiar operating envelopes, but most studies evaluate performance in-distribution rather than with route-wise holdout such as leave-one-route-out [11,30]. The harder regulatory problem is forecasting on unseen routes under distribution shift, where covariate distributions deviate from training and point accuracy can be misleading; classical regression-based confidence intervals also break down on heavy-tailed, multimodal voyage data [31]. Conformal prediction offers distribution-free intervals with finite-sample coverage guarantees [32]. Such robustness matters as targets tighten and operators need to separate weather-driven variability from decision-driven effects in energy-efficiency indicators [12,33]. Uncertainty-aware forecasting is therefore central to compliance and risk reporting rather than a purely technical add-on [2] because under-estimating emissions can trigger penalties and over-estimating can lead to unnecessary costs.
A further challenge is that many operational datasets omit deadweight tonnage (DWT), which is required for official CII computation [34]. In practice, operators often rely on transport-work-agnostic targets such as voyage-level CI (kg CO2/nm), but modeling pipelines should remain explicitly “CII ready” so the target can later be swapped to Annual Efficiency Ratio (AER)/CII when design parameters become available [2,35]. Because intensity metrics influence sustainable-finance signals and carbon-pricing credibility, unstable measures can distort incentives and reporting [36,37]. This risk is amplified when charges or performance thresholds hinge on unobserved factors [37]. Architectures compatible with current indicators and future CII/AER metrics therefore support a smoother transition [38]. Recent work on maritime decarbonization and operational analytics has advanced emissions forecasting and compliance-oriented indicators, but it is still largely evaluated under in-distribution settings and with limited deployment realism for new corridors [10,11]. In parallel, empirical assessments of operational measures (including fuel switching) often remain associational, making estimates sensitive to route- and condition-driven confounding [21,23]. These considerations expose three gaps: unseen-route generalization, distribution-free uncertainty, and causal identification of fuel-switching effects under confounding [21,39].
Existing ML-based maritime emissions studies typically prioritize point prediction accuracy under random or in-distribution train–test splits and rarely test deployment on entirely new corridors where covariate distributions shift [10,11]. In parallel, uncertainty is often omitted or approximated with parametric confidence intervals whose assumptions are brittle in heavy-tailed voyage data and under route-wise domain shift [31]. Moreover, empirical assessments of operational measures such as fuel switching frequently remain correlational, so estimated benefits can be confounded by endogenous speed, loading, and routing decisions [21,22,23]. In contrast, we propose a unified pipeline that combines leave-one-route-out evaluation for unseen-route generalization, distribution-free conformal prediction intervals for finite-sample risk quantification [32,40] and causal estimation (Double/Debiased Machine Learning and Causal Forests) for heterogeneous fuel-switching effects [41,42]. These outputs are further operationalized into an exceedance-risk map and Total Cost Intensity analysis under carbon pricing, providing decision-relevant evidence for corridor-level decarbonization planning [2,7].
From a sustainability and governance standpoint, these gaps translate into practical blind spots. Without tools that generalize to new corridors and quantify risk under deep uncertainty, operators cannot reliably plan compliance margins, and regulators cannot distinguish genuine efficiency improvements from statistical artifacts. Likewise, without causal identification, fuel-switching policies may be rewarded or penalized for the wrong reasons, leading to misallocated investment and distorted signals under carbon pricing. Addressing all three gaps together is therefore essential for decision support that is both operationally useful and policy credible. This is particularly acute in coastal services, where corridors change, operating regimes shift quickly, and data constraints are common.
We analyzed 1440 voyages across four Nigerian coastal routes (2022–2024) from a single operator’s noon-reports and Automatic Identification System (AIS) records, released as a public ship fuel-efficiency dataset on Kaggle [43]. From this base, we built a voyage-level panel and model emission intensity as CI (kg CO2/nm) using engineered operational, environmental, and cost covariates that capture speed and distance effects, route context, weather proxies, and maneuvering conditions. CI is forecast with a physics-informed monotonic LightGBM model under a strict leave-one-route-out protocol; split conformal prediction yields finite-sample intervals on unseen routes; and HFO-to-diesel effects are estimated with Double/Debiased Machine Learning and Causal Forests [41,42]. We then operationalize uncertainty via an Emission Intensity Exceedance Risk metric and integrate Total Cost Intensity (TCI) under alternative carbon prices to support route- and season-specific decisions.
Study design and research questions. This study is an observational, retrospective voyage-level panel analysis of Nigerian short-sea coastal operations (N = 1440 voyages across four routes, 2022–2024) that integrates (i) unseen-route emissions forecasting, (ii) distribution-free uncertainty quantification, and (iii) causal estimation of HFO-to-diesel switching effects under operational confounding. We address three pre-specified research questions:
  • RQ1 (Forecasting): Can voyage-level CO2 intensity be predicted on completely unseen coastal routes under a strict leave-one-route-out design?
  • RQ2 (Uncertainty): Can we provide distribution-free prediction intervals with valid coverage on unseen routes to support uncertainty-aware compliance and planning decisions?
  • RQ3 (Causal fuel switching): What is the causal effect of switching from HFO to diesel on emission intensity, and how heterogeneous is this effect across route–season operating conditions?
The study makes four contributions. (1) Unseen-route emission forecasting: a physics-informed monotonic LightGBM model with leave-one-route-out training delivers robust predictions on unseen coastal corridors while preserving hydrodynamic monotonicities. (2) Distribution-free risk quantification: split conformal prediction provides 90% prediction intervals with finite-sample guarantees and yields an Emission Intensity Exceedance Risk metric for decision support. (3) Causal assessment of fuel switching: Double/Debiased Machine Learning and Causal Forests show a near-zero average HFO-to-diesel effect but substantial route–season heterogeneity. (4) Economic integration via TCI: emissions outcomes are linked to Total Cost Intensity under alternative carbon prices, clarifying when selective diesel adoption becomes cost-effective. Taken together, the framework provides a practically implementable toolkit for fleet managers and regulators to forecast emissions, quantify compliance risk, and evaluate operational measures in a route- and season-specific way.
The rest of this paper is organized as follows. Section 2 presents the materials and methods, including data construction, descriptive context, and the predictive, uncertainty-quantification, and causal components of the proposed pipeline. Section 3 reports empirical results on unseen-route forecasting performance, uncertainty coverage, causal fuel-switching effects (including heterogeneity), and cost implications under carbon pricing. Section 4 discusses interpretation, policy and operational implications, and limitations. Section 5 concludes with practical recommendations and directions for future work.

2. Materials and Methods

All computations were performed in Python (v3.12.11; Python Software Foundation, Wilmington, DE, USA). Model training and forecasting were conducted using LightGBM (v4.6.0) and scikit-learn (v1.6.1). Model interpretability analyses were carried out using SHAP (v0.48.0). Figures and visualizations were generated using Matplotlib (v3.10.3). Causal estimation, including Double/Debiased Machine Learning (DML) and Causal Forests for heterogeneous treatment effects, was implemented using EconML (v0.16.0).
Methods overview. Our workflow proceeds in four steps. First, we constructed an enriched voyage-level panel from operational noon reports and AIS records and defined voyage-level CO2 intensity per nautical mile (CI, kg CO2/nm) as the primary outcome. Second, we trained a physics-informed monotonic LightGBM model under a strict leave-one-route-out (LORO) design to evaluate deployment on completely unseen coastal corridors. Third, we wrapped point forecasts with split conformal inference to obtain distribution-free prediction intervals with finite-sample coverage on unseen routes. Fourth, we estimated the causal effect of switching from HFO to diesel using Double/Debiased Machine Learning Average Treatment Effect (ATE) and Causal Forests heterogeneous Conditional Average Treatment Effect (CATE), relying on overlap/common support diagnostics to support causal interpretation under observed confounding. Finally, we operationalized outputs for decision support through uncertainty-aware risk screening and cost-intensity concepts under carbon pricing scenarios.
Mapping from research questions to methods:
RQ1 → Monotonic LightGBM forecasting under LORO evaluation.
RQ2 → Split conformal prediction intervals and empirical coverage assessment under LORO.
RQ3 → Double/Debiased Machine Learning (DML) for ATE and Causal Forests for heterogeneous CATE, supported by overlap/propensity diagnostics.

2.1. Data Source, Study Setting, and Voyage Panel Construction

This study analyzed a single-operator panel of 1440 completed coastal voyages performed by general cargo and coastal tanker vessels across four Nigerian short-sea routes (Port Harcourt–Lagos, Lagos–Apapa, Escravos–Lagos, and Warri–Bonny) over the period of January 2022 to December 2024.
Figure 1 summarizes the Nigerian coastal study network, showing the four analyzed routes and their associated ports.
The underlying observations originated from the operator’s standard operational reporting system, combining daily noon reports (fuel logs, engine settings, and operational notes) with Automatic Identification System (AIS) records that provide position and speed-over-ground information. Prior to public release, the data owner anonymized and aggregated these operational records and made them available as the “Ship Fuel Efficiency—Kaggle” dataset [43].
In the present work, the Kaggle dataset was used as the operational starting point, and the modeling dataset was constructed as an enriched voyage-level panel. Specifically, we retained the full set of 1440 real-world Nigerian coastal voyages and derived additional variables tailored to corridor-level short-sea operations and the methodological requirements of unseen-route forecasting and causal inference. These engineered variables included route-relative difficulty metrics (e.g., route_difficulty_score), deviation-based indicators (route_median_ratio and intensity_gap), and season- or weather-normalized emission-intensity measures that support both predictive modeling and treatment-effect estimation. The resulting dataset represents a coherent operational environment in the Gulf of Guinea characterized by homogeneous reporting conventions and a shared commercial context. Additional documentation of the source dataset and the derived voyage-level panel is provided in Appendix A and Appendix B.

2.2. Variables, Definitions, and Preprocessing

The voyage-level panel includes route and time identifiers (route_id, month, season_bin), operational variables (distance_nm, fuel_type, fuel_consumption), environmental conditions (weather_category), outcomes (CO2 emissions, CI_kg_per_nm), cost variables (energy_cost and carbon-cost components), and engineered route-relative features (route_difficulty_score, route_median_ratio, intensity_gap). Appendix A and Appendix B provide the complete variable dictionaries and construction steps for both the original and derived datasets.
The main variable groups used in the analyses were:
(i)
Route and time: route identifiers (route_id), seasonal bins (season_bin: Winter, Spring, Summer, Autumn), and cyclic temporal encodings that capture intra-annual variation in operating conditions.
(ii)
Operational conditions: voyage distance in nautical miles (distance_nm) and fuel type ( f u e l _ t y p e H F O , D i e s e l ) .
(iii)
Environmental factors: discrete weather states (weather_category ∈ {Calm, Moderate, Stormy}), constructed from Beaufort-scale logs and significant wave height (Hs) consistent with the operator’s internal reporting conventions for coastal voyages in the Gulf of Guinea. Table 1 summarizes the classification thresholds.
(iv)
Outcomes and cost drivers: voyage-level fuel consumption, total CO2 emissions, and energy cost. Carbon-cost components and Total Cost Intensity (TCI) measures were derived for the scenario-based economic analysis and reported in the Section 3.
Primary outcome. The target variable is Voyage-Level Emission Intensity (CI), defined as total CO2 emissions per nautical mile:
C I = T o t a l   C O 2   E m i s s i o n s   ( kg ) D i s t a n c e   ( nm )
and reported throughout in kg CO2/nm. Deadweight tonnage (DWT) is unavailable in the dataset, preventing direct computation of official CII in g CO2/(dwt·nm). While the IMO’s CII is transport-work-based, CI expressed in kg CO2/nm is widely used in coastal shuttle operations where repeated legs and narrower load variation make per-distance performance monitoring operationally meaningful. Importantly, the modeling pipeline was designed to remain “CII-ready”: once DWT or cargo-mass data become available, the same statistical architecture (feature engineering, prediction, uncertainty quantification, and causal adjustment) can be applied to AER/CII-type outcomes without changing the methodological core.
Preprocessing and data quality. All variables were derived from the operator’s noon-report and AIS workflow. Noon reports are compiled by deck officers once per day or at key voyage stages and include manually entered fuel, weather, and engine settings; AIS provides positional and speed information. As in most operational datasets, occasional inconsistencies and outliers cannot be ruled out. We therefore applied basic plausibility checks and range filters prior to analysis and relied on modeling choices that are robust to residual measurement noise. Load condition (ballast versus laden) was not directly observed; its influence is partially absorbed through route, distance, and seasonal indicators, but transport work was not reconstructed from the available fields.

2.3. Descriptive Statistics and Operational Context

Table 2 summarizes key continuous voyage-level variables (N = 1440). Voyage distances ranged from 20.1 to 498.6 nm (mean 151.8; median 123.5 nm), reflecting fixed-route coastal shuttle operations. The mean emission intensity was 78.9 kg CO2/nm (median 79.2 kg CO2/nm), indicating a stable intensity envelope despite substantial heterogeneity in distance and weather exposure. Total CO2 emissions per voyage ranged from 0.6 to 71.9 tons (mean 13.4; median 8.4 tons). Voyage-level energy cost (bunker expenditure only) ranged from 142 to 14,666 USD (mean 2429; median 1501 USD).
Figure 2 provides an overview of the dataset structure, including route frequencies, seasonal balance (25% per season), fuel shares (Diesel 62.4%, HFO 37.6%), weather distribution (Calm 35.8%, Moderate 32.1%, Stormy 32.1%), and the distance and CI distributions. Operationally, the vessels follow repeated shuttle patterns connecting refinery terminals, river ports, and offshore loading/discharging areas, which implies frequent transitions between open-coast steaming and constrained channel approaches. The four corridors differ in navigational and environmental difficulty: Port Harcourt–Lagos is the longest and involves exposure to bar crossings, offshore currents, and dense traffic; Lagos–Apapa is a very short intra-port leg dominated by low-speed maneuvering and hoteling; Escravos–Lagos and Warri–Bonny involve river-mouth transits and shallow, constrained fairways that increase sensitivity to weather and operational constraints. These corridor characteristics motivate both route-aware predictive evaluation and confounding-aware causal estimation.

2.4. Evidence of Operational Confounding

Figure 3 visualizes emission intensity (CI) versus voyage distance, colored by fuel type. A pronounced negative relationship can be observed: longer voyages typically have a higher share of steady cruising relative to maneuvering and port approaches, leading to lower CI. At short distances—typical of intra-port legs such as Lagos–Apapa—CI is elevated because a substantial share of fuel is consumed in low-speed maneuvers, acceleration–deceleration cycles, and hoteling in confined waters. Although Diesel voyages often exhibit lower CI than HFO voyages, there is substantial overlap between fuels across distances, seasons, and weather states. This overlap indicates that fuel choice is endogenous to operational context: naïve comparisons that ignore distance, route geometry, shallow-water constraints, and weather would yield biased estimates of “fuel switching” benefits. Accordingly, the causal component of our framework explicitly conditions on these covariates and route-relative metrics when estimating the effect of switching from HFO to Diesel.

2.5. Predictive Modeling for Unseen Routes

The forecasting task is to predict voyage-level emission intensity on a completely unseen coastal corridor without retraining, reflecting a deployment-realistic scenario for route-specific compliance planning. We used a Light Gradient Boosting Machine (LightGBM) regressor [44] due to its strong performance on heterogeneous tabular data and its ability to capture nonlinear interactions. Categorical identifiers (e.g., ship_id and route_id) are treated as categorical features so that the model can learn vessel- and corridor-associated patterns without requiring manual encoding.
Leakage-free feature engineering. To generalize to unseen routes, route-level statistics must be constructed without leaking information from the held-out corridor. Naively including route means or medians computed on the full dataset would allow the model to implicitly “recognize” the test route. We therefore implemented a strict out-of-fold feature engineering protocol aligned with leave-one-route-out validation. For each fold, one route is held out entirely; all group-derived summaries are computed only on the remaining routes and then projected to voyage level as relative metrics. Key engineered features include route_median_ratio, route_difficulty_score, and intensity_gap, which summarize operational difficulty as deviations from appropriate baselines rather than absolute route identifiers. To explicitly capture operational stress, these features are constructed as relative deviations. Specifically, route_median_ratio is computed as the ratio of a voyage’s emission intensity to the route-specific median (e.g., a ratio of 1.2 indicates an intensity 20% higher than the corridor norm, likely due to weather or maneuvering). The route_difficulty_score represents the absolute deviation from the route mean ( C I v o y a g e   μ r o u t e ) , retaining the magnitude of intensity spikes, while the intensity_gap measures the deviation from the season-specific median, isolating weather-driven anomalies from baseline seasonal trends. These relative metrics allow the model to learn ‘how difficult’ a voyage is compared to its peers, ensuring that predictions rely on transferable physical signals rather than static route identifiers.
Physics-informed monotonicity. To embed domain knowledge, we imposed a monotonicity constraint on distance_nm, requiring predicted CI to be a non-increasing function of distance (monotone constraint = −1). In coastal shuttle operations, longer voyages tend to allocate more time to efficient cruising and less to fuel-intensive maneuvering and hoteling; thus, CI should not increase with distance, all else being equal. Other covariates (weather, season, route-relative metrics) remained unconstrained to allow for nonlinear interactions.
Hyperparameter configuration. The LightGBM model was configured with n _ e s t i m a t o r s   =   2000 , l e a r n i n g _ r a t e   =   0.01 , m a x _ d e p t h   =   12 , and m i n _ c h i l d _ s a m p l e s   =   20 , with remaining parameters set to defaults unless otherwise stated. This configuration balances capacity and regularization to improve robustness under corridor-level distribution shifts.
Unseen-route evaluation (LORO). We evaluated predictive performance under a strict leave-one-route-out (LORO) protocol: in each fold, one corridor is held out as the test set, while the remaining routes are used for leakage-free feature engineering, model fitting, and uncertainty calibration. Performance metrics (MAE, Root Mean Square Error (RMSE), and R2) are computed on the held-out route only and aggregated across folds.

2.6. Distribution-Free Uncertainty Quantification

Point forecasts alone are insufficient for regulatory risk management; decision-makers require uncertainty ranges that remain valid under distribution shifts across corridors. Classical confidence intervals rely on assumptions (e.g., normality and homoscedasticity of residuals) that are often violated in heavy-tailed, multimodal operational voyage data. We therefore constructed distribution-free prediction intervals using split-conformal prediction, which provides finite-sample coverage guarantees [32,40].
Let f(X) denote the fitted LightGBM predictor. Within the LORO framework, absolute residuals are computed on each held-out route as:
R i = Y i f X i ,
where Yi is the observed emission intensity for voyage i. Residuals from all folds are pooled into a calibration set {Ri}i = 1 … N to characterize prediction error on unseen corridors. For a target coverage level 1 − α = 0.90, we compute the empirical (1 − α)-quantile q ^ of the pooled residuals. For a new voyage with covariates Xnew on a previously unseen route, the split-conformal prediction interval is:
C X n e w = f X new q ^ , f X new + q ^ ,
Under exchangeability within the operational domain, this construction guarantees that the true emission intensity Ynew lies within the interval with probability at least 90%, regardless of the underlying noise distribution. In practice, pooling residuals across folds yields conservative intervals because the calibration incorporates worst-case errors observed on challenging corridors, reducing the risk of underestimation in compliance-oriented settings.

2.7. Causal Inference for Fuel Switching

Beyond forecasting and uncertainty, operational and policy decisions require causal evidence: under comparable conditions, when does switching from HFO to Diesel reduce emission intensity, and when does it increase it? We estimated causal effects in a potential outcomes framework. Let T be a binary treatment indicator with T   =   1   for Diesel and T = 0 for HFO, and let Y denote observed emission intensity (CI, kg CO2/nm). Each voyage has potential outcomes Y(1) and Y(0), corresponding to the emission intensity that would be observed under Diesel or HFO, respectively. The primary estimands are the Average Treatment Effect (ATE, defined as the arithmetic mean difference) and Conditional Average Treatment Effects (CATE, defined as the conditional arithmetic mean of the treatment effect) of Diesel relative to HFO.
Treatment assignment reflects operational fuel policies rather than randomized switching. Many vessel–route pairs are effectively mono-fuel; however, some vessels operate multiple corridors and appear with different fuels across routes. Accordingly, causal effects are interpreted as voyage-level contrasts between Diesel and HFO under comparable route, season, and weather conditions rather than as within-vessel before–after switches on the same leg. Identification relies on overlap (common support) in the covariate space (distance, season, weather category, ship attributes, and engineered route-relative metrics). Propensity-score diagnostics are visualized in Figure 4.
ATE via Double/Debiased Machine Learning (DML). We estimated the ATE using Double/Debiased Machine Learning with cross-fitting [41]. Flexible nuisance models were fitted for the outcome m(X) ≈ E[Y|X] and the propensity score e(X) ≈ P(T = 1|X). Cross-fitting yields orthogonalized residuals that reduce regularization bias when estimating the treatment effect.
CATE via Causal Forests. To capture heterogeneity in fuel-switching benefits, we estimated CATEs using Causal Forests with 1000 trees, following [42]. For a covariate profile X = x, the CATE is:
τ x = E Y 1 Y 0 X = x ,
For interpretability and to avoid data dredging, we summarized heterogeneous effects primarily at the route × season level, focusing on cells with adequate overlap between fuels.
Covariate set and overlap diagnostics. The covariate set X includes core operational variables (distance_nm, season_bin, weather_category), ship-level descriptors (ship_type), and engineered route-relative metrics (route_difficulty_score, intensity_gap). Propensity scores are estimated via logistic regression:
e X = P T = 1 X
and Figure 4 suggests non-zero overlap within the observed covariate space; to assess robustness, we additionally report the ATE/CATE estimates on trimmed samples restricting e X to [0.10, 0.90] and [0.20, 0.80] and we provide complementary overlap/balance diagnostics in Appendix C.

2.7.1. Identification Strategy and Operational Pathways

Fuel choice in coastal shuttle operations is an operational decision and is therefore structurally endogenous to voyage planning. In practice, operators may select Diesel versus HFO based on bunkering availability and contracting, port/terminal constraints, safety-driven maneuvering requirements in constrained approaches, cost planning, and maintenance practices. In our dataset, these operational pathways are proxied by the observed covariate set X capturing route geometry and stress (distance_nm, route_difficulty_score, route_median_ratio, intensity_gap), environmental conditions (weather_category), seasonal operating regimes (season_bin and the seasonal encodings used in the feature set), and vessel class (ship_type). Under the potential-outcomes framework, we interpreted the estimated ATE/CATE as conditional causal contrasts only under three standard assumptions: (i) consistency/SUTVA (each voyage’s outcome corresponds to its observed fuel_type and is not affected by other voyages), (ii) conditional ignorability of fuel choice given X (unconfoundedness), and (iii) positivity (overlap) within the observed covariate space. While unconfoundedness cannot be proven in observational data, the identification argument here is that conditioning on these operational proxies removes the dominant drivers that jointly affect fuel choice and CI_kg_per_nm. Building on the DML formulation above, Double/Debiased Machine Learning targets selection bias by separately learning the treatment-assignment mechanism e ( X )   =   P ( T   =   1 | X ) and the outcome mechanism m ( X )   =   E [ Y | X ] , and then orthogonalizing the residual variation via cross-fitting to reduce bias from flexible nuisance models [41]. To satisfy the strict overlap assumption and mitigate endogeneity, we inspected the propensity score distributions and applied trimming to exclude voyages with extreme treatment probabilities (outside the [0.05, 0.95] range), ensuring that comparisons between HFO and Diesel were made only within a valid common support region. The propensity model in Equation (5) provides an empirical summary of the assignment mechanism and helps detect near-deterministic fuel policies for specific operational profiles. Positivity was assessed empirically through overlap in estimated propensity scores between treated (Diesel) and control (HFO) voyages; Figure 4 suggests non-zero overlap within the observed common-support region (e.g., comparable profiles such as a stormy 100 nm voyage). However, because fuel policies may still be partially structured at corridor and vessel levels, all causal contrasts were interpreted within the observed support (and conditional on X), and we avoided extrapolation to domains where one fuel is rarely or never observed. Appendix C reports additional overlap, balance, trimming, and sensitivity diagnostics that make the supported domain of the causal contrasts explicit.

2.7.2. Sensitivity to Unobserved Confounding

Even after conditioning on X, unobserved factors such as load condition, hull/propeller fouling, draft/trim, and congestion or currents may induce residual confounding in observational voyage data. We therefore complemented the main estimates with a partial-R2-based omitted-variable sensitivity framing that characterizes how strong an unmeasured confounder would need to be (in incremental explanatory power for both the treatment and outcome models) to materially change the ATE conclusions; the calculations are summarized in Table A8. In addition, we performed a leave-one-covariate-block-out robustness check, re-estimating ATE/CATE after removing (i) distance-related terms, (ii) weather/season terms, (iii) ship_type, and (iv) engineered route-stress proxies (route_difficulty_score, route_median_ratio, intensity_gap). Any instability across these perturbations was reported and interpreted as evidence of fragility, whereas stability supports the interpretation of the estimates as conditional contrasts within the observed support. Quantitative results for these checks are reported in Appendix C.

2.8. Summary of Outputs and Reproducibility Notes

The unified pipeline yields five decision-relevant outputs: (i) point forecasts of voyage-level emission intensity on unseen corridors (LightGBM under LORO), (ii) distribution-free prediction intervals with finite-sample coverage guarantees (split-conformal), (iii) inputs for risk-oriented screening based on the uncertainty envelope (e.g., exceedance-type indicators derived from conformal bounds), (iv) causal estimates of fuel switching effects at both the global (ATE) and heterogeneous (CATE) levels (DML and Causal Forests), and (v) an economic integration layer based on cost intensity concepts under alternative carbon-pricing scenarios. Appendix A documents the source dataset structure and limitations, while Appendix B details the construction of the derived voyage-level panel and provides the full variable dictionary, enabling transparent replication of feature construction and modeling inputs.

3. Results

Results at a glance:
  • Under strict Leave-One-Route-Out evaluation, unseen-route emission-intensity forecasting achieved a MAE ≈ 40.7 kg CO2/nm (Table 3), while R2 can be negative due to corridor-level level shifts.
  • The 90% split-conformal prediction intervals achieved 100% empirical coverage on held-out routes, supporting uncertainty-aware compliance planning (Section 3.1).
  • SHapley Additive exPlanations (SHAP) analysis indicates the model relies on route-relative difficulty metrics and a physics-consistent distance effect, supporting robust generalization beyond route identifiers (Section 3.2).
  • In the causal layer, DML estimated a near-zero global average Diesel-vs.-HFO effect (ATE = −0.072 kg CO2/nm, 95% CI [−0.155, 0.011]), while Causal Forests revealed strong route–season heterogeneity (CATEs from −74 g to +29 g CO2/nm) (Section 3.3).
  • A conformal-interval–based Exceedance Risk Map indicates persistently high probabilities (≈81.5–97.3% across route–season cells) that the upper 90% prediction bound exceeds historical HFO median intensity, prioritizing corridors for SEEMP Part III interventions (Section 3.4).
  • In the economic integration, Total Cost Intensity (TCI) results indicate that selective Diesel use on demanding corridors becomes economically attractive when carbon prices approach ~100 USD/tCO2, with further convergence at higher carbon prices (Section 3.5).
This section reports the empirical performance of the proposed causal–conformal framework. We first evaluate the predictive accuracy of the monotonic LightGBM model and its distribution-free uncertainty quantification under a strict leave-one-route-out (LORO) scheme. We then examine how fuel switching from HFO to Diesel affects emission intensity, both on average and across heterogeneous route–season cells. Finally, we translate these environmental results into compliance risk and cost outcomes through an Emission-Intensity Exceedance Risk Map and a Total Cost Intensity (TCI) analysis under alternative carbon pricing scenarios.

3.1. Unseen-Route Forecasting and Uncertainty Quantification

The monotonic LightGBM model exhibited robust generalization when deployed on completely unseen coastal corridors. Across the four leave-one-route-out (LORO) folds, the model attained a global mean absolute error ( M A E )   of   40.7   kg   CO 2 / nm on the held-out routes. Route-level performance metrics are summarized in Table 3.
The highest errors occurred on Escravos–Lagos ( M A E     46.95   kg   CO 2 / nm ,   R M S E     54.72   kg   CO 2 / nm ) and Lagos–Apapa ( M A E     45.00   kg   CO 2 / nm ,   R M S E     52.58   kg   CO 2 / nm ) , reflecting the high operational variability of short, congested river passages and intra-port shuttles where steady cruising is rare. In contrast, the longer Port Harcourt–Lagos route displayed the lowest error ( M A E     33.73   kg   CO 2 / nm ,   R M S E     41.31   kg   CO 2 / nm ) , consistent with a larger share of stable steaming in open-coast conditions. Averaging over folds yielded a global M A E   of   40.68   kg   CO 2 / nm , which is numerically consistent with the rounded value of 40.7 kg CO2/nm reported at the fleet level.
The negative R2 values observed across all folds do not indicate that the model is uninformative; rather, they reflect the severity of the distribution shift inherent to the unseen-route setting. In the LORO protocol, each test fold corresponds to a route whose mean emission intensity can differ substantially from the training-set average. Using the training-route mean as a baseline therefore induces a strong level shift, against which even well-calibrated predictions can yield negative R2 when evaluated in the usual way. In this context, MAE and RMSE remain the more interpretable measures of performance, while negative R2 primarily signals that unseen routes possess different absolute intensity baselines rather than that the model fails to learn meaningful structure. Indeed, the route-level results in Table 3 show that operational difficulty is still ranked in a manner consistent with the underlying geometry and stress profile of each corridor, motivating the distribution-free adjustment provided by the split-conformal framework.
For uncertainty quantification, the LightGBM predictor was wrapped in a 90% split-conformal scheme calibrated under the same LORO protocol. Across all test folds, the split-conformal prediction intervals achieved 100% empirical coverage, with a calibrated interval half-width q ^   =   81.2   kg   CO 2 / nm . These bands are intentionally conservative: the pooled calibration residuals are strongly influenced by difficult segments such as Lagos–Apapa and Escravos–Lagos, where short distances, frequent maneuvering and shallow-water effects generate heavy-tailed errors. The resulting prediction intervals therefore expand to accommodate worst-case covariate shifts on completely unseen corridors. From a regulatory perspective, this conservatism is desirable, as it reduces the risk of underestimating emissions and ensures that regulatory compliance risks are not an understated critical requirement in safety-critical maritime reporting where false negatives carry significant penalties. These intervals underpin the Emission Intensity Exceedance Risk Map discussed in Section 3.4.
Key takeaway: Under strict unseen-route deployment, point-prediction metrics are constrained by corridor-level distribution shifts, whereas conformal intervals provide deployment-relevant uncertainty bounds for compliance-oriented planning.

3.2. Model Interpretability and Physics-Informed Validation

To ensure that the predictive model captured transferable operational physics rather than route-specific artifacts, we analyzed its behavior using SHAP (SHapley Additive Explanations). The global feature-importance plot in Figure 5 shows that route-relative metrics, in particular route_median_ratio and route_difficulty_score (and, to a lesser extent, intensity_gap), dominated the prediction of emission intensity on unseen routes. This dominance indicates that the model relies on deviations from appropriate baselines—how “difficult” a voyage is relative to typical voyages—rather than memorizing absolute route identifiers, which is a key requirement for robust unseen-route generalization.
The SHAP summary plot in Figure 5 provides a direct validation of the physics-informed monotonicity constraint imposed on voyage distance. High values of distance_nm (red points) are consistently associated with negative SHAP values, meaning that longer voyages systematically push the predicted CI downward on a per-nautical-mile basis. This pattern is precisely what the hydrodynamics of coastal operations suggest: longer passages allocate a larger share of time to efficient cruising and a smaller share to fuel-intensive maneuvering and hoteling. The fact that SHAP aligns with the imposed non-increasing constraint on distance, even for routes that were never seen in training, reinforces confidence that the model has internalized a physically meaningful distance–intensity relationship rather than exploiting spurious correlations.
Taken together, the SHAP analysis and the monotonicity constraint show that the LightGBM model respects basic maritime physics while leveraging route-relative features to accommodate heterogeneity in geometry, weather, and congestion across the Nigerian coastal network.
Key takeaway: SHAP patterns confirm that the model relies on route-relative stress features and a physically consistent distance–intensity relationship, supporting transferability beyond route memorization.

3.3. Average Versus Heterogeneous Effects of Fuel Switching

These causal contrasts were interpreted within the observed overlap/common-support region of this single-operator dataset; r o u t e × s e a s o n cells with limited support were not over-interpreted. Complementary overlap, balance, trimming, and sensitivity diagnostics that clarify the supported domain are reported in Appendix C.
The effect of switching from HFO to Diesel on voyage-level emission intensity was next examined after adjusting for operational confounding. Using Double/Debiased Machine Learning, the Average Treatment Effect (ATE, arithmetic mean) of Diesel relative to HFO was estimated as −0.072 kg CO2/nm, with a 95% confidence interval of [−0.155, 0.011]. Because this interval includes zero (p > 0.05), the global average effect is statistically indistinguishable from zero. In other words, when averaged across all voyages in the dataset, no guarantee is provided that a blanket, fleet-wide switch from HFO to Diesel would lead to a systematic reduction in emission intensity.
This absence of a strong global effect does not imply that fuel switching is irrelevant; rather, it indicates that positive and negative effects coexist and largely cancel out in aggregate. The Causal Forest analysis revealed substantial treatment-effect heterogeneity across route–season cells. The CATE map in Figure 6 shows that emission reductions concentrated in specific high-stress operational regimes, while other segments exhibited neutral or even adverse outcomes. The CATE map in Figure 6 illustrates these contrasts: on the Port Harcourt–Lagos route during Autumn, switching from HFO to Diesel reduced emission intensity by 74 g CO2/nm, corresponding to a substantial improvement in CI on a long, hydrometeorologically demanding passage. In contrast, on the short Lagos–Apapa route during Summer, Diesel increased emissions by 29 g CO2/nm, reflecting the inefficiencies of Diesel engines under low-load, maneuvering-dominated conditions.
These heterogeneous patterns are consistent with the underlying route geometry and operating environment. On Port Harcourt–Lagos, particularly in Autumn, vessels encounter long stretches of steady steaming interspersed with current-affected bar crossings. Such conditions push the main engine closer to its optimal loading regime and allow Diesel’s thermodynamic advantages over HFO to translate into the observed emission savings of up to 74 g CO2/nm. In contrast, the Lagos–Apapa leg is a very short intra-port shuttle; it is dominated by low-speed maneuvering, short acceleration–deceleration cycles, and extended periods at or near idle power. Under these part-load and stop–start conditions, Diesel engines operate far below their design efficiency bands and may exhibit higher specific fuel consumption than HFO, explaining the positive CATE values (emission penalties) observed on this segment.
Overall, the joint ATE–CATE results suggest that uniform fuel switching policies are inefficient. Instead, Diesel should be deployed selectively on high-stress routes and seasons where it provides clearly beneficial emission reductions, while HFO can remain preferable on low-stress, maneuvering-dominated legs where Diesel offers little or no environmental advantage.
Key takeaway: The decision-relevant signal is heterogeneity: despite a near-zero global ATE, route–season CATEs indicate where Diesel is beneficial versus counterproductive under comparable operating conditions.

3.4. Compliance Risk and Economic Implications: Exceedance Risk

To connect predictive uncertainty with regulatory compliance, we constructed an Emission-Intensity Exceedance Risk Map based on the split-conformal intervals. For each route–season cell, the map reports the probability that the upper bound of the 90% conformal prediction interval exceeds the historical HFO median emission intensity on that corridor.
As shown in Figure 7, the resulting probabilities were high across the operational domain, ranging from 81.5% on Lagos–Apapa in winter to 97.3% on Warri–Bonny in Summer. This apparent saturation is a direct consequence of the conservative, distribution-free calibration: to guarantee coverage on completely unseen routes, the conformal procedure constructs intervals wide enough to encompass the worst-case residuals observed in the most challenging operating regimes.
From a risk-management perspective, high exceedance probabilities should therefore be interpreted as a feature rather than a flaw. They indicate that given current operational practices, even the upper-bound scenarios for many route–season combinations are likely to surpass historical HFO medians and thus warrant close monitoring in a CII-oriented compliance framework. Crucially, the Exceedance Risk metric integrates both the central tendency and the uncertainty of predicted CI, providing a more robust signal than point forecasts alone.
In practical terms, the Exceedance Risk Map can be embedded into Ship Energy Efficiency Management Plan (SEEMP) Part III processes as a route–season screening tool. Cells with persistently high exceedance probabilities flag segments where proactive intervention is most needed. Fleet managers can use these signals to prioritize speed optimization, weather routing, trim and engine-tuning checks, or targeted fuel switching prior to dispatch, thereby integrating statistically rigorous risk information into routine voyage planning and compliance reporting. In the next subsection, we complement this risk-based view with an explicit cost analysis through the Total Cost Intensity metric under alternative carbon pricing regimes.
Key takeaway: Carbon pricing and uncertainty together shape actionability—TCI identifies when selective switching becomes economically plausible, while exceedance risk prioritizes where compliance margins are most likely to be strained.

3.5. Total Cost Intensity

To integrate environmental outcomes with economic considerations, we analyzed Total Cost Intensity (TCI) in USD per nautical mile under three carbon pricing scenarios: 50, 100, and 150 USD/tCO2. Figure 8 displays the TCI distributions for HFO and Diesel voyages. Under the 50 USD/tCO2 scenario, HFO retained a clear cost advantage: its lower bunker price dominated the modest emission-based surcharge, and the TCI distributions of the two fuels remained well separated. As the carbon price increased to 100 and then 150 USD/tCO2, however, the TCI gap narrowed substantially. At 150 USD/tCO2, the Diesel and HFO distributions exhibited significant overlap, indicating that Diesel becomes economically viable for a non-trivial subset of voyages—particularly those high-stress routes identified in Section 3.3 where Diesel also delivers the largest physical emission reductions.
To illustrate the economic magnitude of these differences, consider a typical 150 nm coastal voyage. A TCI gap of 4–6 USD/nm between fuels translates into an additional 600–900 USD per trip, which is material for short-sea operators facing thin margins and frequent sailings. The threshold of approximately 100 USD/tCO2 at which Diesel becomes cost-competitive on the most challenging routes is broadly consistent with projected carbon price ranges under emerging EU ETS and FuelEU Maritime regimes. This alignment suggests that the crossover points observed in our TCI analysis are not mere theoretical artifacts but fall within plausible near-term regulatory scenarios. In combination with the heterogeneous CATE estimates, the TCI results support a strategy of selective Diesel deployment on high-stress routes where both environmental and economic incentives are aligned.

4. Discussion

The results were intentionally generated under a deployment-like design: forecasting and uncertainty quantification were evaluated under a strict leave-one-route-out protocol, while fuel-switching effects were interpreted as conditional contrasts rather than randomized outcomes. The aim of this discussion is therefore not to “celebrate” a single predictive metric, but to clarify what the combined predictive–uncertainty–causal outputs mean for corridor-level decision-making under compliance pressure and carbon-cost exposure. We synthesize three messages. First, when deployment involves domain shift, uncertainty-aware bounds can be more decision-relevant than point accuracy on unseen routes. Second, heterogeneity (CATE) should be treated as an operational decision map rather than a statistical curiosity, and it can overturn blanket “average-effect” rules. Third, these corridor-level tools matter because the sector’s decarbonization trajectory is increasingly shaped by a policy stack that rewards auditable, risk-aware operational planning. We close by translating findings into actionable ship-side and port-side implications, and by clarifying scope conditions under which the claims should (and should not) be generalized.

4.1. Uncertainty Guarantees and Physics-Informed Robustness

The results demonstrate that accurate voyage-level emission-intensity forecasting is achievable even under a stringent leave-one-route-out protocol. The physics-informed monotonic LightGBM model attained a global mean absolute error of 40.7 kg CO2/nm on completely unseen corridors.
The observation of negative $R2$ values and an MAE of ~40.7 kg CO2/nm warrants critical context. In a Leave-One-Route-Out (LORO) setting, the model faces a complete distribution shift; the ‘unseen’ route often operates at a fundamentally different intensity baseline than the training routes due to static geometric factors (e.g., river depth vs. open sea). A naïve baseline model predicting the global training mean for all test voyages would yield comparable absolute errors but would fail to capture intra-voyage dynamics. The practical utility of the proposed framework, therefore, does not rely on achieving a high $R2$ on unseen routes—which is mathematically impossible without target data—but on two other capabilities: (i) Risk Management, where conformal prediction intervals successfully adapt to this uncertainty providing a statistically guaranteed upper bound; and (ii) Causal Sensitivity, where the ML model, unlike a static baseline, learns valid physical derivatives allowing it to correctly estimate the marginal effect of fuel switching.
Consequently, the 90% split-conformal prediction intervals achieved 100% empirical coverage on the held-out routes. While coverage exceeding the nominal level might initially appear overly conservative, it is an expected and desirable outcome when calibration is performed under covariate shift and the method is required to remain valid on unseen corridors. By aggregating residuals across all folds, including the most challenging legs such as Lagos–Apapa and Escravos–Lagos, the calibration procedure inflates the conformal half-width to a global value of 81.2 kg CO2/nm that reflects worst-case conditions. In cells where hydrometeorology, congestion, or river-mouth constraints induce strong variability, wider intervals act as a statistical safety buffer, reducing the risk of under-reporting emissions or misclassifying voyages near compliance thresholds.
We specifically examined the validity of the monotonic distance constraint on short, congested segments like the Lagos–Apapa corridor. Although maneuvering adds noise, our diagnostics on the raw data confirm that the correlation between distance and emission intensity remains strongly positive; thus, the monotonic constraint serves as an effective regularizer that prevents the model from overfitting to transient congestion anomalies.
Two design choices reinforce that these outputs are deployment-oriented rather than retrospective curve-fitting: (i) the strict leave-one-route-out protocol, and (ii) leakage-aware out-of-fold constructs, which force the model to operate under genuine domain shift. The SHAP analysis further reinforces this robustness. Relative features—such as route-median deviations and route-level difficulty scores—consistently dominated the global SHAP-importance rankings, indicating that meaningful, route-relative operational physics are captured rather than mere memorization of raw route identifiers. At the same time, the monotonicity constraint imposed on distance_nm is clearly reflected in the SHAP sign pattern: longer voyages systematically exhibited negative SHAP contributions, confirming that the model has internalized the hydrodynamic principle whereby an increased share of steady cruising reduces emission intensity on a per-nautical-mile basis.
Key implication: Even when unseen-route point forecasts exhibit level shifts, uncertainty-aware upper bounds remain practically actionable for compliance-oriented screening and planning.

4.2. What “Heterogeneity” Means Operationally (ATE vs. CATE)

Average treatment effects (ATE, arithmetic mean) provide a compact summary, but operational decisions are made on specific corridor–season profiles where marginal impacts can vary. In this setting, heterogeneous effects (CATE) are best read as a decision map: they indicate where switching from HFO to Diesel is expected to reduce emission intensity and where it may be neutral or even adverse, conditional on the observed operational context. This is why heterogeneity should be treated as a feature, not a nuisance—especially in coastal operations where route geometry, maneuvering intensity, and weather regimes jointly shape both fuel selection and emissions response. A single global “fuel switching rule” can therefore be misleading, even when the average effect appears benign.
A physically interpretable way to read the heterogeneity is to link fuel effects to operating regimes. On longer, higher-stress passages with a larger share of steady open-coast steaming, main engines are more likely to operate nearer their efficient loading band, so changes in fuel properties can translate into measurable shifts in intensity per nautical mile. In contrast, on short, maneuvering-dominated legs—where acceleration–deceleration cycles, low-speed segments, and idling near terminal approaches are more prominent—engines spend more time in part-load regimes, and the expected direction of a fuel switch can attenuate or even reverse. In other words, the “same” fuel switch can behave differently depending on whether the voyage is dominated by steady cruising or by stop–start maneuvering and constrained approaches.
Accordingly, the interpretative value of the causal layer is prioritization, not universal prescriptions. The operationally appropriate use is to target switching (and complementary levers) to the corridor–season cells where conditional contrasts are favorable and counterfactual support is credible, while recognizing that other cells may be better served by speed management, maintenance planning, or port-side measures.
Key implication: The CATE layer is most useful as a targeted screening tool that supports selective interventions rather than uniform fleet-wide fuel-switching mandates.

4.3. International Context and Sector Evidence

Coastal shipping decarbonization is increasingly shaped by a portfolio of policy and market pressures that reward measurable efficiency and penalize carbon-intensive operations, including intensity-based schemes, emerging carbon-pricing mechanisms, and procurement expectations from cargo owners. Across international assessments of the energy transition, a recurring theme is that hard-to-abate transport segments require layered strategies: near-term operational efficiency gains, medium-term fuel and engine transitions constrained by bunkering and port readiness, and robust monitoring and verification that remains credible under real-world variability. In that sense, sector narratives (often grounded in global energy-transition evidence) consistently emphasize decision robustness and implementation feasibility over any single “best” model metric.
Within that broader framing, the contribution of this study is not a new sector-wide statistic but a corridor-level decision lens that makes variability, domain shift, and heterogeneity explicit. The leave-one-route-out results highlight that deployment often entails distribution shift, where absolute baselines can differ materially across corridors; this motivates conservative uncertainty quantification rather than over-reliance on point forecasts. In parallel, the cost-intensity framing shows how carbon pricing can change the attractiveness of operational choices, enabling the same emissions insights to be interpreted as cost exposure under plausible policy trajectories (including scenarios where carbon prices approach levels that narrow the cost gap between fuels on the most challenging routes). Together, these elements align with the structure of international sector discussions: “measure” (credible monitoring under variability), “manage” (risk-aware decision thresholds), and “mitigate” (targeted levers where the marginal benefit is plausibly positive).
Key implication: Corridor-level uncertainty bounds, heterogeneity maps, and cost-intensity stress tests translate broad policy and sector pressures into operationally interpretable route–season prioritization.

4.4. Policy and Sustainability Implications (Ship-Side + Port-Side)

From a policy and operations perspective, the key step is to treat uncertainty and heterogeneity as first-class planning inputs rather than residual noise. The exceedance-risk outputs based on conformal upper bounds can be used to flag corridor-season cells where the probability of breaching an intensity benchmark is high, while the carbon-pricing analysis frames how the same physical variability propagates into cost exposure. Fuel-switching guidance should then be applied selectively—only within the operational profiles where the conditional contrast is favorable and where comparable counterfactual voyages exist—while other cells may be better served by operational levers or port-side mitigation.

4.4.1. Economic Sensitivity and Market Drivers

The economic viability of fuel switching is fundamentally driven by the spread between HFO and Marine Diesel prices relative to the imposed carbon price. Our finding that Diesel becomes cost-competitive only at ~100 USD/tCO2 is therefore conditional on the baseline market conditions assumed in this study, rather than a universal threshold. Beyond carbon pricing itself, vessel- and regime-specific operating conditions (engine load, maneuvering share, and auxiliary/hoteling demand) shape the Total Cost Intensity (TCI) by affecting effective specific fuel consumption and the extent to which carbon-cost components accumulate under different route–season operating profiles. Table 4 summarizes the key economic variables used in this study, their baseline values, qualitative “range/scenario” descriptors (without additional computations), and the direction by which each factor is expected to shift the breakeven (crossover) carbon price.
As shown, the ~100 USD/tCO2 crossover should be interpreted as conditional on the baseline bunker price spread (≈300 USD/ton) observed during the study period, not as a universal constant. If global oil market dynamics were to compress this spread—for instance, due to higher refinery supply of distillates or regulatory shifts that make HFO structurally more expensive—the carbon price required to incentivize switching would decrease. Conversely, in a market with relatively cheap HFO and a widespread, even aggressive carbon pricing may struggle to bridge the gap without complementary operational or institutional incentives (e.g., differentiated port dues or corridor-level programs).
Seasonality and route regime matter because they alter the operating envelope in which costs are realized (e.g., maneuvering share, congestion exposure, and hoteling time), thereby shifting the effective breakeven in a corridor–season–specific way. Thus, the crossover should be interpreted at the corridor–season level, not as a single global constant.

4.4.2. Port-Side Enablers and Near-Coastal Mitigation

Complementary port-side measures can amplify the benefits of ship-side interventions, particularly for short sea routes with frequent turnaround. Energy management systems can reduce idle and auxiliary loads, while shore power and near-terminal electrification can decouple berth emissions from onboard generation. In parallel, integrating renewables into port microgrids can lower the effective carbon intensity of electrified operations, thereby improving the total system-level benefit of fuel switching and speed management. These port-side levers are particularly relevant in corridors where operational constraints limit the feasibility of aggressive ship-side retrofits.
A practical “how to use” synthesis for operators and regulators is as follows:
  • Use exceedance-risk outputs (based on the conformal upper bound) as a route–season screening layer for voyage planning and internal compliance checks, focusing attention on corridor–season cells where compliance margins are most fragile.
  • Apply selective fuel-switching guidance through heterogeneous-effect summaries (CATE) rather than a blanket fleet-wide rule, and explicitly avoid over-interpreting cells where support is limited.
  • Pair fuel decisions with routinely available ship-side levers (e.g., speed management, maintenance scheduling, and weather-aware routing), which can reduce variability and narrow decision uncertainty.
  • Coordinate with port and terminal stakeholders so that ship-side actions are not undermined by infrastructure constraints (e.g., turnaround pressures, auxiliary-load demand at berth, or limited availability of enabling port-side measures).
Methodologically, the framework is designed as a transport-work–agnostic decision architecture: in the absence of deadweight tonnage or detailed cargo data, the outcome is defined as voyage-level CI in kg CO2/nm, while the same workflow can be retargeted to AER/CII-style indicators if transport-work data become available. This framing supports a pragmatic pathway in which operators can start with auditable voyage-level screening and progressively refine both monitoring and decision rules as richer operational and cargo proxies are integrated.
Finally, the Nigerian short-sea case illustrates a broader relevance for emerging maritime regions: data-driven, route-specific tools can help align local operational realities with global decarbonization expectations without assuming that one-size-fits-all prescriptions transfer across infrastructure regimes. In that sense, corridor-level, uncertainty-aware decision support is compatible with sustainability goals that emphasize climate action, resilient infrastructure, and the protection of coastal and marine environments while remaining grounded in operational feasibility.

4.4.3. Operational Usage Scenario (Worked Example)

To illustrate the practical utility of these uncertainty intervals for SEEMP Part III planning, consider a fleet manager scheduling a vessel for the ‘Warri–Bonny’ route during the Summer season—a high-risk cell in our Risk Map. Suppose the model predicts a central emission intensity of 90 kg CO2/nm. With the calibrated conformal half-width of ~81 kg, the 90% prediction interval extends up to an upper bound of 171 kg CO2/nm.
While this interval is wide, it provides a critical “safety signal”. If the vessel’s internal CII reference line (converted to voyage-equivalent intensity) is, for instance, 120 kg CO2/nm, the fact that the upper bound (171 kg) significantly exceeds this limit serves as an immediate “Red Flag”. It warns the operator that under plausible worst-case conditions (e.g., adverse weather or congestion typical of unseen routes), this voyage could severely damage the ship’s annual CII rating. Consequently, the manager would be advised to intervene before departure—by enforcing a speed reduction, scheduling hull cleaning, or selecting a more efficient vessel for this specific leg—rather than risking a non-compliant voyage. In this context, the wide interval is not a lack of precision, but a necessary buffer against regulatory penalties.

4.4.4. Probability-Based Decision Rule (Link to Exceedance Risk)

In practice, operators and regulators may prefer a probability trigger rather than relying only on a single upper bound. In our framework, the exceedance risk map provides exactly this. For each route–season cell, it summarizes the probability that the uncertainty-aware bound exceeds a chosen baseline (e.g., a historical median or an internal compliance line). A simple SEEMP Part III rule can then be defined: if the exceedance probability is persistently high (e.g., above an operator-defined trigger), the voyage is classified as “high-risk” and requires pre-departure mitigation (speed discipline, maintenance scheduling, or selective fuel strategy); otherwise, standard operating procedures apply. This makes the wide interval interpretable as a quantified compliance-risk signal rather than a vague loss of precision.
Key implication: The most credible near-term gains come from pairing uncertainty-aware targeting (what/where/when)—operationalized via conformal upper bounds and exceedance probabilities—with feasible ship-side actions and enabling port-side measures (how).

4.5. Limitations and Scope

This work is grounded in observational voyage data and therefore inherits the usual limitations of deployment-oriented empirical studies. Some determinants of emission intensity—such as load condition, hull/propeller fouling, draft/trim state, dry-docking cycles, and localized congestion or current effects—may be imperfectly observed, which can contribute to residual uncertainty in both forecasting and conditional fuel-switching contrasts. We therefore interpreted the fuel-switching results as conditional contrasts under the stated assumptions, and we treated uncertainty quantification as a conservative risk-management layer rather than a guarantee of point accuracy.
Because the analysis relied on a single-operator coastal network and fuel assignment may be partially structured at the route and vessel levels, the ATE/CATE estimates should be interpreted as conditional contrasts within the observed covariate support (positivity region) of this dataset. Extrapolating these estimates to other fleets, regions, or bunkering/operational regimes requires re-estimation and renewed overlap/balance assessment on the target deployment data. Appendix C makes the supported domain explicit and reports robustness to trimming and omitted-variable sensitivity.
Future work should prioritize: (i) multi-operator and multi-region validations to assess external validity under different fuel policies and infrastructure regimes; (ii) richer operational covariates and improved sensing/proxy design for load, hull condition, and congestion/currents to reduce residual confounding and tighten uncertainty bounds; and (iii) extension to additional operational treatments (e.g., speed reduction or alternative fuels) and systematic retargeting to evolving intensity indicators as transport-work data become available.
Since the Annual Efficiency Ratio (AER) and CII differ from our outcome metric primarily by a static denominator (DWT), our distance-based predictions Y ^ can be directly converted to CII estimates via the transformation C I I ^ = Y ^ / D W T , assuming the vessel’s capacity remains constant.
Key implication: The study’s claims are strongest within the observed operational support of the analyzed network and should be extended through re-estimation and richer data rather than extrapolated.

5. Conclusions

Under a strict Leave-One-Route-Out (LORO) evaluation, the physics-informed monotonic LightGBM model attained a global MAE of 40.7 kg CO2/nm on completely unseen corridors. While route-level R2 values can be negative due to baseline level shifts (see Table 3), this highlights that point forecasts on unseen routes serve primarily as screening tools rather than stable baselines. To manage this uncertainty, distribution-free split-conformal prediction yielded 90% prediction intervals with 100% empirical coverage on held-out routes. These bounds, mapped via Exceedance Risk, provide concrete “safety triggers” for SEEMP Part III-oriented planning (see Figure 7), ensuring that compliance risks are identified before departure.
Regarding operational measures, the estimated average effect of switching from HFO to Diesel was statistically indistinguishable from zero (−0.072 kg CO2/nm), but this global average masks significant heterogeneity. Causal Forests revealed that treatment effects vary widely by route and season (CATE range: −74 g to +29 g per nm), supporting a context-specific strategy where switching is treated as a conditional lever within the observed support rather than a universal prescription.
Linking emissions with economic realities, the Total Cost Intensity (TCI) analysis indicates that widespread Diesel adoption becomes economically attractive only when carbon prices approach ~100 USD/tCO2 (see Figure 8). As summarized in Table 4, this breakeven point is dynamic and shifts based on market drivers and operating regimes.
These conclusions are drawn from an observational panel (N = 1440) covering four Nigerian coastal routes and should be interpreted within the observed support of this network. Because transport-work data (e.g., DWT) are unavailable, we modeled voyage-level CI in kg CO2/nm; however, the proposed architecture is explicitly “CII-ready” and can be retargeted to AER/CII-style indicators when such data become available. For practical use, the workflow provides a unified toolkit: compliance planning via conformal bounds, operational choice via heterogeneous causal patterns, and economic stress-testing via TCI under evolving carbon-price scenarios.

Author Contributions

Conceptualization, M.Y.; methodology, M.Y. and G.E.; software, A.A. and M.Y.; validation, A.A. and G.E.; formal analysis, M.Y. and A.A.; investigation, M.Y.; resources, M.Y.; data curation, M.Y.; writing—original draft preparation, M.Y.; writing—review and editing, G.E.; visualization, M.Y.; supervision, G.E. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The operational data presented in this study are available in the publicly accessible “Ship Fuel Efficiency” repository on Kaggle at “https://www.kaggle.com/code/oumaimagasmi/ship-fuel-efficiency/ (accessed on 15 September 2025)”. The derived voyage-level panel constructed for this study is described in Appendix B.

Acknowledgments

The authors thank the anonymous reviewers for their constructive comments.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AERAnnual Efficiency Ratio
AISAutomatic Identification System
ATEAverage Treatment Effect (arithmetic mean)
CATEConditional Average Treatment Effect (arithmetic mean)
CIICarbon Intensity Indicator
CO2Carbon Dioxide
DMLDouble/Debiased Machine Learning
DWTDeadweight Tonnage
EEDIEnergy Efficiency Design Index
EU ETSEuropean Union Emissions Trading System
GHGGreenhouse Gas
HFOHeavy Fuel Oil
IMOInternational Maritime Organization
LightGBMLight Gradient Boosting Machine
LOROLeave-One-Route-Out
MAEMean Absolute Error
MDOMarine Diesel Oil
MLMachine Learning
RMSERoot Mean Square Error
SDGSustainable Development Goals
SEEMPShip Energy Efficiency Management Plan
SHAPSHapley Additive exPlanations
TCITotal Cost Intensity

Appendix A

Appendix A.1. Data Origin and Purpose

The first dataset, ship_fuel_efficiency [43] is a publicly available real-world fuel-efficiency dataset describing maritime operations in Nigeria. It reports fuel consumption and CO2 emissions for several ship types operating on typical coastal routes, together with basic operational and environmental descriptors. The dataset was originally designed for exploratory analysis of fuel-usage patterns and emissions across common ship categories (e.g., Fishing Trawler, Oil Service Boat, Surfer Boat, Tanker Ship) and was used in this study as the conceptual starting point for constructing the enriched voyage-level operational panel described in Appendix B.
In its raw form, ship_fuel_efficiency.csv contains 1440 voyage records with ship-level and voyage-level information consistent with short-sea coastal shuttle operations in Nigeria. While the main causal–conformal modeling in Section 3, Section 4 and Section 5 did not directly use this file as the modeling dataset, its structure and variables informed the design of the derived voyage-level panel on which all machine-learning and causal analyses were based.

Appendix A.2. Original Variables and Units

Table A1 lists all variables contained in ship_fuel_efficiency, including their units, type, and data level.
Table A1. Original variables in the “ship_fuel_efficiency” dataset.
Table A1. Original variables in the “ship_fuel_efficiency” dataset.
Variable NameDescriptionUnitTypeData Level
ship_idUnique identifier assigned to each real-world vessel operating in the panel(string label)CategoricalShip/voyage identifier
ship_typeCategorical indicator of ship class (e.g., Oil Service Boat, Tanker Ship, Surfer Boat, Fishing Trawler)(category)CategoricalShip-type level, repeated across voyages
route_idIdentifier of the coastal corridor served (e.g., Port Harcourt–Lagos, Lagos–Apapa, Escravos–Lagos, Warri–Bonny)(string label)CategoricalVoyage-level (route for each record)
monthCalendar month of operation (January–December), used as a proxy for seasonality(category)CategoricalVoyage-level time index
distanceLength of the voyage performed by the vessel on the given route and monthNautical miles (nm)ContinuousVoyage-level
fuel_typeType of fuel used on the voyage (Heavy Fuel Oil, HFO; or Marine Diesel)(category)CategoricalVoyage-level treatment indicator
fuel_consumptionTotal fuel consumed during the voyageLitersContinuousVoyage-level aggregate
CO2_emissionsCarbon dioxide emissions associated with the reported fuel consumptionKilograms (kg)ContinuousVoyage-level aggregate
weather_conditionsDiscrete description of the prevailing weather during the voyage (Calm, Moderate, Stormy)(category)CategoricalVoyage-level environmental descriptor
engine_efficiencyEffective engine efficiency during the voyage, expressed as a percentagePercent (%)ContinuousVoyage-level technical/operational flag

Appendix A.3. Pre-Processing and Limitations

The ship_fuel_efficiency dataset [43] is real-world and primarily organized at the voyage level, with each record representing the operation of a given ship on a specific route and month under a given fuel type and weather condition. Although it contains realistic ranges for distance, fuel consumption and CO2 emissions, it does not by itself include the full set of route-relative metrics, seasonal encodings, cost variables or engineered features required by the causal–conformal framework developed in the main text.
For this reason, the empirical analysis in Section 2, Section 3, Section 4 and Section 5 was conducted on the derived voyage-level operational panel ship_fuel_efficiency_derived, documented in Appendix B. The original public dataset remains important as the conceptual and numerical foundation of the study and is documented here for transparency and reproducibility.

Appendix B

The second dataset, ship_fuel_efficiency_derived, was the core voyage-level operational panel used in all machine-learning, conformal, and causal analyses reported in Section 3, Section 4 and Section 5.

Appendix B.1. Construction from the Source Dataset

The panel in ship_fuel_efficiency_derived was constructed from, and conceptually aligned with, the public fuel-efficiency dataset described in Appendix A, enriched with additional operational, environmental and cost information tailored to Nigerian short-sea coastal shuttle trades. Starting from ship-, route- and month-level records, the following steps were applied:
  • explicit voyage identifiers are introduced (ship_id, route_id, month, distance_nm) together with fuel choice (fuel_type);
  • fuel consumption in tons and associated CO2 emissions in both kilograms and tons are computed;
  • a voyage-level emission-intensity measure (CI_kg_per_nm) consistent with the main text is derived;
  • weather categories and corresponding numeric encodings are added;
  • seasonality features are constructed (calendar month numbers, sine–cosine encodings, and season_bin categories);
  • route-relative difficulty metrics are created (route_difficulty_score, route_median_ratio, intensity_gap);
  • energy and carbon costs, as well as Total Cost Intensity (TCI) under alternative carbon-price scenarios, are computed.
The resulting dataset is organized at the voyage level and is fully consistent with the description in Section 2 (“Data and Descriptive Statistics”).

Appendix B.2. Panel Overview and Scope

The derived panel comprised N = 1440 voyages performed by real-world general cargo and coastal tanker vessels operating across four Nigerian coastal routes (Port Harcourt–Lagos, Lagos–Apapa, Escravos–Lagos, Warri–Bonny) over the period 2022–2024. Each row corresponds to a single voyage characterized by a specific ship, route, month, fuel type and prevailing weather category.
Voyage-level emission intensity (CI_kg_per_nm) was the main outcome variable used in the predictive and causal models. Fuel type (fuel_type, with a corresponding binary dummy) served as the treatment in the causal analysis. All machine-learning, conformal and causal-inference results reported in Section 3, Section 4 and Section 5 were based exclusively on this derived voyage-level dataset.

Appendix B.3. Variable Dictionary for the Derived Panel

Table A2 lists all variables contained in the ship_fuel_efficiency_derived dataset, together with their definitions, units and analytical roles.
Table A2. Variables in the “ship_fuel_efficiency_derived” dataset (voyage-level panel).
Table A2. Variables in the “ship_fuel_efficiency_derived” dataset (voyage-level panel).
Variable NameDescriptionUnitRole
ship_idUnique identifier for each vessel in the panel(string label)Identifier
ship_typeShip class (Oil Service Boat, Fishing Trawler, Surfer Boat, Tanker Ship)(category)Covariate (ship attribute)
route_idRoute identifier (Warri–Bonny, Port Harcourt–Lagos, Lagos–Apapa, Escravos–Lagos)(category)Identifier/Covariate
monthCalendar month of the voyage(category)Time covariate
distance_nmLength of the voyageNautical miles (nm)Core operational covariate
log_distanceNatural logarithm of distance_nmlog-nmEngineered covariate
fuel_typeFuel used on the voyage (HFO vs. Diesel)(category)Treatment
fuel_consumption_literTotal fuel consumed during the voyageLitersOutcome/cost driver
fuel_consumption_tonFuel consumption converted to mass using standard density factorsTonsOutcome/cost driver
CO2_kgCO2 emissions associated with the voyageKilograms (kg)Outcome (emissions level)
CO2_tonCO2 emissions expressed in tonsTonsOutcome (emissions level)
weather_categoryDiscrete weather state (Calm, Moderate, Stormy)(category)Environmental covariate
weather_codeNumeric code corresponding to weather_category (e.g., 0 = Calm, 1 = Moderate, 2 = Stormy)(integer code)Engineered covariate
engine_efficiency_rawRaw engine efficiency indicator for the voyagePercent (%)Technical covariate
fuel_intensity_ton_per_nmFuel consumption intensity per nautical mile (fuel_consumption_ton/distance_nm)Tons/nmDerived outcome/covariate
CI_kg_per_nmVoyage-level emission intensity: CO2 per nautical mile (CO2_kg/distance_nm)kg CO2/nmPrimary outcome
route_CI_meanMean emission intensity for the corresponding route, used as a route-level baselinekg CO2/nmEngineered covariate
route_difficulty_scoreDifference between a voyage’s CI_kg_per_nm and the route-level baseline (route_CI_mean)kg CO2/nmEngineered covariate (route difficulty)
route_median_ratioRatio of voyage emission intensity to a global or route-specific median CIDimensionless ratioEngineered covariate
intensity_gapDeviation of voyage CI from the season-specific median intensitykg CO2/nmEngineered covariate
CI_AER_readyRescaled emission intensity prepared for potential conversion to AER/CII metricskg CO2/nm (proxy)Engineered outcome proxy
month_numberNumeric encoding of calendar month (1–12)(integer)Time covariate
season_binSeasonal category (Winter, Spring, Summer, Autumn)(category)Time/environmental covariate
energy_costFuel cost for the voyage at baseline bunker pricesUSDOutcome (cost)/covariate
carbon_cost_lowCarbon cost for the voyage at a low CO2 price scenario (e.g., 50 USD/tCO2)USDEngineered cost covariate
carbon_cost_100Carbon cost for the voyage at 100 USD/tCO2USDEngineered cost covariate
carbon_cost_highCarbon cost for the voyage at a high CO2 price scenario (e.g., 150 USD/tCO2)USDEngineered cost covariate
total_cost_50Sum of energy_cost and carbon_cost_low (total cost at low carbon price)USDOutcome (cost scenario)
total_cost_100Sum of energy_cost and carbon_cost_100 (total cost at 100 USD/tCO2)USDOutcome (cost scenario)
total_cost_150Sum of energy_cost and carbon_cost_high (total cost at high carbon price)USDOutcome (cost scenario)
TCI_50Total Cost Intensity per nautical mile at low carbon price (total_cost_50/distance_nm)USD/nmOutcome (TCI)
TCI_100Total Cost Intensity per nautical mile at 100 USD/tCO2 (total_cost_100/distance_nm)USD/nmOutcome (TCI)
TCI_150Total Cost Intensity per nautical mile at high carbon price (total_cost_150/distance_nm)USD/nmOutcome (TCI)
weather_stress_indexInteger index reflecting the operational severity of weather (e.g., 0 = Calm, 1 = Moderate, 2 = Stormy)(integer)Engineered environmental covariate
engine_efficiency_normNormalized engine efficiency, scaled for use in machine-learning modelsDimensionlessEngineered technical covariate
sin_monthSine transformation of month_number for cyclic seasonality encodingDimensionlessEngineered time covariate
cos_monthCosine transformation of month_number for cyclic seasonality encodingDimensionlessEngineered time covariate
fuel_diesel_dummyBinary indicator for Diesel usage (1 = Diesel, 0 = HFO)(0/1)Treatment encoding
fuel_weather_interactionEncoded interaction between fuel type and weather stress (e.g., Diesel × Stormy)(integer code)Engineered interaction covariate
engine_distance_interactionInteraction term between engine efficiency and distance, capturing scale effects on performanceMixed (efficiency × nm)Engineered interaction covariate

Appendix C. Causal Diagnostics and Robustness Checks

This appendix documents where causal fuel-switching contrasts are empirically supported in the data and how robust they are to overlap, balance, trimming, and omitted-variable concerns. Operationally, a causal contrast requires comparable voyages—i.e., voyages with similar distance_nm, weather_category, season_bin, ship_type, and route-stress proxies—where either fuel_type could plausibly be chosen. When one fuel is effectively deterministic for a corridor/season (e.g., due to bunkering constraints, terminal requirements, or operator policy), the counterfactual comparison becomes extrapolation rather than comparison. We therefore report support counts and propensity overlap by route, assess covariate balance before/after trimming to common-support regions, and summarize how ATE/CATE estimates change under trimming and sensitivity perturbations. These diagnostics are not presented as proof of causality; instead, they clarify the empirical domain over which the identification assumptions are most defensible and the estimated contrasts are most interpretable.
Table A3 confirms that all route–season combinations contained non-zero counts for both fuel types (i.e., support is available). This ensures that CATE estimates are based on observed contrasts rather than pure extrapolation.
Table A3. Route × season support counts by fuel type. Columns: route_id (or route name), season_bin, n_Diesel, n_HFO. Note: CATE summaries are restricted to cells with non-zero counts for both fuels.
Table A3. Route × season support counts by fuel type. Columns: route_id (or route name), season_bin, n_Diesel, n_HFO. Note: CATE summaries are restricted to cells with non-zero counts for both fuels.
Route (route_id)Season (season_bin)n_Dieseln_HFOSupported?
Escravos–LagosAutumn6532Yes
Spring5835Yes
Summer4730Yes
Winter6240Yes
Lagos–ApapaAutumn5140Yes
Spring6530Yes
Summer5753Yes
Winter5933Yes
Port Harcourt–LagosAutumn7034Yes
Spring6933Yes
Summer6236Yes
Winter5332Yes
Warri–BonnyAutumn4622Yes
Spring3832Yes
Summer4332Yes
Winter5427Yes
Figure A1 Propensity-score distributions by route (four panels).
Figure A1. Route-wise distributions of the estimated propensity score e(X) = P(T = 1|X) for Diesel (treated) and HFO (control). Vertical lines indicate trimming thresholds e(X) ∈ [0.10, 0.90] and e(X) ∈ [0.20, 0.80] used for robustness checks.
Figure A1. Route-wise distributions of the estimated propensity score e(X) = P(T = 1|X) for Diesel (treated) and HFO (control). Vertical lines indicate trimming thresholds e(X) ∈ [0.10, 0.90] and e(X) ∈ [0.20, 0.80] used for robustness checks.
Sustainability 18 00723 g0a1
Table A4 shows that prior to trimming, significant imbalances existed (e.g., ship_type SMD > 1.0), reflecting the endogenous nature of fuel choice. Trimming the sample to the common support region (0.1 < e(X) < 0.9) drastically reduced these imbalances, bringing almost all key covariates, including route metrics and vessel types, below or near the 0.10 diagnostic threshold.
Table A4. Covariate balance diagnostics (Standardized Mean Differences, SMD). Panels: (i) before trimming, (ii) after trimming to e(X) ∈ [0.10, 0.90], and (iii) after trimming to e(X) ∈ [0.20, 0.80].
Table A4. Covariate balance diagnostics (Standardized Mean Differences, SMD). Panels: (i) before trimming, (ii) after trimming to e(X) ∈ [0.10, 0.90], and (iii) after trimming to e(X) ∈ [0.20, 0.80].
CovariateSMD (Before Trimming)SMD (After Trim 0.1–0.9)SMD (After Trim 0.2–0.8)
distance_nm0.1940.0250.025
route_difficulty_score0.6260.0710.071
route_median_ratio0.6250.0720.072
intensity_gap0.6220.0710.071
ship_type_Surfer Boat1.0610.0000.000
ship_type_Tanker Ship0.3380.0540.054
ship_type_Oil Service0.2260.0730.073
weather_Moderate0.1190.1210.121
season_Summer0.1070.1310.131
Figure A2 visually demonstrates the effectiveness of the identification strategy. In the Full Sample, variables such as ship_type_Surfer Boat and route_difficulty_score showed distinct imbalances, with standardized mean differences well above the 0.10 threshold, confirming that fuel choice is endogenous. However, applying the trimming criteria brought the SMD for all key operational covariates below or near the 0.10 diagnostic threshold. This visual convergence confirms that within the trimmed subsample, Diesel and HFO voyages possess comparable covariate distributions.
Figure A2. Love plot summarizing covariate balance diagnostics before and after propensity score trimming. The figure displays the absolute Standardized Mean Difference (SMD) between Diesel (treated) and HFO (control) voyages for each covariate. Blue circles represent the full, unmatched sample, revealing significant imbalances (SMD > 0.1). Orange crosses and green squares represent the samples after trimming to the regions of common support (0.1 < e(X) < 0 and 0.2 < e(X) < 0.8, respectively). The vertical red dashed line at SMD = 0.1 indicates the conventional threshold for negligible imbalance; the shift of all covariates to the left of this threshold confirms satisfactory balance.
Figure A2. Love plot summarizing covariate balance diagnostics before and after propensity score trimming. The figure displays the absolute Standardized Mean Difference (SMD) between Diesel (treated) and HFO (control) voyages for each covariate. Blue circles represent the full, unmatched sample, revealing significant imbalances (SMD > 0.1). Orange crosses and green squares represent the samples after trimming to the regions of common support (0.1 < e(X) < 0 and 0.2 < e(X) < 0.8, respectively). The vertical red dashed line at SMD = 0.1 indicates the conventional threshold for negligible imbalance; the shift of all covariates to the left of this threshold confirms satisfactory balance.
Sustainability 18 00723 g0a2
Table A5 shows that the Average Treatment Effect (ATE) remained statistically indistinguishable from zero across all specifications. The stability of the estimate (~−0.05 g/nm) after trimming confirms that the result is not driven by lack of overlap in extreme propensity regions.
Table A5. ATE robustness: full sample vs. trimmed samples.
Table A5. ATE robustness: full sample vs. trimmed samples.
SampleATE Estimate (kg CO2/nm)95% CI Lower95% CI UpperSignificant?
Full Sample (N = 1440)−0.00004−0.000150.00008No
Trimmed (0.1–0.9)−0.00005−0.000210.00010No
Trimmed (0.2–0.8)−0.00005−0.000240.00013No
Table A6 shows that the Trimmed CATE model focused on the region of common support (N = 1116), removing voyages where fuel choice was effectively deterministic. The wider range in the Trimmed model reflects better identification of the heterogeneous effects (e.g., high savings in specific conditions) that were previously smoothed out by the full model’s noise.
Table A6. CATE robustness summary across supported route × season cells.
Table A6. CATE robustness summary across supported route × season cells.
Summary MetricFull Model (N = 1440)Trimmed Model (N = 1116)
Mean CATE−0.00002−0.00006
Min CATE−0.00022−0.00088
Max CATE0.000150.00027
CorrelationLow (Due to noise in Full model)
Table A7 shows that the propensity model is primarily driven by vessel type (Surfer Boat), which aligns with operational reality (smaller, faster vessels rely on MDO). Route difficulty and seasonal factors also play a role, confirming that fuel choice is endogenous to operational stress.
Table A7. Propensity model drivers.
Table A7. Propensity model drivers.
FeatureCoefficientDirectionOperational Interpretation
ship_type_Surfer Boat2.51(+)Surfer boats are highly likely to use Diesel.
sin_month−0.35(−)Seasonal cyclicity influences fuel allocation.
season_Spring0.32(+)Diesel usage increases slightly in Spring.
ship_type_Tanker−0.25(−)Tankers have a slight preference for HFO.
route_median_ratio0.18(+)Higher-intensity (difficult) voyages favor Diesel.
Table A8 shows that the sensitivity analysis indicates that the near-zero ATE is robust to the omission of key covariate blocks. The estimated effect does not flip sign or become significant even when major factors like weather or distance are removed, suggesting that the main result is stable.
Table A8. Sensitivity summary for unobserved confounding.
Table A8. Sensitivity summary for unobserved confounding.
Sensitivity QuantityValueInterpretation
Robustness Value
(RVq = 1RVq = 1)
1.65%Unobserved confounders would need to explain 1.65% of the residual variance to explain away the estimated effect.
ATE without Distance0.0000No change. Distance is well-balanced.
ATE without Weather0.0000No change. Weather impact is handled.
ATE without Route Metrics0.0024Slight shift, confirming route metrics capture important confounding.

References

  1. International Maritime Organization. Resolution Mepc.352(78) (Adopted on 10 June 2022) 2022 Guidelines on Operational Carbon Intensity Indicators and The Calculation Methods (CII Guidelines, G1); International Maritime Organization: London, UK, 2022. [Google Scholar]
  2. Kotzampasakis, M. Maritime Emissions Trading in the EU: Systematic Literature Review and Policy Assessment. Transp. Policy 2025, 165, 28–41. [Google Scholar] [CrossRef]
  3. United Nations. Unctad Review of Maritime Transport: Navigating Maritime Chokepoints; United Nations Publication: New York, NY, USA, 2024. [Google Scholar]
  4. Anantharaman, M.; Sardar, A.; Islam, R. Decarbonization of Shipping and Progressing Towards Reducing Greenhouse Gas Emissions to Net Zero: A Bibliometric Analysis. Sustainability 2025, 17, 2936. [Google Scholar] [CrossRef]
  5. Tadros, M.; Ventura, M.; Soares, C.G. Review of Current Regulations, Available Technologies, and Future Trends in the Green Shipping Industry. Ocean Eng. 2023, 280, 114670. [Google Scholar] [CrossRef]
  6. Psaraftis, H.N.; Kontovas, C.A. Decarbonization of Maritime Transport: Is There Light at the End of the Tunnel? Sustainability 2020, 13, 237. [Google Scholar] [CrossRef]
  7. Dominioni, G. Carbon Pricing for International Shipping, Equity, and WTO Law. Rev. Eur. Comp. Int. Environ. Law 2024, 33, 19–30. [Google Scholar] [CrossRef]
  8. Christodoulou, A.; Cullinane, K. The Prospects for, and Implications of, Emissions Trading in Shipping. Marit. Econ. Logist. 2024, 26, 168–184. [Google Scholar] [CrossRef]
  9. Jamasb, T. Regulatory Economics and the Future of Ports as Energy Hubs. In Ports as Energy Transition Hubs; Sornn-Friese, H., Ed.; CBS Maritime: Copenhagen, Denmark, 2024; pp. 34–42. [Google Scholar]
  10. Alshareef, M.H.; Alghanmi, A.F. Optimizing Maritime Energy Efficiency: A Machine Learning Approach Using Deep Reinforcement Learning for EEXI and CII Compliance. Sustainability 2024, 16, 10534. [Google Scholar] [CrossRef]
  11. Gkerekos, C.; Lazakis, I. A Novel, Data-Driven Heuristic Framework for Vessel Weather Routing. Ocean Eng. 2020, 197, 106887. [Google Scholar] [CrossRef]
  12. Li, Z.; Fei, J.; Du, Y.; Ong, K.L.; Arisian, S. A near Real-Time Carbon Accounting Framework for the Decarbonization of Maritime Transport. Transp. Res. E Logist. Transp. Rev. 2024, 191, 103724. [Google Scholar] [CrossRef]
  13. Issa-Zadeh, S.B.; Garay-Rondero, C.L. Decarbonizing Seaport Maritime Traffic: Finding Hope. World 2025, 6, 47. [Google Scholar] [CrossRef]
  14. Kerton, F.M. UN Sustainable Development Goals 14 and 15—Life below Water, Life on Land. RSC Sustain. 2023, 1, 401–403. [Google Scholar] [CrossRef]
  15. Vakili, S.; Insel, M.; Singh, S.; Ölçer, A. Decarbonizing Domestic and Short-Sea Shipping: A Systematic Review and Transdisciplinary Pathway for Emerging Maritime Regions. Sustainability 2025, 17, 7294. [Google Scholar] [CrossRef]
  16. Serra, P.; Fancello, G. Towards the IMO’s GHG Goals: A Critical Overview of the Perspectives and Challenges of the Main Options for Decarbonizing International Shipping. Sustainability 2020, 12, 3220. [Google Scholar] [CrossRef]
  17. Oloruntobi, O.; Chuah, L.F.; Mokhtar, K.; Gohari, A.; Rady, A.; Abo-Eleneen, R.E.; Akhtar, M.S.; Mubashir, M. Decarbonising ASEAN Coastal Shipping: Addressing Climate Change and Coastal Ecosystem Issues through Sustainable Carbon Neutrality Strategies. Environ. Res. 2024, 240, 117353. [Google Scholar] [CrossRef]
  18. Govindan, K.; Dua, R.; Mehbub Anwar, A.; Bansal, P. Enabling Net-Zero Shipping: An Expert Review-Based Agenda for Emerging Techno-Economic and Policy Research. Transp. Res. E Logist. Transp. Rev. 2024, 192, 103753. [Google Scholar] [CrossRef]
  19. Zadeh, S.B.I.; Soltani, H.R. Revamping Seaport Operations with Renewable Energy: A Sustainable Approach to Reducing Carbon Footprint. GMSARN Int. J. 2024, 18, 315–324. [Google Scholar]
  20. Müller-Casseres, E.; Leblanc, F.; van den Berg, M.; Fragkos, P.; Dessens, O.; Naghash, H.; Draeger, R.; Le Gallic, T.; Tagomori, I.S.; Tsiropoulos, I.; et al. International Shipping in a World below 2 °C. Nat. Clim. Change 2024, 14, 600–607. [Google Scholar] [CrossRef]
  21. Zis, T.; Psaraftis, H.N. Operational Measures and Logistical Considerations for the Decarbonisation of Maritime Transport. In Proceedings of the 7th Symposium of the European Association for Research in Transportation, Athens, Greece, 5–7 September 2018. [Google Scholar]
  22. Lindstad, H.; Asbjørnslett, B.E.; Strømman, A.H. Reductions in Greenhouse Gas Emissions and Cost by Shipping at Lower Speeds. Energy Policy 2011, 39, 3456–3464. [Google Scholar] [CrossRef]
  23. Krantz, G.; Moretti, C.; Brandão, M.; Hedenqvist, M.; Nilsson, F. Assessing the Environmental Impact of Eight Alternative Fuels in International Shipping: A Comparison of Marginal vs. Average Emissions. Environments 2023, 10, 155. [Google Scholar] [CrossRef]
  24. Pettit, S.; Wells, P.; Haider, J.; Abouarghoub, W. Revisiting History: Can Shipping Achieve a Second Socio-Technical Transition for Carbon Emissions Reduction? Transp. Res. D Transp. Environ. 2018, 58, 292–307. [Google Scholar] [CrossRef]
  25. Boviatsis, M.; Alexopoulos, A.B.; Vlachos, G.P. Evaluation of the Response to Emerging Environmental Threats, Focusing on Carbon Dioxide (CO2), Volatile Organic Compounds (VOCs), and Scrubber Wash Water (SOx). Euro-Mediterr. J. Environ. Integr. 2022, 7, 391–398. [Google Scholar] [CrossRef]
  26. Graham, D.J. Causal Inference for Transport Research. Transp. Res. Part A Policy Pract. 2025, 192, 104324. [Google Scholar] [CrossRef]
  27. Sung, I.; Zografakis, H.; Nielsen, P. Multi-Lateral Ocean Voyage Optimization for Cargo Vessels as a Decarbonization Method. Transp. Res. D Transp. Environ. 2022, 110, 103407. [Google Scholar] [CrossRef]
  28. Caldeirinha, V.; Felício, J.A.; Pinho, T. Role of Cargo Owner in Logistic Chain Sustainability. Sustainability 2023, 15, 10018. [Google Scholar] [CrossRef]
  29. Cepowski, T.; Drozd, A. Measurement-Based Relationships between Container Ship Operating Parameters and Fuel Consumption. Appl. Energy 2023, 347, 121315. [Google Scholar] [CrossRef]
  30. Haque, T.; Bin Syed, M.A.; Das, S.; Ahmed, I. Advancing Marine Surveillance: A Hybrid Approach of Physics Infused Neural Network for Enhanced Vessel Tracking Using Automatic Identification System Data. J. Mar. Sci. Eng. 2024, 12, 1913. [Google Scholar] [CrossRef]
  31. Tang, L.; Luo, R.; Zhou, Z.; Colombo, N. Enhanced Route Planning with Calibrated Uncertainty Set. Mach. Learn. 2025, 114, 129. [Google Scholar] [CrossRef]
  32. Angelopoulos, A.N.; Bates, S. Conformal Prediction: A Gentle Introduction. Found. Trends Mach. Learn. 2023, 16, 494–591. [Google Scholar] [CrossRef]
  33. Jasmi, M.F.A.; Fernando, Y. Drivers of Maritime Green Supply Chain Management. Sustain. Cities Soc. 2018, 43, 366–383. [Google Scholar] [CrossRef]
  34. Yuan, Q.; Wang, S.; Peng, J. Operational Efficiency Optimization Method for Ship Fleet to Comply with the Carbon Intensity Indicator (CII) Regulation. Ocean. Eng. 2023, 286, 115487. [Google Scholar] [CrossRef]
  35. Agand, P.; Kennedy, A.; Harris, T.; Bae, C.; Chen, M.; Park, E.J. Fuel Consumption Prediction for a Passenger Ferry Using Machine Learning and In-Service Data: A Comparative Study. Ocean. Eng. 2023, 284, 115271. [Google Scholar] [CrossRef]
  36. Bolton, P.; Adrian, T.; Kleinnijenhuis, A. The Great Carbon Arbitrage; IMF Working Papers; International Monetary Fund: Washington, DC, USA, 2022; Volume 2022, pp. 1–67. [Google Scholar] [CrossRef]
  37. Pedersen, L.H. Carbon Pricing versus Green Finance. SSRN Electron. J. 2023. [Google Scholar] [CrossRef]
  38. Millischer, L.; Evdokimova, T.; Fernandez, O. The Carrot and the Stock: In Search of Stock-Market Incentives for Decarbonization. Energy Econ. 2023, 120, 106615. [Google Scholar] [CrossRef]
  39. Burgass, M.J.; Halpern, B.S.; Nicholson, E.; Milner-Gulland, E.J. Navigating Uncertainty in Environmental Composite Indicators. Ecol. Indic. 2017, 75, 268–278. [Google Scholar] [CrossRef]
  40. Papadopoulos, H.; Proedrou, K.; Vovk, V.; Gammerman, A. Inductive Confidence Machines for Regression. In European Conference on Machine Learning; Springer: Berlin/Heidelberg, Germany, 2002; pp. 345–356. [Google Scholar]
  41. Chernozhukov, V.; Chetverikov, D.; Demirer, M.; Duflo, E.; Hansen, C.; Newey, W.; Robins, J. Double/Debiased Machine Learning for Treatment and Structural Parameters. Econom. J. 2018, 21, C1–C68. [Google Scholar] [CrossRef]
  42. Wager, S.; Athey, S. Estimation and Inference of Heterogeneous Treatment Effects Using Random Forests. J. Am. Stat. Assoc. 2018, 113, 1228–1242. [Google Scholar] [CrossRef]
  43. Gasmi, O. Ship Fuel Efficiency Kaggle. Available online: https://www.kaggle.com/code/oumaimagasmi/ship-fuel-efficiency/ (accessed on 15 September 2025).
  44. Ke, G.; Meng, Q.; Finley, T.; Wang, T.; Chen, W.; Ma, W.; Ye, Q.; Liu, T.-Y. LightGBM: A Highly Efficient Gradient Boosting Decision Tree. In Proceedings of the 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA, 4–9 December 2017. [Google Scholar]
Figure 1. Study corridor network in coastal Nigeria (Gulf of Guinea). The six ports used in the voyage-level panel are shown as nodes (Lagos, Apapa, Escravos, Warri, Port Harcourt, Bonny), and the four analyzed routes are indicated with colored arrows: Port Harcourt–Lagos, Lagos–Apapa, Escravos–Lagos, and Warri–Bonny. The figure highlights the corridor-level geometry underlying the leave-one-route-out evaluation and the route–season stratification used in the predictive, uncertainty, and causal analyses.
Figure 1. Study corridor network in coastal Nigeria (Gulf of Guinea). The six ports used in the voyage-level panel are shown as nodes (Lagos, Apapa, Escravos, Warri, Port Harcourt, Bonny), and the four analyzed routes are indicated with colored arrows: Port Harcourt–Lagos, Lagos–Apapa, Escravos–Lagos, and Warri–Bonny. The figure highlights the corridor-level geometry underlying the leave-one-route-out evaluation and the route–season stratification used in the predictive, uncertainty, and causal analyses.
Sustainability 18 00723 g001
Figure 2. Descriptive characteristics of the dataset (N = 1440). (a) Frequency of voyages by route. (b) Seasonal distribution (25.0% per season). (c) Fuel type proportions (Marine Diesel 62.4%, HFO 37.6%). (d) Weather conditions (Calm 35.8%, Moderate 32.1%, Stormy 32.1%). (e) Distribution of voyage distance (nm). (f) Distribution of emission intensity (kg CO2/nm), showing a near-normal distribution centered around 79 kg CO2/nm. Together, these panels depict a fleet engaged in balanced, year-round coastal shuttle operations, with no single route, season or weather state dominating the sample. This multimodality is typical of short-sea coastal trade and ensures that the subsequent analysis reflects a broad range of operational regimes rather than a single, idiosyncratic corridor.
Figure 2. Descriptive characteristics of the dataset (N = 1440). (a) Frequency of voyages by route. (b) Seasonal distribution (25.0% per season). (c) Fuel type proportions (Marine Diesel 62.4%, HFO 37.6%). (d) Weather conditions (Calm 35.8%, Moderate 32.1%, Stormy 32.1%). (e) Distribution of voyage distance (nm). (f) Distribution of emission intensity (kg CO2/nm), showing a near-normal distribution centered around 79 kg CO2/nm. Together, these panels depict a fleet engaged in balanced, year-round coastal shuttle operations, with no single route, season or weather state dominating the sample. This multimodality is typical of short-sea coastal trade and ensures that the subsequent analysis reflects a broad range of operational regimes rather than a single, idiosyncratic corridor.
Sustainability 18 00723 g002
Figure 3. Emission intensity vs. voyage distance by fuel type. The pronounced negative curve highlights the transport-work efficiency of longer voyages, where steady cruising dominates maneuvering and port approaches. The significant overlap between Diesel (red) and HFO (blue) clusters illustrates the confounding role of operational factors. This pattern is consistent with operational experience: as passage distance increases, the relative share of fuel-intensive maneuvering and port approaches decreases compared to steady cruising, so longer voyages naturally exhibit lower emission intensity per nautical mile despite occasionally harsher environmental conditions.
Figure 3. Emission intensity vs. voyage distance by fuel type. The pronounced negative curve highlights the transport-work efficiency of longer voyages, where steady cruising dominates maneuvering and port approaches. The significant overlap between Diesel (red) and HFO (blue) clusters illustrates the confounding role of operational factors. This pattern is consistent with operational experience: as passage distance increases, the relative share of fuel-intensive maneuvering and port approaches decreases compared to steady cruising, so longer voyages naturally exhibit lower emission intensity per nautical mile despite occasionally harsher environmental conditions.
Sustainability 18 00723 g003
Figure 4. Estimated propensity-score distributions for Diesel (treated) and HFO (control) voyages. The figure illustrates the degree of empirical overlap in e ( X ) within the observed covariate space; complementary overlap, balance, trimming, and sensitivity diagnostics are reported in Appendix C.
Figure 4. Estimated propensity-score distributions for Diesel (treated) and HFO (control) voyages. The figure illustrates the degree of empirical overlap in e ( X ) within the observed covariate space; complementary overlap, balance, trimming, and sensitivity diagnostics are reported in Appendix C.
Sustainability 18 00723 g004
Figure 5. Model interpretability via SHAP values. (a) Global feature importance based on mean absolute SHAP values, highlighting the dominance of route-relative metrics (e.g., route_median_ratio, route_difficulty_score) in predicting emission intensity on unseen routes. (b) SHAP summary plot illustrating the direction of impact for key features. Consistent with the imposed monotonicity constraint on distance, longer voyages (red points for distance_nm) are associated with lower emission intensity (negative SHAP values), while adverse operational conditions and high route difficulty increase predicted CI.
Figure 5. Model interpretability via SHAP values. (a) Global feature importance based on mean absolute SHAP values, highlighting the dominance of route-relative metrics (e.g., route_median_ratio, route_difficulty_score) in predicting emission intensity on unseen routes. (b) SHAP summary plot illustrating the direction of impact for key features. Consistent with the imposed monotonicity constraint on distance, longer voyages (red points for distance_nm) are associated with lower emission intensity (negative SHAP values), while adverse operational conditions and high route difficulty increase predicted CI.
Sustainability 18 00723 g005
Figure 6. Conditional Average Treatment Effect (CATE) of switching from HFO to Diesel (g CO2/nm), estimated via Causal Forest. Values represent the gram-level change in CO2 per nautical mile when Diesel is used instead of HFO, conditional on route and season. Blue regions indicate emission savings, peaking at −74 g/nm on the Port Harcourt–Lagos route in Autumn. Red/orange regions indicate emission increases, reaching +29 g/nm on the Lagos–Apapa route in Summer. The heterogeneity confirms that Diesel is not universally superior but yields substantial abatement only under specific high-stress operating conditions.
Figure 6. Conditional Average Treatment Effect (CATE) of switching from HFO to Diesel (g CO2/nm), estimated via Causal Forest. Values represent the gram-level change in CO2 per nautical mile when Diesel is used instead of HFO, conditional on route and season. Blue regions indicate emission savings, peaking at −74 g/nm on the Port Harcourt–Lagos route in Autumn. Red/orange regions indicate emission increases, reaching +29 g/nm on the Lagos–Apapa route in Summer. The heterogeneity confirms that Diesel is not universally superior but yields substantial abatement only under specific high-stress operating conditions.
Sustainability 18 00723 g006
Figure 7. Emission-Intensity Exceedance Risk Map. Values denote the probability that the upper 90% conformal prediction bound exceeds the route-specific historical HFO median emission intensity for each route–season cell. Probabilities range from 81.5% (Lagos–Apapa, winter) to 97.3% (Warri–Bonny, Summer), reflecting the conservative calibration required to maintain coverage guarantees on strictly unseen routes. High exceedance probabilities identify segments where operational or fuel switching interventions should be prioritized within SEEMP Part III.
Figure 7. Emission-Intensity Exceedance Risk Map. Values denote the probability that the upper 90% conformal prediction bound exceeds the route-specific historical HFO median emission intensity for each route–season cell. Probabilities range from 81.5% (Lagos–Apapa, winter) to 97.3% (Warri–Bonny, Summer), reflecting the conservative calibration required to maintain coverage guarantees on strictly unseen routes. High exceedance probabilities identify segments where operational or fuel switching interventions should be prioritized within SEEMP Part III.
Sustainability 18 00723 g007
Figure 8. Total Cost Intensity (USD/nm) by fuel type under three carbon pricing scenarios p 50,100,150   USD t CO 2 . The boxplots show that while HFO retains a median cost advantage due to lower bunker prices, the TCI distributions converge as the carbon price rises. At 150 USD/tCO2, the significant overlap indicates that Diesel becomes a cost-competitive option for operationally efficient voyages, despite having a higher baseline fuel cost.
Figure 8. Total Cost Intensity (USD/nm) by fuel type under three carbon pricing scenarios p 50,100,150   USD t CO 2 . The boxplots show that while HFO retains a median cost advantage due to lower bunker prices, the TCI distributions converge as the carbon price rises. At 150 USD/tCO2, the significant overlap indicates that Diesel becomes a cost-competitive option for operationally efficient voyages, despite having a higher baseline fuel cost.
Sustainability 18 00723 g008
Table 1. Operational weather categorization criteria.
Table 1. Operational weather categorization criteria.
CategoryBeaufort RangeApprox. Hs (m)Operational Meaning
Calm0–3<1.25Normal coastal steaming, negligible impact on speed.
Moderate4–51.25–2.5Increased resistance, occasional slamming, minor speed reductions.
Stormy≥6>2.5Heavy-weather maneuvering, enforced speed reductions.
Table 2. Summary statistics of key voyage-level variables (N = 1440). Energy cost reflects bunker expenditure only.
Table 2. Summary statistics of key voyage-level variables (N = 1440). Energy cost reflects bunker expenditure only.
VariableMeanMedianStd. Dev.MinMax
Distance (nm)151.8123.5108.520.1498.6
Emission Intensity (kg CO2/nm)78.979.228.025.6146.6
Total CO2 Emissions (tons)13.48.413.60.671.9
Energy Cost (USD)24291501252214214,666
Table 3. Route-level predictive performance on unseen test sets (Leave-One-Route-Out Cross-Validation).
Table 3. Route-level predictive performance on unseen test sets (Leave-One-Route-Out Cross-Validation).
FoldUnseen Route (Test Set)MAE (kg/nm)RMSE (kg/nm)R2
1Port Harcourt–Lagos33.7341.31–1.14
2Lagos–Apapa45.0052.58–2.51
3Escravos–Lagos46.9554.72–2.70
4Warri–Bonny37.0344.07–1.72
Avg.Global Average40.6848.17–2.02
Table 4. Summary of Economic Variables and Sensitivity Drivers (descriptive transparency table; no additional sensitivity computations are introduced).
Table 4. Summary of Economic Variables and Sensitivity Drivers (descriptive transparency table; no additional sensitivity computations are introduced).
VariableBaseline ValueRange/Scenario (Qualitative)Operational ImpactSensitivity/Effect on Crossover Point
HFO Price400 USD/tonMarket-dependent; cheaper HFO higher crossoverHighA decrease in HFO price widens the cost gap, pushing the required carbon tax above ~100 USD/tCO2.
Diesel Price700 USD/tonMarket-dependent; narrower spread lower crossoverHighThe bunker price spread (≈300 USD/ton in our data) is the primary barrier; if the spread narrows, the crossover carbon price drops proportionally.
Carbon Price0–150 USD/tCO2Policy lever (0–150 in scenarios)HighAt carbon prices around and above ~100 USD/tCO2, the carbon-cost component begins to outweigh the bunker price spread.
Specific Fuel Consumption (SFOC)Variable (engine load dependent)Load-dependent; port/maneuvering increases breakevenModerateDiesel engines lose efficiency at low loads (port/maneuvering), increasing the breakeven carbon price on short routes such as Lagos–Apapa.
Auxiliary Load~10–15% of total~10–15%; higher hoteling favors HFOLow–ModerateAssumed constant proportional to hoteling time; higher auxiliary loads generally favor HFO due to lower base cost.
CO2 emission factor (HFO vs. Diesel)Fuel-specific (standard factors used in cost construction)Fuel-specific; secondaryLow–ModerateSecondary driver; the bunker price spread dominates the crossover point under the assumed market conditions.
Operational regime (route × season: weather, maneuvering share, hoteling time)Captured via season_bin, weather_category, and route/distance proxiesSeasonal/route-specific; shifts breakevenModerateShort/maneuvering-heavy legs raise the effective breakeven; long steady-steaming legs lower it.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Yildiz, M.; Akgundogdu, A.; Elmas, G. Decarbonizing Coastal Shipping: Voyage-Level CO2 Intensity, Fuel Switching and Carbon Pricing in a Distribution-Free Causal Framework. Sustainability 2026, 18, 723. https://doi.org/10.3390/su18020723

AMA Style

Yildiz M, Akgundogdu A, Elmas G. Decarbonizing Coastal Shipping: Voyage-Level CO2 Intensity, Fuel Switching and Carbon Pricing in a Distribution-Free Causal Framework. Sustainability. 2026; 18(2):723. https://doi.org/10.3390/su18020723

Chicago/Turabian Style

Yildiz, Murat, Abdurrahim Akgundogdu, and Guldem Elmas. 2026. "Decarbonizing Coastal Shipping: Voyage-Level CO2 Intensity, Fuel Switching and Carbon Pricing in a Distribution-Free Causal Framework" Sustainability 18, no. 2: 723. https://doi.org/10.3390/su18020723

APA Style

Yildiz, M., Akgundogdu, A., & Elmas, G. (2026). Decarbonizing Coastal Shipping: Voyage-Level CO2 Intensity, Fuel Switching and Carbon Pricing in a Distribution-Free Causal Framework. Sustainability, 18(2), 723. https://doi.org/10.3390/su18020723

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop