1. Introduction
Shipyards operate under complex and demanding production conditions shaped by engineering-to-order manufacturing, labour-intensive workflows, tight spatial constraints, and exposure to global demand volatility. These characteristics make shipbuilding one of the most complex and least automated manufacturing sectors, with productivity levels that have historically lagged behind other heavy industries [
1]. At the same time, shipyards face intensifying economic pressures, including rising labour costs, global competition, and thin operating margins, alongside emerging decarbonisation requirements driven by the International Maritime Organisation (IMO) Greenhouse Gas (GHG) Strategy, the European Union (EU) Emissions Trading System expansion, the EU Green Deal, the EU Circular Economy Action Plan, and national industrial-emissions regulations [
2,
3,
4,
5,
6]. Together, these forces are compelling shipyards to modernise production processes, reduce energy and material waste, and adopt more data-driven decision-making frameworks.
Digitalisation technologies have therefore become essential for improving efficiency, reducing cost, and enabling the transition toward smarter, lower-carbon shipyard operations. While many digital tools were already common in other industries, shipbuilding began adopting them systematically only in the last decade. Before 2010, most applications were confined to isolated workshop simulations or custom diagnostic tools. This changed in the early 2010s with advances in computing, the maturation of commercial Discrete-Event Simulation (DES) software, Industry 4.0 developments, and the growth of the Internet of Things (IoT) and Machine Learning (ML). These technologies enabled integrated product–process–resource (PPR) structures, optimisation linkages, and early forms of real-time data interoperability [
1,
7].
As a result, the period from roughly 2010 to 2025 marks the first sustained phase in which digitalisation in shipbuilding evolved into coherent cyber-physical systems capable of supporting integrated planning, multi-resource scheduling, logistics coordination, energy-aware optimisation, and Digital Twin (DT)-enabled operational control. This represents a clear departure from pre-2010 work, where DES was used mainly for isolated bottleneck or line-level analyses and lacked the integrated PPR structures, optimisation linkages, and interoperable data environments that characterise contemporary research.
The purpose of this paper is to present a structured review of simulation methods, optimisation techniques, Artificial Intelligence (AI)- and ML-based approaches, and DT technologies as they relate specifically to shipyard production. Rather than surveying the broader maritime or supply-chain domains, this study looks directly inside the shipyard, tracing how digital tools support modelling, scheduling, logistics, and emerging forms of real-time operational control across workshops, stockyards, block-handling systems, and the cranes and transporters that connect them. Although previous reviews have discussed DES, Industry 4.0 concepts, or selected automation technologies, and recent studies examined shipyards from sustainability and system-transformation perspectives, with emphasis on regulatory frameworks, environmental technologies, and organisational change [
8], none integrate these methodological streams into a unified account of how shipyard-production modelling has evolved over the past fifteen years, particularly with respect to digital production-planning methods. This study fills that gap by tracing the progression from pre-2010 isolated workshop simulations, through hybrid simulation–optimisation methods (2010–2016), to the emergence of AI-enhanced and DT-enabled cyber-physical systems (2017–2025). By situating these developments within a single framework, the review clarifies how modelling foundations, computational techniques, and data infrastructures have co-evolved and identifies the structural gaps that remain.
To achieve this aim, the review examines research published between 2010 and 2025 and synthesises developments across the major analytical and digitalisation trends in shipyard production.
Section 2 assesses the impacts of DES and discrete-time simulation (DTS), Multi-Objective Optimisation (MOO), and hybrid simulation–optimisation methods on shipyard systems, tracing how these tools have shaped bottleneck analysis, scheduling, layout planning, and environmental performance.
Section 3 evaluates the expansion of AI, ML, reinforcement learning (RL), and DT technologies, highlighting their role in cyber-physical integration, predictive modelling, and adaptive decision support.
Section 4 reviews the software ecosystem that enables these methods, commercial DES environments, open-source Python/SimPy frameworks, optimisation engines, and Manufacturing Execution System (MES), Enterprise Resource Planning (ERP), and Product Lifecycle Management (PLM) platforms that supply the structured product, process, resource, and execution-layer information required for high-fidelity simulation and DT synchronisation.
Section 5 analyses the availability and limitations of industrial datasets, identifying the confidentiality barriers that constrain reproducibility and benchmarking.
Section 6 synthesises the major research gaps and outlines future directions, while
Section 7 concludes by summarising the key findings and their implications for next-generation, data-rich, DT-enabled shipyard systems.
The literature was identified through targeted searches in Scopus, Web of Science, and Google Scholar using Boolean combinations of keywords such as shipyard production, discrete-event simulation, optimisation, machine learning, digital twin, and Industry 4.0. Additional papers were identified from the reference lists of key studies. Publications from 2010 to 2025 were considered. Only English-language, peer-reviewed papers and high-quality conference proceedings directly addressing shipyard manufacturing or production digitalisation were included. To complement academic sources, industry reports, white papers, and official company websites demonstrating practical DT or MES/PLM implementations were also reviewed. In total, approximately 190 sources were screened, and 92 met the inclusion criteria.
This paper is positioned as a structured, domain-specific methodological review rather than a PRISMA-style systematic review. This choice reflects the heterogeneous reporting practices and context-dependent validation characteristics of shipyard production research, which spans conceptual architectures, simulation-based implementations, and partially validated cyber-physical systems. Accordingly, the review emphasises critical comparative evaluation rather than exhaustive enumeration. Beyond literature identification, an explicit qualitative evaluative framework is applied to support comparative synthesis. Each study is assessed along three analytical dimensions: (i) model maturity (conceptual framework, prototype implementation, validation with industrial data, or live deployment); (ii) validation basis (simulation-only testing, historical data replay, partial shop-floor validation, or live operational use); and (iii) decision scope (diagnostic analysis, planning support, scheduling optimisation, or real-time control). Numerical scoring is deliberately avoided because validation contexts across studies are not comparable. Contributions lacking explicit validation depth, scalability discussion, or clear decision relevance are retained for completeness but interpreted primarily as exploratory or architectural within the comparative synthesis. This structured perspective guides the organisation of the review and underpins the identification of methodological limitations, limits of applicability, and future research directions presented later in the paper.
2. Foundations of Shipyard Production Modelling: DES, Optimisation, and Hybrid Methods
Shipyard production research has evolved substantially over the past two decades, driven by the need to manage highly complex, spatial–temporal, multi-resource manufacturing systems. Across this literature, three methodological streams have emerged as the dominant analytical foundations: DES/DTS, optimisation and metaheuristics, and hybrid simulation–optimisation approaches (
Figure 1). These streams differ not only in modelling formalisms, but also in analytical role, validation depth, decision scope, and degree of operational maturity.
To distinguish methodological robustness and operational credibility, the reviewed studies in this section are interpreted using a four-level hierarchy of evidence: (i) simulation-only evaluation using stylised or synthetic inputs; (ii) calibration or replay using historical industrial data; (iii) partial shop-floor or Advanced Planning and Scheduling (APS)/MES-integrated validation; and (iv) sustained live operational deployment with demonstrable production key performance indicators (KPIs). This hierarchy enables differentiation between diagnostic capability and validated industrial performance, rather than treating all reported improvements as equivalent.
In the first methodological stream, DES captures system evolution through event-driven state changes, such as welding completion and transporter arrivals, thereby representing stochastic interactions, spatial interference, and resource contention that static or purely analytical models cannot resolve. DTS advances the system at fixed temporal intervals and is therefore better suited to planning contexts governed by periodic decision cycles, such as shift-based scheduling and synchronised APS updates. The second methodological family comprises MOO and related optimisation approaches, which formalise trade-offs among competing objectives such as makespan, material-handling cost, energy consumption, and CO2 emissions. A third and increasingly prominent category is hybrid simulation–optimisation, which integrates DES with Constraint Programming (CP), evolutionary algorithms, or RL to reconcile mathematical optimality with operational feasibility under spatial and temporal constraints.
While these methods collectively form the analytical backbone of contemporary shipyard production research, their comparative credibility depends on validation depth, representational fidelity, and demonstrated applicability beyond isolated workshops. Across the literature, these dimensions are addressed unevenly.
Early DES research in shipyard production modelling focused primarily on bottleneck identification, throughput maximisation, and lead-time reduction. These studies consistently demonstrate the diagnostic strength of DES in revealing inefficiencies that remain invisible to aggregate or deterministic planning methods. For example, Caprace et al. [
9] reported improved lead time and labour efficiency in liquified natural gas (LNG) block erection through revised block-splitting strategies evaluated via DES. Ozkok and Helvacioglu [
10] showed substantial throughput gains in double-bottom fabrication when DES was combined with Optimised Production Technology. Rouco-Couzo et al. [
11] identified welding capacity as a critical bottleneck, with capacity expansion reducing makespan, and Tamer et al. [
12] demonstrated measurable reductions in material-handling time through layout-based DES scenarios.
From an evaluative standpoint, however, most of these studies fall within the first two levels of the evidence hierarchy: simulation-based experimentation or validation against limited historical datasets. Although Hadjina et al. [
13] reported acceptable deviation between simulated and observed behaviour in a robotised profile line, the validation scope remained workshop-specific and temporally bounded. Consequently, while DES demonstrates high behavioural fidelity and strong feasibility-analysis capability, the majority of reported improvements reflect controlled scenario testing rather than sustained yard-scale operational deployment. DES can therefore be considered methodologically mature in representational terms, but only partially mature in execution-layer integration.
Following early bottleneck-focused applications, DES-based research expanded toward coordination across planning levels and interconnected subsystems. Cebral Fernandez et al. [
14] showed that misalignment between planning tiers could increase makespan by up to 30%, and Sender et al. [
15] demonstrated the necessity of explicitly modelling production–logistics coupling to avoid systematic delay propagation. These contributions extend DES beyond local diagnostics toward planning-level integration; however, they remain predominantly validated through offline experimentation rather than real-time execution-layer synchronisation.
Taken together, the DES/DTS literature demonstrates high behavioural fidelity and strong diagnostic capability at workshop and subsystem levels. However, consistent evidence of yard-scale, closed-loop optimisation under continuous real-time disturbance is largely absent. These gains demonstrate analytical capability but do not constitute evidence of sustained yard-scale operational transformation.
Beyond classical DES applications, several studies extend simulation formalisms to address uncertainty and reallocative decision-making. However, their methodological maturity and validation basis vary considerably. Hadžić et al. [
16] employed a Markov and probabilistic simulation framework for a fabrication line and reported a 13% production-rate increase through improved resource reallocation. While analytically valuable, the study is validated at the workshop level and relies on simulation-based experimentation rather than live deployment, placing it within a simulation-only evidence tier. Similarly, Shablykova [
17] applied simulation to early-stage layout and material-flow planning, demonstrating that buffer placement and routing logic decisively influence throughput before detailed scheduling is fixed. This contribution is conceptually important, shifting the decision scope toward pre-scheduling design, but remains exploratory in validation depth. Okubo and Mitsuyuki [
18] introduced a discrete-time simulation model linking high-level production plans to operational constraints; their results show that lower-resolution temporal models can expose sequencing conflicts overlooked by deterministic planning. Yet the abstraction required for tractability limits spatial realism and constrains applicability to yard-scale coordination.
Collectively, these studies confirm that simulation supports decision-making across layout design, planning coordination, and workshop flow control. Yet most remain confined to offline environments using historical or synthetic datasets. As a result, modelling sophistication has advanced more rapidly than empirical validation under live execution-layer conditions.
A similar differentiation emerges in sustainability-oriented research. In DES-based studies such as Ahmad [
19], environmental metrics are incorporated at the evaluation layer, allowing estimation of energy consumption under alternative scenarios without modifying scheduling logic. This approach broadens performance assessment but does not alter decision structure, meaning sustainability remains diagnostically informative rather than optimisation-driving. In contrast, optimisation-focused studies (e.g., Guo et al. [
20]; Jiang et al. [
21]; Guo et al. [
22]; Li et al. [
23]; Meng et al. [
24]) embedded carbon emissions, energy use, and noise directly into multi-objective formulations. These models demonstrate mathematical capability in balancing environmental and productivity objectives and often report measurable percentage improvements relative to baseline heuristics. However, most are validated through simulation benchmarking or controlled case studies rather than partial or live industrial deployment. As a result, while optimisation maturity is high at the algorithmic level, deployment maturity remains limited. Furthermore, spatial interference and dynamic execution-layer variability are frequently abstracted to maintain computational tractability, constraining their applicability to tightly coupled yard-scale systems.
Labour utilisation and material-handling optimisation illustrate a similar pattern. DES-based lean reconfiguration studies such as Hassan and Kajiwara [
25] demonstrated meaningful reductions in waiting time and congestion through rule-based adjustments, but these improvements are scenario-dependent and lack systematic robustness testing under stochastic disturbances. Algorithmic optimisation approaches, including Tokola et al. [
26], Gunawan et al. [
27], Pambudi et al. [
28], and Kafali [
29], reported significant reductions in peak labour load or material-handling cost. Yet many assume simplified routing constraints or static layouts, and several compare performance only against heuristic baselines rather than real operational schedules. Consequently, while reported gains are substantial, the hierarchy of evidence indicates simulation-based validation rather than live execution-layer confirmation.
Hybrid simulation–optimisation frameworks provide comparatively stronger methodological credibility than standalone DES or purely analytical optimisation. In disruption-sensitive scheduling contexts, studies such as Kwak et al. [
30] combined CP with DES-based feasibility validation, demonstrating measurable makespan reductions while preserving due-date compliance under realistic constraints. Guan et al. [
31] applied Mixed-Integer Linear Programming (MILP) with column generation to syncrolift operations and reported substantial waiting-time reductions relative to baseline rules. The distinguishing feature of these approaches is that optimisation outputs are explicitly tested within simulated execution environments, allowing mathematical solutions to be stress-tested against spatial, temporal, and resource-interference constraints. Although validation typically remains at the case-study or simulation level rather than sustained live deployment, the integration of optimisation logic with behavioural simulation places hybrid methods higher in evidentiary robustness than optimisation models evaluated purely against analytical benchmarks or DES studies used solely for diagnosis.
More operationally grounded contributions appear in APS–DES integration research. Ju et al. [
32] and related works [
33,
34] embedded backward process-centric DES within planning systems, demonstrating substantial delay reduction and improved schedule robustness. These studies move closer to partial validation through alignment with real planning systems, indicating higher deployment maturity relative to simulation-only studies. However, even here, public disclosure of detailed datasets remains limited, constraining reproducibility and cross-study comparability.
A differentiated hierarchy of methodological credibility emerges from the reviewed literature. Standalone DES provides high behavioural realism but remains predominantly diagnostic and offline. Pure optimisation offers rigorous trade-off exploration but often abstracts spatial and disturbance dynamics. Hybrid simulation–optimisation frameworks demonstrate comparatively greater operational credibility because optimisation logic is tested against behavioural constraints. Nevertheless, most evidence remains confined to simulation or historical-data tiers, and sustained yard-scale deployment is rare. Thus, while analytical progress is substantial, extrapolation to full yard-wide operational transformation remains constrained.
Spatial–temporal modelling constitutes a decisive methodological boundary between analytically elegant schedules and operationally feasible production plans. Early DES research predominantly emphasised temporal dynamics, queue lengths, processing times, and resource utilisation, while treating spatial configuration as fixed or exogenous. However, a subset of studies demonstrates that spatial feasibility is not a secondary modelling refinement but a structural precondition for credible scheduling. Jeong et al. [
35] showed that transporter routing distances, blocking constraints, and stockyard geometry materially alter logistics behaviour; schedules derived without explicit spatial embedding systematically underestimate congestion and access conflicts. Similarly, Scarlat et al. [
36] illustrated that geospatial modelling reveals routing incompatibilities and movement bottlenecks that remain invisible in purely temporal formulations. Studnev and Burmistrov [
37] further argued that production-preparation models lacking accurate physical layout representation cannot reliably support downstream optimisation. These contributions are predominantly validated through structured simulation experiments at the workshop or subsystem scale. Their principal strength lies in enhancing spatial feasibility realism; their limitation lies in restricted decision scope and limited integration with execution-layer systems. Within the defined evidence hierarchy, spatially embedded DES models therefore possess greater operational credibility than optimisation formulations that abstract spatial constraints.
As hybrid approaches expanded, simulation began to adopt multi-resolution and integrative modelling philosophies. Cha and Roh [
38] proposed a combined DES–DTS framework capable of representing both event-driven shop-floor logic and periodic planning cycles. Although developed prior to contemporary DT discourse, their architecture anticipates key cyber-physical requirements, including asynchronous state updates and multi-temporal synchronisation. More recent work by Iwańkowicz and Rutkowski [
39] explicitly situated simulation within DT architectures, demonstrating how DES can function as an integrative behavioural core linking planning, monitoring, and execution data. However, it is important to distinguish architectural capability from validated operational impact. These DES–DT-integrated studies remain at the prototype or simulation-synchronised level, rather than demonstrating sustained yard-wide deployment under live operational conditions. Validation typically relies on controlled scenarios or partial shop-floor alignment rather than longitudinal performance assessment under real disturbance conditions. Accordingly, increasing architectural sophistication does not automatically correspond to higher evidentiary robustness. A differentiated hierarchy of methodological robustness emerges from the reviewed literature:
Standalone DES models provide high behavioural fidelity and strong diagnostic insight, particularly for bottleneck analysis and feasibility testing. Their primary weakness lies in offline execution, reliance on static or historically calibrated inputs, and limited adaptability to dynamic (real-time) disturbances.
Pure optimisation and metaheuristic models offer rigorous exploration of multi-objective trade-offs and can deliver substantial performance gains in structured case studies. However, many formulations simplify spatial–temporal interactions and assume stable system parameters, limiting credibility when extrapolated to disturbance-prone yard environments.
Hybrid simulation–optimisation frameworks demonstrate comparatively greater operational credibility because optimisation outputs are subjected to behavioural validation. By reconciling mathematical efficiency with spatial–temporal realism, these approaches reduce the risk of infeasible schedules. Nevertheless, even these contributions are predominantly validated through structured case data or bounded experimentation rather than sustained live industrial deployment.
As summarised in
Table 1, reported improvements in makespan, energy consumption, material-handling cost, and labour utilisation are often substantial. Yet these gains must be interpreted within their evidentiary context. The majority derive from simulation-only studies, historical data replay, or bounded case scenarios at the workshop scale. Yard-wide integration across multiple workshops, docks, subcontractors, and energy systems is rarely demonstrated. Therefore, while the literature convincingly establishes analytical potential, it does not provide conclusive evidence of sustained, full-scale operational transformation.
A key inference from the current body of work is that behavioural realism, particularly spatial embedding and disturbance-aware validation, constitutes the principal differentiator of methodological credibility. However, due to confidentiality constraints and limited access to openly reproducible industrial datasets, cross-study benchmarking remains weak, and comparative claims must be interpreted cautiously. The existing literature allows strong conclusions regarding feasibility enhancement and structured-case performance gains, but it does not yet support definitive claims about yard-scale, long-term productivity transformation.
In summary,
Section 2 documents significant methodological maturation from diagnostic DES toward integrated simulation–optimisation frameworks with increasing spatial realism and disruption sensitivity. Hybrid approaches currently constitute the most operationally credible methodological class within this stream. Nonetheless, the prevailing evidence base remains predominantly simulation-bound and subsystem-focused, underscoring the need for tighter execution-layer integration and real-time adaptability, issues that motivate the data-driven and cyber-physical developments examined in
Section 3.
3. Emerging Data-Driven and Cyber-Physical Approaches in Shipyard Production Optimisation: AI, ML, and DTs
Over the past decade, digitalisation, AI, and cyber-physical integration have become increasingly prominent in shipyard production optimisation, extending the simulation–optimisation trajectory established in
Section 2. These approaches rarely replace DES or MOO; instead, they augment them by enabling richer data capture, predictive modelling, and more responsive decision support. In this context, DTs are typically framed as virtual representations of physical production systems synchronised with operational data streams, although the degree of synchronisation, integration, and decision authority varies widely across the literature.
Because these technologies are often presented with broad claims (e.g., “real-time optimisation” or “intelligent control”), their contribution must be interpreted through explicit criteria rather than architectural description alone. Accordingly, the studies reviewed in
Section 3 are evaluated along three operational dimensions: (i) integration depth with execution systems (e.g., MES/Supervisory Control and Data Acquisition (SCADA)/IoT connectivity and actuation pathways), (ii) validation basis (simulation-only, historical replay/calibration, partial shop-floor testing, or sustained live use with production KPIs), and (iii) decision scope (prediction/estimation, planning, scheduling, dispatching/control). This lens is applied explicitly to DT and DT-enabling studies via comparative classification (
Table 2) and is used to critically interpret supporting digitalisation and AI contributions discussed below.
3.1. Machine Learning and Reinforcement Learning for Prediction, Simulation-Driven Scheduling, and Control
Across shipyard applications, ML contributes primarily to upstream modelling rather than direct system-level optimisation. The most operationally credible ML contributions in this literature are those that improve process representation where analytic parameterisation is weak, such as forming, heating, or nesting, thereby strengthening the inputs to downstream DES and optimisation. For example, Luo et al. [
42] used a Sparrow Search Algorithm-Backpropagation neural network (SSA-BP) model to predict forming deformation; Moon et al. [
43] applied deep learning with optimisation to inverse design of line-heating patterns; and Son et al. [
44] used ML-based nesting to improve material utilisation. Methodologically, these studies are valuable because they replace coarse empirical assumptions with learned surrogates; however, their decision scope remains component- or process-level, and their system-level impact is indirect. Their operational maturity therefore depends on (a) training-data representativeness and (b) demonstrated integration into planning/scheduling loops, which is often not fully evidenced in publicly available results.
In contrast, RL is positioned as a control and scheduling mechanism, typically trained through interaction with simulated environments that resemble DES. Woo et al. [
45] used deep RL for load balancing in a planning simulation; Le et al. [
46] trained RL policies for robotic water-blasting; and Nam et al. [
47] used deep RL to multi-objective parallel-machine scheduling within a DES-based environment. These studies demonstrate that RL can generate adaptive dispatching policies through interaction with simulated environments, reducing reliance on manually specified scheduling rules. However, across the reviewed literature, validation is predominantly conducted in simulation-based settings or bounded case studies rather than through sustained live industrial deployment. Consequently, reported performance improvements are contingent on the fidelity of the underlying training and testing environments, including the adopted state representation, reward formulation, and disturbance modelling assumptions. While these contributions establish algorithmic feasibility and adaptability under controlled conditions, the current evidence base remains limited in demonstrating robustness under yard-scale complexity, heterogeneous resource interactions, and long-term operational variability. In several cases, the term “real-time” refers to responsiveness within simulation or test environments rather than verified integration into continuously operating execution systems.
A third strand integrates learning with optimisation to improve scalability. Kwak et al. [
48] integrated ML to guide constraint satisfaction for long-term block planning, with feasibility checked through DES; Jeong et al. [
49] similarly linked predictive models to management-level decision support; and Li et al. [
23] illustrated combining forecasting with optimisation for project scheduling contexts. These hybrid configurations are methodologically promising because they couple predictive ML (to reduce uncertainty or prune the search space) with prescriptive optimisation (to generate schedules) and, in some cases, simulation (to verify feasibility). However, in most reported cases, validation is conducted through simulation experiments or bounded industrial case studies rather than longitudinal live-system deployment, which constrains the strength of inferences that can be drawn regarding long-term operational robustness or yard-scale generalisability.
Overall, the AI literature reviewed here supports a cautious but clear conclusion: ML is most mature as a surrogate-modelling and parameter-estimation layer, whereas RL is most mature as a simulation-trained policy generator. In both cases, evidence is strongest for feasibility and performance in controlled experimental settings, and weakest for sustained yard-wide operational impact, which requires integration with execution systems, disturbance-rich validation, and longitudinal KPI reporting.
Figure 2 summarises how ML-based predictive models, RL-based adaptive policies, and optimisation frameworks interact with DES/DTS, highlighting learning’s role in improving model fidelity, policy adaptability, and optimisation efficiency within simulation-based shipyard production planning.
3.2. DT Architectures and Cyber-Physical Integration
DT research extends the AI trajectory by focusing on cyber-physical integration: continuous state reconstruction, synchronisation, and (in some cases) closed-loop decision support. In principle, DT architectures can shift DES from offline analysis toward execution-linked monitoring and rescheduling. In practice, however, DT contributions span a wide maturity range, and a recurring methodological issue is the conflation of architectural completeness with valid1ated operational performance. For the purposes of this review, DT maturity is therefore judged primarily by demonstrable (i) data integration with execution systems, (ii) validation beyond simulation-only studies, and (iii) decision authority implemented within a loop that could plausibly affect operations. At the architectural end of the spectrum, Pang et al. [
50] proposed a yard-wide DT and Digital Thread framework intended to connect design, planning, and execution data. Such contributions clarify integration logic and information flows, but their evidentiary tier is typically conceptual/architectural, meaning they support conclusions about system design requirements rather than verified optimisation gains. More operationally oriented studies demonstrate partial cyber-physical linkage at the workshop scale. Wang et al. [
51] for example presented a resource-allocation DT for hull-part picking and processing in which a DES model was synchronised using IoT and QR-based tracking, enabling dynamic task reallocation. This class provides stronger evidence because it moves beyond static calibration toward state synchronisation and execution-linked dispatching, even if validation remains bounded to a specific workshop context. Similarly, Liang et al. [
52] described an IoT-driven DT for workshop monitoring, improving temporal resolution of equipment and process states; however, its decision scope was primarily situational awareness and early warning rather than closed-loop optimisation.
A central insight across DT implementations is that data mapping and synchronisation quality often dominate algorithmic sophistication; Wang et al. [
53] focused on multi-source fusion for DT–DES synchronisation and report performance-related improvements, while Sun et al. [
54] embedded optimisation within a DT–GA–DES loop for transporter scheduling under disruption scenarios. Liu et al. [
55] also formalised DT synchronisation across heterogeneous sources. These contributions indicate that DT performance claims are credible only insofar as (a) latency, missingness, and mapping errors are managed and (b) the resulting decisions are validated under realistic disturbance conditions rather than idealised scenarios.
Taken together, the DT literature reviewed here exhibits high architectural activity but uneven evaluative depth. Several studies report measurable improvements, yet the dominant evidence base remains prototype, simulation-synchronised, or partially validated rather than sustained yard-wide deployment with longitudinal production KPIs. This distinction matters because cyber-physical capability “in principle” does not, by itself, justify claims of operational transformation. To make this heterogeneity explicit,
Table 2 classifies DT and DT-enabling studies by system type, validation basis, and decision scope, thereby making differences in practical maturity and operational credibility explicit. Within this landscape, RL-based schedulers classified as DT-enabling studies demonstrate adaptability in simulated environments but typically rely on simplified representations and limited disturbance modelling, leaving open questions regarding robustness under yard-scale complexity and real operational uncertainty.
Table 2.
Comparative classification of Digital Twin and DT-enabling studies in shipyard production, by system maturity, validation basis, and decision scope (2015–2025).
Table 2.
Comparative classification of Digital Twin and DT-enabling studies in shipyard production, by system maturity, validation basis, and decision scope (2015–2025).
| Study | System Type | Validation Basis | Decision Scope |
|---|
| Pang et al. [50] | Core DT | Conceptual/architectural | Yard-wide (conceptual) |
| Wang et al. [51] | Core DT | Partial shop-floor validation (IoT, QR tracking) | Workshop-level control |
| Wang et al. [53] | Core DT | Partial shop-floor validation | Workshop-level scheduling and control |
| Sun et al. [54] | Core DT | Simulation with disruption scenarios | Transporter scheduling |
| Liang et al. [52] | Core DT | Partial shop-floor sensing | Workshop monitoring |
| Liu et al. [55] | Core DT | Conceptual/prototype | Yard/workshop (architectural) |
| Woo et al. [45] | DT-enabling | Simulation-only | Scheduling/load balancing |
| Nam et al. [47] | DT-enabling | Simulation-only | Scheduling optimisation |
| Le et al. [46] | DT-enabling | Simulation-only | Component-level control |
| Kwak et al. [48] | DT-enabling | Simulation with historical data | Long-term planning |
| Zhong et al. [56] | DT-enabling | Simulation-only | Stockyard layout and scheduling |
3.3. Digitalisation and Data Integration as Enablers of Intelligent Shipyard Optimisation
Digitalisation initiatives at organisational, logistics, and supply-chain levels form the infrastructural precondition for AI-, ML-, and DT-enabled optimisation in shipyards. Unlike DES or optimisation studies that target specific scheduling or layout problems, this strand of literature primarily addresses data architecture, interoperability, and information integration. Its contribution therefore lies less in immediate optimisation gains and more in enabling the conditions under which optimisation can become operationally credible.
Studies such as Strandhagen et al. [
57,
58,
59] analysed digitalised manufacturing logistics in engineering-to-order and maritime supply-chain contexts, outlining required information flows, system interfaces, and Industry 4.0 technologies for smart-yard development. Slobodian et al. [
60] and Slobodian and Kharytonov [
61] proposed Shipbuilding 4.0 information models and digital-transformation roadmaps, formalising enterprise- and cluster-level data platforms intended to support simulation, DT, and optimisation readiness. Rubio et al. [
62] similarly developed process-oriented PLM architectures that structure PPR data in formats compatible with DES and scheduling tools. These works demonstrate architectural feasibility and conceptual integration depth; however, validation is typically architectural or at the prototype-level rather than tied to longitudinal production KPIs. Their operational maturity therefore lies primarily in system design and interoperability specification rather than demonstrated optimisation impact.
At smaller scales, Studnev and Burmistrov [
37] showed that incremental digitalization, beginning with structured documentation and layout modelling, can progressively enable simulation-driven planning. This contribution clarifies a practical pathway toward optimisation readiness, particularly for small or medium shipyards. Nevertheless, evidence remains bounded to staged implementation examples rather than yard-wide performance transformation.
However, deeper digital integration also introduces structural vulnerabilities. Diaz et al. [
63] demonstrated that highly interconnected maritime and shipbuilding information ecosystems increase exposure to cybersecurity threats, data breaches, and system-level disruptions. From an evaluative standpoint, this implies that digital maturity and optimisation readiness are contingent not only on data richness and integration depth, but also on governance, cybersecurity robustness, and resilience. Consequently, architectural sophistication alone cannot be equated with operational credibility unless system security and stability are demonstrably ensured. A complementary strand focuses on spatial and measurement infrastructures that enhance data fidelity. Scarlat et al. [
36] highlighted Geographic Information System (GIS)-based geospatial modelling for maritime logistics and yard-layout planning, reinforcing the importance of spatially explicit data for realistic scheduling. Huang et al. [
64] propose high-precision spatial measurement-field technologies capable of feeding DT systems with continuous geometric and positional data. These contributions strengthen the physical-data layer necessary for spatial–temporal modelling credibility; however, they typically validate sensing accuracy or representation fidelity rather than optimisation outcomes. Their decision scope is therefore infrastructural and enabling, not directly prescriptive.
When synthesised across the AI-, ML-, and DT-oriented contributions reviewed above, a recurring architectural pattern becomes visible. At the physical layer, sensors, IoT devices, GIS platforms, and measurement-field technologies capture high-frequency operational states and spatial information [
36,
50,
52,
64]. At the virtual layer, DES models, predictive ML surrogates, and hybrid simulation frameworks represent workshop dynamics, logistics interactions, and process behaviour [
42,
43,
44,
51,
53]. At the intelligent layer, optimisation engines, constraint solvers, reinforcement-learning agents, and forecast–optimise pipelines translate these representations into scheduling and allocation decisions [
23,
45,
46,
47,
48].
This layered interpretation does not imply that all reviewed studies implement a complete three-tier architecture; rather, it reflects a structural convergence across otherwise heterogeneous contributions. The most operationally credible DT implementations are those in which these layers are tightly coupled through real-time synchronisation and behavioural validation, as demonstrated in disruption-aware scheduling or resource-allocation studies [
53,
54]. However, the evidentiary basis varies substantially: many digitalisation works remain at the conceptual or prototype level [
50,
57,
58,
59,
60,
61], some demonstrate partial shop-floor integration [
51,
52,
53], and only a limited subset report explicit optimisation KPIs under realistic operational conditions [
53,
54].
A methodological distinction therefore emerges between traditional DES/MOO studies and digitalisation-oriented research. DES- and optimisation-focused papers frequently quantify makespan, throughput, or cost improvements under structured case studies (
Table 1), whereas many digitalisation and DT publications emphasise interoperability, architectural integration, synchronisation fidelity, or information modelling rather than explicit production-performance gains [
50,
57,
58,
59,
60,
61]. Similarly, process-level ML studies report prediction accuracy or defect reduction [
42,
43], and RL studies report policy convergence, adaptability, or workload smoothing metrics [
45,
47], rather than yard-scale productivity indicators. Even within DT research, only a subset of works embed optimisation within cyber-physical loops and publish measurable operational improvements [
53,
54], while others prioritise integration feasibility and data mapping.
This divergence reflects differing research objectives rather than inherent methodological weakness. Nevertheless, it constrains inference. The current literature robustly supports the claim that digitalisation enhances data fidelity, integration depth, and the technical feasibility of closed-loop optimisation architectures. What remains less substantiated is sustained yard-wide productivity transformation attributable solely to digital integration, given limited longitudinal deployment evidence, restricted cross-yard benchmarking, and confidentiality constraints surrounding industrial KPIs [
63,
65].
Accordingly, digitalisation and data integration should be interpreted as enabling infrastructures rather than optimisation mechanisms in themselves. Their maturity is highest at the architectural and data-integration level, moderate at partial shop-floor synchronisation, and comparatively limited at the level of validated yard-scale optimisation impact.
Table 3 summarises the reviewed AI-, ML-, and DT-based technologies within this evidentiary landscape, illustrating a clear trajectory toward increasingly cyber-physical and data-driven production architectures built upon DES-based behavioural cores.
3.4. Shipyard-Oriented DT Architecture
DTs have increasingly been proposed as a comprehensive cyber-physical framework for real-time optimisation and production management in complex shipyard environments. Unlike standalone DES, ML, or optimisation tools, DT systems aim to integrate behavioural simulation, predictive analytics, and optimisation logic with live operational data streams and execution systems. The architectural requirements of such systems therefore warrant explicit articulation.
Synthesis of the reviewed literature suggests that an operationally credible shipyard-oriented DT requires at least six interacting layers: Physical, Execution, Engineering, Planning, Simulation & AI, and Human–Interface.
Figure 3 and
Figure 4 conceptualise and operationalise these layers at the workshop scale.
Across DT-enabled scheduling and execution-linked simulation studies, a recurring requirement for cyber-physical coherence is continuous synchronisation of shop-floor states, machine status, transporter positions, Work in Progress (WIP) progression, and operator activity, via MES and SCADA infrastructures [
32,
53,
54]. Engineering-layer consistency, supported by PLM/Product Data Management (PDM) systems and Bill of Materials (BOM)/Bill of Operations (BOP) structures, is similarly essential to prevent divergence between virtual process logic and actual production configurations. Studies addressing block logistics and stockyard coordination further demonstrate the necessity of semantic data mapping and state-reconstruction mechanisms capable of unifying heterogeneous yard data into a consistent digital representation [
35,
56]. The Simulation & AI layer, comprising DES/DTS models, optimisation solvers, and AI/ML components, relies on bidirectional data exchange with this synchronisation core to support predictive analytics, disturbance-aware scheduling, and scenario evaluation [
53,
54,
67]. Finally, decision-support interfaces provide the human–machine interaction layer through which optimised decisions are validated and enacted.
Importantly, while this multilayer architecture reflects structural convergence across the reviewed literature, not all studies implement a complete end-to-end DT stack. Validation depth varies considerably, and architectural completeness should not be conflated with operational maturity.
While the multi-layer DT architectures illustrated in
Figure 3 and
Figure 4 reflect structural convergence in the literature, the reported DT implementations differ substantially in modelling philosophy, learning integration, and evidentiary maturity. The reviewed studies can be grouped into three conceptual configurations.
The first category, analytical twins (DES-based DTs), consists of synchronised DES or discrete-time models grounded in rule-based or physics-based process logic without embedded learning components. Representative examples include the DT–GA–DES transporter-scheduling framework of Sun et al. [
54] and the MES-synchronised workshop DT implementations of Wang et al. [
51,
53].
The second category, ML-enhanced DTs, integrates trained surrogate models into explicit process representations to improve prediction fidelity in nonlinear subprocesses (e.g., Luo et al. [
42]).
The third category, RL-driven DTs (DES–RL hybrids), combines DES-based simulation environments with reinforcement-learning agents that learn adaptive dispatching or allocation policies (e.g., Woo et al. [
45]; Nam et al. [
47]).
The defining characteristics of these DT configurations are summarised in
Figure 5.
3.5. Comparison Between Traditional Simulation/Optimisation Approaches and DT/Digital Thread Frameworks
To clarify the methodological transition identified across
Section 2 and
Section 3,
Table 4 contrasts traditional DES/MOO approaches with emerging DT and Digital Thread frameworks through the lenses of integration depth, validation depth, and operational scope. Rather than implying categorical superiority, the comparison highlights structural differences in modelling assumptions, data coupling, execution-layer connectivity, and evidentiary tier. The objective is not to position DT as a replacement for DES/MOO, but to identify how digital-thread-enabled architectures reconfigure the relationship between simulation, optimisation, and live operational data.
Section 3 shows that AI, ML, RL, and DT technologies extend shipyard production modelling beyond the static, execution-decoupled character of traditional DES and optimisation. ML improves model fidelity through learned behavioural relationships; RL enables simulation-trained adaptive dispatching; and DT architectures deepen integration by synchronising virtual models with real-time IoT and MES data, supporting monitoring and, in some cases, closed-loop optimisation.
However, assessed in terms of validation depth, integration maturity, and decision scope, the evidence remains uneven. Many DT studies emphasise architectural integration and synchronisation accuracy rather than longitudinal KPI validation, while RL and ML applications are typically tested in bounded simulation environments. Conversely, traditional DES/optimisation studies report quantified gains but often lack live execution-layer integration. Thus, the literature convincingly establishes technical feasibility and integration potential, yet provides limited evidence of sustained yard-wide productivity transformation. The central tension lies between architectural sophistication and empirically validated operational impact.
4. Software Ecosystem for Simulation and Optimisation in Shipyard Production Research
The evolution of shipyard production modelling is shaped not only by methodological innovation in simulation and optimisation, but also by the software ecosystems through which these methods are operationalised. Across the reviewed literature, three interdependent tool categories recur: (i) discrete-event and hybrid simulation platforms that represent spatial–temporal production dynamics; (ii) optimisation engines and algorithmic frameworks that generate candidate schedules and allocation decisions; and (iii) execution and data-integration systems, such as MES, ERP, and PLM/PDM platforms, that provide the structured, time-resolved data required for model calibration, synchronisation, and, in advanced cases, DT implementation. The maturity and applicability of shipyard modelling approaches are therefore strongly conditioned by tool capabilities, interoperability, and data accessibility.
Section 4.1 reviews dominant simulation environments;
Section 4.2 examines optimisation engines and algorithmic paradigms; and
Section 4.3 analyses MES and enterprise platforms as enabling infrastructures that mediate between research prototypes and industrial deployment.
4.1. Simulation Tools in Shipyard Research
Commercial DES platforms, most prominently Tecnomatix Plant Simulation [
68] DELMIA QUEST [
69], ExtendSim [
70], and Arena [
71] continue to underpin a significant share of shipyard modelling studies. Their prevalence reflects modelling stability, mature event libraries, and compatibility with industrial data structures. However, their use also correlates with the availability of structured industrial datasets and institutional partnerships, indicating a linkage between software choice and evidentiary depth.
Tecnomatix Plant Simulation is particularly prominent in mid-term planning, energy-aware modelling, and DT-oriented research [
13,
19,
33,
53,
54]. Its alignment with PPR structures and the availability of the Simulation Toolkit Shipbuilding (STS) facilitate consistent modelling of cranes, transporters, stockyards, and welding systems. Advanced modules such as the Energy Analyzer extend performance evaluation beyond throughput toward energy and CO
2 metrics. These features support higher representational fidelity and partial execution-layer integration. Nevertheless, high licencing costs, data-structuring requirements, and modelling complexity limit accessibility and tend to confine its use to industrially supported case studies rather than exploratory academic experiments.
DELMIA QUEST (Dassault Systèmes) has played a central role in spatially intensive modelling, particularly block erection and crane-interference analysis [
9,
41,
72]. Its 3D kinematic capabilities enable high-fidelity validation of clearance, trajectory, and handling feasibility under constrained geometries. Such applications demonstrate strong spatial realism but typically remain within workshop-scale or scenario-based validation tiers. Proprietary architecture and limited openness to external optimisation libraries have contributed to gradual migration toward more scriptable and integration-friendly environments in recent DT-oriented research.
ExtendSim (ANDRITZ Inc.) and Arena (Rockwell Automation) occupy a second tier of DES tools, frequently used for workshop-level bottleneck analysis, layout comparison, and queue-dynamics studies [
10,
11,
12,
73]. Their strengths lie in rapid prototyping, accessibility, and suitability for diagnostic or educational applications. However, limited shipyard-specific libraries, weaker CAD/PLM integration, and constrained 3D capability restrict scalability toward yard-wide or DT-synchronised implementations. Accordingly, these tools are most prevalent in simulation-only or structured case-study tiers rather than execution-integrated research.
A notable methodological shift is the increasing adoption of Python- and SimPy-based DES environments. These open-source frameworks offer full programmability and seamless integration with ML, RL, and optimisation libraries, making them particularly suited to hybrid DES–AI research [
45,
47,
48]. In such studies, simulation environments serve not only as evaluation tools but also as generators of state-transition data for policy training. This architecture supports rapid experimentation and algorithmic innovation; however, the absence of native graphical interfaces and limited industrial-system integration often confine validation to simulation-based tiers.
Multi-method environments such as Simio, AnyLogic, and MATLAB/SimEvents appear in studies requiring agent-based modelling (ABM) or hybrid DES–ABM configurations. While conceptually suited to modelling decentralised crane coordination, Automated Guided Vehicle fleets, or distributed decision-making, their adoption in shipbuilding remains comparatively limited and frequently exploratory. Licencing constraints, modelling complexity, and the absence of shipyard-specific component libraries restrict their prevalence in validated industrial applications.
Table 5 summarises the principal simulation platforms identified in the reviewed literature, highlighting their functional strengths, integration characteristics, and practical constraints. Collectively, the evidence indicates that simulation fidelity is strongly influenced by tool choice, yet higher visual or modelling sophistication does not automatically correspond to higher validation maturity or operational integration.
4.2. Optimisation Algorithms and Engines in Shipyard Research
In parallel with simulation platforms, shipyard optimisation research employs a diverse range of computational engines and algorithmic paradigms. However, these approaches differ substantially in decision scope, scalability, validation depth, and operational maturity. When evaluated through these lenses, a differentiated methodological hierarchy becomes visible.
Mathematical optimisation engines such as CPLEX and Gurobi are widely used for mixed-integer and linear programming (MILP/LP) formulations of block sequencing, transporter scheduling, berth allocation, and resource coordination [
20,
21,
23,
31]. Their principal strength lies in formal optimality guarantees or tight bounds for well-structured subproblems. For narrowly defined scheduling contexts with manageable instance sizes, they provide high analytical credibility. However, the nondeterministic polynomial-hard (NP)-hard nature of shipbuilding planning, characterised by precedence networks, crane interference, blocking constraints, and spatial–temporal coupling, typically restricts these solvers to reduced-scale formulations or benchmark generation rather than full yard-scale deployment. Consequently, while mathematically rigorous, their operational scope is often bounded by abstraction and instance tractability.
CP occupies a complementary niche in problems dominated by combinatorial feasibility and rule consistency. In hybrid CP–DES studies (e.g., Kwak et al. [
30]), CP generates schedules satisfying precedence, shift, and due-date constraints, after which DES evaluates spatial feasibility and interaction effects. This “optimise first, simulate second” pattern enhances behavioural realism relative to pure analytical models. Nevertheless, validation typically remains at the simulation or bounded case-study level, and integration with live execution systems is uncommon. CP therefore demonstrates strong expressive power but moderate operational maturity in practice.
Metaheuristics constitute the most frequently adopted optimisation class in the reviewed literature. Genetic Algorithm (GA), NSGA-II/III, multi-objective Tabu Search (MOTS), improved Grey Wolf Optimisation (IGWO), improved whale-based algorithms (IGWOA), and related population-based methods are applied across layout optimisation, green logistics, welding scheduling, block-flow coordination, and yard design [
20,
21,
22,
23,
24,
27,
29]. Their strength lies in robustness under nonlinear, multi-objective, and spatially constrained conditions, particularly when DES-based fitness evaluation is embedded. This compatibility with simulation enables richer representation of crane interference, congestion, and stochastic arrivals than purely analytical approaches. However, metaheuristic performance is problem-dependent, parameter-sensitive, and lacks guarantees of global optimality. Reported improvements are frequently demonstrated in structured simulation experiments rather than disturbance-rich longitudinal deployment. Thus, while metaheuristics exhibit high adaptability and broad applicability, their evidentiary strength is typically empirical rather than formally validated.
ML and RL represent the most recent additions to the optimisation ecosystem and introduce adaptive policy-generation capabilities. RL-based schedulers trained within DES environments can learn disturbance-responsive dispatching strategies and outperform static heuristics under stochastic test conditions [
45,
46,
47,
53]. ML-based predictors improve duration estimation, resource forecasting, or forming accuracy, thereby strengthening downstream optimisation reliability [
48,
66]. From an evaluative standpoint, these approaches demonstrate algorithmic feasibility and adaptability in simulation-based environments. However, validation is predominantly confined to controlled or workshop-scale contexts, and robustness under heterogeneous yard-scale operations remains less substantiated. Their maturity therefore depends critically on integration with execution systems and disturbance-rich validation rather than purely on algorithmic sophistication.
Hybrid DES–optimisation architectures represent the most methodologically mature configurations in the reviewed literature. Examples include DES-validated GA scheduling [
41], CP-DES hybrid assembly planning [
30], APS + DES mid-term scheduling [
32], and DT-integrated DES-GA and DES-RL decision systems [
53,
54]. In these configurations, optimisation engines generate candidate solutions while DES enforces behavioural feasibility under spatial and temporal constraints. This coupling reduces the risk of infeasible analytical solutions and strengthens operational credibility relative to standalone optimisation. Nevertheless, even in hybrid systems, validation depth varies: many implementations remain simulation-synchronised or at the workshop scale rather than having sustained yard-wide deployment.
Table 6 summarises the optimisation engines and algorithmic approaches identified across the literature, organised by methodological type, representative applications, and practical strengths. Importantly, the comparison should not be interpreted as a linear progression toward superiority; rather, it reflects differing trade-offs between analytical rigour, scalability, adaptability, integration maturity, and validation depth.
Overall, the optimisation ecosystem in shipyard research is methodologically diverse but uneven in operational maturity. Mathematical solvers provide strong analytical guarantees under structured formulations, yet are often applied to reduced-scale or abstracted representations. Metaheuristics offer flexibility for high-dimensional and multi-objective problems, though their performance evidence is typically empirical and scenario-dependent. ML and RL introduce adaptive and data-driven capabilities, but validation remains predominantly simulation-based or workshop-bounded. Hybrid DES–optimisation architectures demonstrate comparatively higher behavioural credibility by embedding feasibility checks within optimisation loops; however, longitudinal yard-wide deployment remains sparsely documented.
Taken together,
Table 5 and
Table 6 indicate that while the software and algorithmic ecosystem enable increasingly sophisticated modelling configurations, methodological sophistication does not uniformly translate into validated operational impact. The principal differentiator across approaches lies not in algorithmic novelty alone, but in the depth of integration with execution systems and the robustness of empirical validation under realistic disturbance conditions.
4.3. MES, Production-Control and Data-Integration Platforms
Although MES, ERP, and PLM/PDM systems are not optimisation tools in themselves, they constitute the infrastructural layer upon which simulation, optimisation, and DT architectures depend. Across the reviewed digitalisation and DT literature [
39,
50,
55,
57,
58,
59,
60,
61,
62], these platforms are consistently positioned as prerequisites for execution-layer integration and state synchronisation. However, the mere presence of such systems does not, by itself, imply validated optimisation capability; rather, their relevance depends on how deeply they are coupled with simulation and decision-support loops.
In operational shipyards, production data are embedded within enterprise systems that manage WIP states, BOM/Bill of Resources (BOR) structures, routing definitions, resource calendars, and task assignments. MES platforms such as Floor2Plan provide time-stamped execution data that can calibrate DES models or support DT synchronisation. Similarly, AVEVA MES integrates routing and configuration information inherited from CAD and PLM environments, enabling structured PPR representations that can be transferred into simulation models. These contributions enhance data granularity and consistency; however, published studies rarely provide systematic evidence of longitudinal yard-wide optimisation gains directly attributable to MES–DES or MES–DT coupling.
ERP systems, including industrial and financial systems (IFSs) encode higher-level constraints such as material availability, labour capacity, and cost structures. These parameters define the feasible search space within MILP, CP, and metaheuristic optimisation models. Their influence is therefore structural rather than algorithmic: they shape boundary conditions for planning but do not guarantee improved scheduling performance unless tightly integrated with validated simulation or control layers.
PLM/PDM systems, most prominently Siemens Teamcenter, maintain product–process definitions that support PPR-based model generation. Integration between Teamcenter and platforms such as Tecnomatix or DELMIA reduces modelling effort and improves engineering–planning consistency. From an evaluative standpoint, this strengthens architectural coherence and reduces data-fragmentation risk. Yet, similar to MESs, published evidence predominantly documents interoperability feasibility and prototype integration rather than disturbance-rich, yard-scale optimisation validation.
Table 7 summarises the principal enterprise platforms referenced in the reviewed studies and clarifies their functional role relative to simulation and optimisation.
Overall, the software ecosystem underpinning shipyard simulation and optimisation research exhibits high technical diversity but uneven operational integration maturity. Commercial DES platforms provide behavioural richness and industrial alignment, yet are often constrained by data-availability and deployment scale. Open-source environments facilitate methodological innovation and AI integration, though validation frequently remains simulation-bounded. Mathematical solvers and metaheuristics enable formal and large-scale search capabilities, but their practical credibility depends on behavioural validation through DES or DT coupling. Enterprise platforms such as MES, ERP, and PLM establish the necessary data infrastructure for cyber-physical integration; however, the published literature more often demonstrates architectural connectivity than sustained yard-wide optimisation impact.
Consequently, the principal constraint across the ecosystem is not algorithmic capability but integration depth and validation structure. While toolchains capable of supporting fully integrated digital-thread or DT architectures now exist in principle, empirical evidence of longitudinal, disturbance-rich, yard-scale deployment remains limited. This gap between infrastructural readiness and validated operational transformation directly motivates the data-access and evaluation challenges discussed in
Section 5 and informs the research gaps articulated in
Section 6.
5. Availability and Role of Real Industrial Data in Shipyard Simulation and Optimisation Research
A structural constraint across shipyard production-modelling research is the limited availability of openly accessible, high-resolution industrial datasets. Although many studies report the use of “real shipyard data,” the underlying numerical parameters are rarely published in reproducible form. Confidentiality agreements, the proprietary architecture of APS, MES, ERP, and PLM systems, and the commercial or defence sensitivity of shipbuilding operations prevent disclosure of detailed processing times, routing matrices, crane trajectories, transporter capacities, spatial layouts, and interference rules. As a consequence, only a small subset of publications provide sufficient quantitative detail to allow independent reconstruction of test cases for benchmarking DES engines, optimisation solvers, or hybrid architectures.
From an evaluative perspective, this opacity directly constrains the validation tier that studies can legitimately claim. Where numerical datasets cannot be independently reconstructed, evidence is necessarily case-specific and non-replicable, limiting cross-study comparability and weakening the development of a cumulative performance hierarchy. This limitation is particularly consequential in shipbuilding, where system behaviour is highly sensitive to block geometry, yard topology, resource-interference patterns, congestion dynamics, and stochastic arrival processes. Without transparent parameter sets, methodological robustness cannot be systematically compared across alternative modelling paradigms.
5.1. Availability of Reconstructible Industrial Data and Confidentiality Barriers
Only a limited number of shipyard-modelling studies publish quantitative data with sufficient granularity to enable independent reconstruction, benchmarking, or methodological reuse. These studies, summarised in
Table 8, constitute the primary openly accessible numerical foundations for DES-, MOO-, or hybrid simulation–optimisation research in this domain. Examples include workshop-level cutting and welding parameters, transporter-routing networks and stockyard occupancy configurations, and explicit processing and setup times for production workflows [
11,
35,
67]. Certain DT-oriented contributions provide partial operational detail, such as transporter cooperation constraints, time-window definitions, workstation states, or workload distributions, which permits approximate reconstruction under bounded assumptions [
53,
54]. Dynamic scheduling studies, including Zhong et al. [
56], further describe stockyard layouts, crane operating zones, and arrival patterns relevant to spatial–temporal modelling.
However, even within this subset, data disclosure is typically confined to isolated workshops or subsystem-level processes. Fully integrated yard-wide datasets, linking engineering structures, execution states, logistics flows, and disturbance dynamics, are not publicly available. Consequently, while these studies enable methodological benchmarking under structured conditions, they do not support validation of large-scale cyber-physical or end-to-end optimisation architectures.
By contrast, the majority of shipyard-modelling studies rely on confidential industrial datasets embedded within enterprise platforms. These datasets contain detailed product structures, routing logic, thousands of operation records, shift calendars, resource constraints, and spatial configurations that cannot be disclosed without violating commercial or defence obligations. The resulting opacity severely constrains reproducibility and cross-study benchmarking. While structural abstractions such as WBS-driven schemas, PPR frameworks, and automated DES model-generation pipelines improve conceptual standardisation [
62,
74,
75,
76,
77], they do not substitute for access to disturbance-rich numerical data.
Importantly, confidentiality does not eliminate the possibility of conceptual evaluation; architectural coherence, integration logic, disturbance modelling assumptions, and validation design remain assessable. However, the absence of shared benchmark datasets limits the development of a stratified evidence hierarchy capable of distinguishing between exploratory demonstrations, bounded industrial pilots, and sustained yard-scale optimisation impact. In practice, rigorous DES, optimisation, and DT research therefore continues to depend on close industry collaboration under non-disclosure agreements, reinforcing the gap between architectural capability and reproducible empirical validation identified in earlier sections.
5.2. Global Landscape of Shipyard Digitalisation and Data-Driven Technologies
Evidence from academic publications, industry reports, and corporate disclosures indicates that digitalisation, smart-yard initiatives, and DT programmes are being actively pursued across major shipbuilding regions.
Figure 6 presents a non-exhaustive geographical overview of DT-related activities synthesised from peer-reviewed studies, technical reports, and publicly available industrial communications. Academic research output is most concentrated in South Korea, China, Japan, and several European countries, consistent with the regional distribution of shipbuilding capacity and research investment identified throughout this review.
Industrial initiatives are progressing in parallel, though the evidentiary basis differs substantially from that of peer-reviewed studies. In South Korea, major shipbuilders including HD Hyundai Heavy Industries, Daewoo Shipbuilding & Marine Engineering (DSME), and Samsung Heavy Industries have publicly reported smart-yard and DT-oriented programmes focused on automation, IoT integration, and cyber-physical coordination [
78,
79,
80,
81]. In China, large yards under the China State Shipbuilding Corporation (CSSC), such as Jiangnan Shipyard, Hudong-Zhonghua Shipbuilding, Guangzhou Shipyard International, Dalian Shipbuilding Industry Company, and Shanghai Waigaoqiao Shipbuilding, are actively engaged in IoT-enabled workshop digitalisation, spatial sensing, and data-driven production optimisation [
82]. In Japan, members of the DT Project Consortium (including Mitsui E&S, Sumitomo Heavy Industries Marine & Engineering, and Kyokuyo Shipyard) are pursuing DT-enhanced design-production integration [
83]. Additional initiatives are documented in Singapore through the ABS “Smart Technologies for Shipyards” programme, and across Europe through activities reported by Meyer Werft, Navantia, Fincantieri, and related stakeholders [
84,
85,
86].
Mapping these developments geographically therefore serves primarily to contextualise the structural constraint identified in
Section 5.1. Despite widespread adoption narratives and visible investment in digital infrastructures, openly accessible, reconstructible industrial datasets remain extremely scarce. Most documented progress appears to occur within confidential industry–academia collaborations or proprietary enterprise environments, limiting cross-yard benchmarking and comparative methodological evaluation.
Taken together,
Section 5 indicates a dual reality: digitalisation and DT adoption are advancing globally at architectural and strategic levels, yet the empirical evidence base required for reproducible benchmarking and rigorous comparative validation remains underdeveloped. This asymmetry reinforces the central tension identified throughout the review between technological capability in principle and demonstrable, transferable operational impact in practice.
6. Discussion: Research Gaps, Future Directions, and Limitations of the Present Work
To synthesise the heterogeneous body of reviewed work, the analysed methods can be positioned along four evaluative dimensions: (i) methodological role (behavioural modelling, optimisation, prediction, or control), (ii) decision level (layout design, planning, scheduling, or operational control), (iii) data dependency (offline, historical-data-driven, or real-time synchronised), and (iv) degree of system integration (standalone analytical tools, hybrid simulation–optimisation frameworks, or DT-enabled cyber-physical systems).
These dimensions are not merely classificatory; they enable differentiation of methodological maturity, validation depth, and operational credibility across the literature. By mapping approaches against these axes, it becomes possible to distinguish experimentally validated optimisation frameworks from architecturally complete yet weakly validated cyber-physical proposals, and to identify where claims of “intelligent” or “real-time” optimisation are supported by demonstrable execution-layer evidence. This framework therefore underpins the critical assessment of research gaps, future directions, and the limitations of the present review.
6.1. Research Gaps: Methodological Fragmentation, Limited Integration, and Data Constraints
Across the reviewed studies, three structural research gaps consistently emerge: methodological fragmentation, limited system integration, and severe data constraints. These gaps are not merely organisational shortcomings; they directly affect the strength of inference that can be drawn regarding yard-scale operational transformation.
First, a pronounced divide persists between traditional DES/MOO research and the more recent DT-, AI-, and ML-oriented literature. Classical DES and optimisation studies typically report quantifiable improvements in makespan, material-handling cost, energy use, or resource utilisation under clearly defined case conditions [
20,
21,
24,
27]. Their evidentiary strength lies in explicit performance metrics, even when execution-layer integration remains limited. In contrast, much of the DT literature emphasises architectural completeness, data-integration frameworks, and synchronisation fidelity, with benefits expressed in terms of mapping accuracy, cyber-physical coherence, or real-time visibility rather than sustained production KPIs [
50,
57,
58,
59,
60,
61,
63]. Only a small subset of studies embed optimisation or learning-based control within DT-synchronised DES environments and report measurable operational gains (e.g., Sun et al. [
54] and Wang et al. [
53]). As a result, the field exhibits a hierarchy of maturity: optimisation-driven DES hybrids demonstrate validated performance within bounded contexts, whereas many DT architectures remain conceptually robust but operationally under-evaluated. The two streams continue to evolve largely in parallel, and fully integrated, yard-scale, performance-validated cyber-physical systems remain rare.
Second, the scope of most studies remains structurally narrow. DES-, optimisation-, and ML/RL-based contributions predominantly target isolated workshops or subprocesses, cutting–welding shops, profile lines, stockyards, forming cells, nesting operations, or robotic surface preparation [
11,
13,
35,
42,
43,
44,
46]. While methodologically rigorous within their defined boundaries, these models seldom capture the multi-workshop, dock-level, supplier-linked, and logistics-coupled interactions characteristic of engineering-to-order shipbuilding. Vertical integration between APS planning layers and execution-level DES/DT/RL control is limited [
32,
33], and horizontal integration across docks, subcontractors, and yard logistics networks remains embryonic [
57,
58,
65]. Consequently, many reported improvements are locally valid but systemically untested. The absence of multi-scale, yard-wide models restricts the ability to evaluate emergent bottlenecks, cross-workshop coupling effects, and energy-logistics interdependencies that define real shipyard complexity.
Third, the near-absence of openly available, high-resolution industrial datasets constitutes a structural constraint on scientific consolidation. Only a small minority of studies publish reconstructible quantitative parameters, and these are typically confined to workshop-level subsystems rather than integrated yard-scale configurations. Most DES, optimisation, and DT research relies on confidential APS, MES, and ERP data that cannot be disclosed. Although modelling standards such as WBS/PPR frameworks enhance structural consistency, they do not resolve the lack of numerical transparency. The result is limited reproducibility, restricted cross-study benchmarking, and fragmented methodological comparison. Without anonymised, simulation-ready benchmark datasets calibrated to realistic shipyard patterns, claims of algorithmic superiority or DT-enabled optimisation remain difficult to generalise beyond case-specific collaborations.
Taken together, the evidence suggests a differentiated maturity landscape. Hybrid DES–optimisation frameworks currently represent the most methodologically validated and operationally credible class of approaches within bounded planning contexts. AI- and DT-based systems, while architecturally promising and increasingly sophisticated, function primarily as enabling infrastructures whose yard-scale optimisation impact has yet to be demonstrated through longitudinal, disturbance-rich validation. Bridging this maturity gap remains the central methodological challenge for the field.
6.2. Future Research Directions
The gaps identified above suggest a structured research agenda centred on integration maturity, evidentiary depth, and system scale rather than incremental algorithmic refinement.
Transition from workshop demonstrators to yard-scale, multi-layer DT architectures. Future DT research must move beyond isolated workshop pilots toward integrated, yard-wide cyber-physical systems. Architectures should synchronise planning, logistics, workshops, stockyards, docks, and energy subsystems within unified data environments supported by robust semantic mapping and real-time MES/ERP/PLM/IoT connectivity. Interoperable PPR/WBS representations capable of automated model generation and yard-scale state reconstruction are critical prerequisites. Although European initiatives such as ECOSHIPYARD and DISY-UK indicate growing institutional momentum [
87,
88], most implementations remain exploratory, and validated whole-yard DT deployments are still limited.
Embed optimisation and RL within execution-linked DT loops. While ML, RL, and metaheuristics demonstrate promising performance in controlled DES environments, few studies report sustained closed-loop deployment within live DT infrastructures [
22,
45,
46,
47]. Future research should prioritise explainable, safety-aware, and operator-supervised RL controllers embedded within hybrid DT–DES ecosystems. Validation should extend beyond simulation-only evaluation to disturbance-rich, execution-layer-linked trials capable of demonstrating robustness under operational uncertainty.
Establish anonymised benchmark datasets and standardised test problems. Methodological comparability is currently constrained by the absence of shared datasets reflecting shipyard-specific constraints such as crane interference, transporter blocking, stockyard geometry, spatial congestion, and stochastic arrivals. Developing benchmark families derived from anonymised industrial patterns, template-based synthetic generation, or parameterised reference yards would enable reproducible evaluation of DES, optimisation, and AI methods. Without such benchmarks, claims of superiority remain largely case-specific.
Integrate sustainability and energy modelling into production optimisation. Although green scheduling and energy-aware logistics are emerging themes [
20,
21,
22,
24], shipyard-wide Digital Energy Twins remain largely unexplored. Future DES/DT frameworks should incorporate electrification pathways, hybrid-energy configurations, CO
2 accounting, and life-cycle performance metrics. Linking production scheduling with energy-system optimisation would support decarbonisation strategies and regulatory compliance while improving long-term infrastructure planning.
Expand modelling boundaries to multi-yard and supply-network scales. Current models are predominantly yard-internal. Extending DES/DT architectures to supplier-linked, multi-yard, and cluster-level ecosystems could enable coordinated optimisation of material flows, subcontractor production, inventory buffers, port interfaces, and energy exchanges. Building on work in digitalised maritime supply networks [
36,
57,
58], such integration would align shipyard optimisation with emerging visions of digitally connected maritime industrial ecosystems.
Couple energy-system design tools with production simulation. No reviewed study explicitly integrates energy-system planning tools within shipyard production models. Platforms such as HOMER or DER-CAM could be coupled with DES/DT environments to evaluate electrification scenarios, distributed generation, storage strategies, and low-carbon infrastructure configurations. This integration would enable holistic assessment of production–energy interdependencies at the yard scale.
6.3. Limitations of the Present Review and Directions Beyond Its Scope
While this review provides a comprehensive synthesis of simulation-, optimisation-, AI/ML-, and DT-oriented production-modelling methods, its scope is intentionally confined to analytical and decision-support frameworks. It does not systematically examine organisational, economic, or socio-technical transformation processes in shipyards. Nor does it attempt comparative analysis by ship type, yard scale, production strategy (e.g., ETO/CTO/STO), or geographic region, as such comparisons would require consistent operational metadata that is rarely disclosed in the literature.
A further limitation arises from the evidentiary structure of the field itself. Many studies are based on confidential industrial datasets and report limited contextual detail regarding yard characteristics, workforce structure, legacy IT maturity, or capital constraints. Consequently, cross-yard benchmarking and comparative generalisation remain constrained. The review therefore emphasises methodological capability and architectural maturity rather than empirical benchmarking of transformation outcomes.
Industry-oriented analyses indicate that digital adoption is shaped by fragmented legacy infrastructures, limited MES/ERP/PLM interoperability, high capital intensity, skills shortages, and uncertainty regarding transformation sequencing [
37,
57,
60,
62,
89]. These factors directly influence the integration depth and validation maturity of simulation and DT implementations but are only partially visible in modelling-focused publications. As a result, conclusions drawn in this review pertain primarily to methodological feasibility rather than organisational readiness or implementation success.
Socio-technical dimensions, including operator acceptance, training requirements, ergonomic implications, safety integration, and human-in-the-loop decision support, are increasingly emphasised in Industry 5.0 paradigms promoting human-centric and resilient manufacturing [
90,
91,
92]. However, these aspects fall largely outside the modelling scope of the reviewed studies and therefore beyond the analytical boundary of this review.
Finally, the literature suggests a practical trade-off between model fidelity and deployability. High-resolution DES and DT implementations often require extensive data preparation, site-specific configuration, and iterative calibration against real operations [
32,
33,
57,
58]. In some cases, modular or reduced-complexity models may offer greater practical value than highly detailed architectures that are difficult to maintain or synchronise [
18,
37]. Data quality, latency, and MES/ERP/PLM integration constraints further affect the sustainability of real-time decision-support systems [
50,
61,
65]. These factors are rarely quantified explicitly, yet they substantially shape whether digital modelling approaches transition from research prototypes to sustained industrial use.
7. Conclusions
This review has examined fifteen years of methodological development in shipyard production planning and control, tracing the progression from isolated workshop-level simulation toward increasingly integrated, data-driven, and cyber-physical production architectures. By synthesising research on DES/DTS, mathematical optimisation and metaheuristics, AI/ML and RL, and DT technologies, the paper clarifies how these streams have co-evolved within the constraints of shipyard-specific spatial, logistical, and organisational complexity.
A consistent finding is that DES remains the behavioural backbone of shipyard production modelling. It uniquely captures stochastic interactions, spatial interference, and resource coupling, which are structurally difficult to encode in purely analytical formulations. Over time, DES has evolved from a diagnostic bottleneck-analysis tool into a decision-support core supporting scheduling, layout planning, energy-aware optimisation, and DT synchronisation.
Optimisation methods, particularly multi-objective formulations and metaheuristics, complement DES by enabling systematic exploration of trade-offs among productivity, cost, energy, and environmental performance. Among the reviewed approaches, hybrid simulation–optimisation frameworks exhibit the highest operational credibility, as they reconcile mathematical search with behavioural feasibility validation. These methods consistently report measurable performance gains under structured case conditions.
AI/ML and RL extend this foundation by introducing predictive and adaptive capabilities. ML enhances model fidelity by replacing static parameters with learned surrogates, particularly in process-intensive domains. RL demonstrates adaptive scheduling potential within simulation environments. However, validation remains predominantly simulation-based, and evidence of sustained yard-scale deployment is limited. Thus, while algorithmic capability is advancing rapidly, operational robustness at the full-yard scale remains insufficiently demonstrated.
DT technologies represent the most ambitious integration paradigm, combining simulation, optimisation, and real-time data streams within cyber-physical architectures. The literature shows clear progress in architectural integration and state synchronisation. Nevertheless, a recurring methodological tension persists: architectural sophistication often exceeds longitudinal KPI validation. Cyber-physical integration alone does not guarantee optimisation impact; demonstrable performance improvements remain closely tied to rigorous simulation–optimisation foundations and disturbance-aware validation.
Across all methodological families, structural constraints continue to limit field maturity. Fragmented execution-layer integration, predominance of workshop-scale studies, and severe data confidentiality barriers constrain benchmarking, reproducibility, and yard-wide validation. Moreover, high-fidelity DES and DT models impose significant calibration and maintenance burdens, reinforcing trade-offs between model realism and deployability.
Viewed through the lenses of methodological role, decision scope, validation depth, and integration maturity, the field reveals a clear evolutionary trajectory, but also a differentiated evidence hierarchy. Hybrid DES–optimisation systems currently represent the most empirically substantiated class of solutions, while AI- and DT-enabled architectures function primarily as enabling infrastructures whose full yard-scale optimisation impact remains to be robustly demonstrated.
The intellectual contribution of this review lies in making this hierarchy explicit. Rather than treating simulation, optimisation, AI, and DT technologies as uniformly progressive layers of innovation, the analysis distinguishes between architectural promise and validated operational performance. Future advancement will depend less on introducing new algorithmic techniques and more on strengthening integration depth, disturbance-rich validation, cross-yard benchmarking, and sustained industry–academic collaboration capable of translating methodological innovation into demonstrable production transformation.