Next Article in Journal
Regulatory Feedback and Adaptive Constraints in Publicly Funded R&D Project Management Systems: A Multicriteria Decision Analysis
Previous Article in Journal
Subsample Analysis of Oil Revenue Shocks and Macroeconomic Policy Transmission
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Industry 4.0/5.0 Maturity Models: Empirical Validation, Sectoral Scope, and Applicability to Emerging Economies

by
Dayron Reyes Domínguez
1,*,
Marta Beatriz Infante Abreu
2 and
Aurica Luminita Parv
1,*
1
Department of Manufacturing Engineering, Transilvania University of Brasov, 500036 Braşov, Romania
2
Department of Business Informatics, Technological University of Havana “José Antonio Echeverría” (CUJAE), La Habana 19390, Cuba
*
Authors to whom correspondence should be addressed.
Systems 2026, 14(2), 134; https://doi.org/10.3390/systems14020134
Submission received: 23 December 2025 / Revised: 20 January 2026 / Accepted: 22 January 2026 / Published: 27 January 2026
(This article belongs to the Section Systems Practice in Social Science)

Abstract

This article presents an academic literature analysis of 75 Industry 4.0 (I4.0) and Industry 5.0 (I5.0) maturity models published between 2020 and 2024, examining their empirical validation, sectoral scope, geographical origin, and stated applicability to developing-country contexts. The study combines descriptive profiling, contingency-table analyses with exact tests and effect sizes, and a large-scale synthesis of 562 research gaps reported by model authors. Knowledge production is highly concentrated in single-country studies (77.3%) and in developed economies, while most models do not explicitly or implicitly document applicability to developing-country settings (approximately 83%). Empirical validation practices are uneven, with multiple-case studies (33.3%) and surveys (24.0%) dominating, and sectoral coverage is strongly skewed toward manufacturing, limiting transferability to other sectors relevant for emerging economies. A statistically detectable association is observed between the development level of the model’s country of origin and the presence of applicability statements ( χ 2 = 17.13, p < 0.05 , moderate effect size), whereas authorship configuration shows no substantive association. Thematic analysis of reported gaps highlights persistent deficits in empirical rigor, sectoral breadth, SME orientation, operationalization of human-centric and sustainability dimensions associated with Industry 5.0, availability of implementation tools, and longitudinal or predictive evidence. The article concludes by outlining a research agenda focused on context-aware validation designs, broader sectoral grounding, and greater transparency and reproducibility, supported by open access to all underlying data, codebooks, and taxonomies.

1. Introduction

The digital transformation landscape in developing countries is both complex and dynamic, offering opportunities for growth and diversification amid structural challenges and limited strategic capabilities [1,2]. The paradigm shift brought by Industry 4.0/5.0 enables these countries to accelerate their economic development and, in some cases, leapfrog stages that developed nations have experienced sequentially [3]. Digital policies in African countries, for example, explicitly frame digitalization as a lever for productivity, job creation, and a transition toward knowledge-based economies [1]. At the same time, Artificial Intelligence (AI) and related Fourth Industrial Revolution (4IR) technologies are emerging as new factors of production, complementing human and physical capital and enabling more efficient, data-driven processes [4]. These technologies have the potential to expand internet access, increase the number of knowledge workers, and enhance the use of smart products with fewer errors, thereby strengthening competitiveness and innovation in developing-country contexts [4].
Digitalization also affects labor markets, trade, and social inclusion. Evidence points to positive contributions to job creation and reductions in unemployment, including a favourable impact on female employment through more flexible work arrangements such as teleworking [4]. At the same time, digitalization can help reduce poverty and improve living standards by generating new occupations for socially and economically marginalized groups [4]. In terms of trade, it reduces transaction costs and facilitates diversification across tasks, products, and sectors [1,4], helping to overcome persistent specialization patterns and enabling developing countries to leverage latent comparative advantages and expand exports to developed markets [1].
Despite this high potential, developing countries face substantial barriers that hinder the adoption of Industry 4.0/5.0 technologies and the realization of their benefits [1,2]. These barriers include inadequate digital infrastructure, high ICT service costs, limited scientific and technical capabilities, skills mismatches, and institutional and governance challenges. In addition, the complexity of 4IR technologies, their concentration among a small number of leading firms, and the need for complementary investments in skills, services, and organizational change further constrain absorption in developing-country contexts [2]. These challenges underscore the need for structured, context-aware approaches to guide digital transformation.
In this context, maturity models (MMs) have become widely used tools to support organizations in navigating digital transformation pathways [5]. MMs provide structured frameworks that allow organizations to assess their current digital capabilities and define staged trajectories toward more advanced states [6,7,8,9,10,11]. In resource-constrained environments, such as those prevalent in many developing countries, maturity models play a particularly important role by supporting prioritization, phased investment, and strategic planning, rather than ad hoc or overly ambitious transformation initiatives.
Beyond technological assessment, more advanced maturity models incorporate organizational, human, and sustainability dimensions, aligning with the broader Industry 5.0 agenda [9,11,12]. By integrating aspects such as workforce capabilities, resilience, and social and environmental sustainability, these models aim to support more holistic and inclusive transformation pathways. However, the extent to which existing maturity models adequately reflect diverse sectoral realities, development contexts, and empirical evidence remains an open question.
Recent literature reviews on Industry 4.0/5.0 maturity and readiness models have mapped available frameworks and their conceptual dimensions, and have begun to incorporate Industry 5.0 perspectives into this landscape [13,14,15,16]. These reviews have highlighted persistent sectoral and geographical biases, with most models designed and validated in manufacturing firms located in industrialized economies. They also note the limited availability of standardized instruments, open datasets, and robust empirical validation. Nevertheless, most existing reviews rely primarily on descriptive comparisons and narrative synthesis, offering limited systematic analysis of empirical validation practices, sectoral coverage, or explicit consideration of developing-country applicability.
Building on this literature, our previous study published in Sustainability [17] developed an academic literature analysis of 75 Industry 4.0/5.0 maturity models, focusing on their conceptual structure, standardized dimensions, maturity levels, and enabling technologies, and proposing a meta-typology of hybrid I4.0–I5.0 designs. While that work concentrated on conceptual design choices and classificatory properties, it did not examine how these models are distributed geographically, how they have been empirically validated, or to what extent they explicitly address applicability in developing-country contexts.
The present article makes a distinct and non-overlapping contribution. Although it shares an initial review design and document-retrieval workflow with the Sustainability study, its analytical scope and research questions differ substantially. Specifically, this article shifts the focus from conceptual typologies to contextual, empirical, and geographical dimensions of maturity-model research. It introduces original analyses that are not addressed in the companion paper, including: (i) a sensitivity analysis of relevance-based database ranking; (ii) inferential and exact-statistical testing of associations between authorship configuration, development level, sectoral focus, and explicit statements on developing-country applicability; (iii) a normalized sectoral analysis with goodness-of-fit testing; and (iv) a large-scale, auditable synthesis of research gaps reported by model authors, supported by AI-assisted extraction and systematic human verification. As such, the present manuscript is designed to be read as a standalone contribution that extends existing reviews toward evidence-based insights on the transferability and contextual grounding of Industry 4.0/5.0 maturity models, particularly for emerging economies.
Against this background, this study analyzes 75 maturity models published between 2020 and 2024 to address the following research questions: where these models are produced and how developing countries are represented; how authorship type and development context relate to explicit applicability statements; how extensively models have been empirically validated and in which sectors; and which research gaps and limitations authors themselves identify as barriers to Industry 4.0/5.0 adoption in developing economies (Table 1). The remainder of the paper is organized as follows. Section 2 describes the review protocol and coding procedures. Section 3 presents the empirical findings. Section 4 discusses implications for theory and practice, with particular attention to developing-country contexts and Industry 5.0 dimensions. Section 5 concludes and outlines directions for future research.

2. Materials and Methods

The research was carried out based on an exhaustive analysis of existing maturity models. The method consists of five sequential phases with parallel branches (see Figure 1). First, the review design (1) is defined, establishing the sources, time frame, and inclusion/exclusion criteria to build the corpus. Then, the variables and categories are operationalized (2). Subsequently, data extraction, normalization, and coding are performed (3), yielding a dataset that serves as input for the analysis. Next, the analysis phase is conducted, comprising four branches that represent four groups of tools and techniques (4): descriptive statistics; visualization and bibliometrics; inferential statistics; and text mining and qualitative analysis. To facilitate traceability and reproducibility of the method, each phase has been named as shown in Figure 1. Each phase is detailed below, along with the tools, methodologies, and software that supported them.
In both this article and our previous Academic Literature Analysis of Industry 4.0/5.0 maturity typologies and hybrid I4.0–I5.0 designs, published in Sustainability [17], we apply the same five-phase protocol to the same curated corpus of publications. Here we briefly summarise the shared design and data-collection phases (Phases 1–3 in Figure 1) so that the study remains self-contained, and then focus on analytical procedures (Phase 4) that are specific to empirical validation, sectoral scope, and applicability to developing-country contexts. As in [17], our analysis is deliberately confined to maturity models proposed in the peer-reviewed academic literature: we map the academic design space of Industry 4.0 and 5.0 maturity models as conceptual artefacts, rather than evaluating their diffusion or effectiveness in practice. The workflow for review design, document retrieval, and data extraction is shared across both studies, but in the present article the coding scheme and analytical focus are extended to address a different set of questions concerning empirical validation, sectoral scope, and applicability to developing-country contexts.

2.1. Document Retrieval Process (RQ1–RQ5, Review Foundation)

Five key research questions were defined (see Table 1) to guide the analysis of the maturity models. These questions examine: (1) the geographical origin of the models and the developed–emerging country bias; (2) the influence of the type of authorship (single-country, international collaboration, or global approach) on the extent to which emerging contexts are considered; (3) the impact of the development level of the model’s country of origin on its stated applicability; (4) the degree of empirical validation and sectoral coverage achieved; and (5) the knowledge gaps identified by the literature itself as barriers to the adoption of Industry 4.0/5.0.
The search terms were structured to capture relevant literature using Boolean operators:
(“Industry 4.0” OR “Fourth Industrial Revolution” OR “Industry 5.0” OR “Fifth Industrial Revolution”) AND (“maturity model” OR “adoption model” OR “framework”)
Web of Science and Scopus were selected due to their recognized coverage of high-quality scientific literature. The method followed for retrieving reference documents is illustrated in Figure 2.
An initial search returned 6909 and 7537 records in Web of Science and Scopus, respectively. Results were first restricted to the 2020–2024 period, English-language, peer-reviewed articles, conference papers, and review articles, and to the subject categories listed in Table 2. Because each database still yielded more than 4000 items, we relied on the built-in relevance ordering to obtain a manageable but information-rich subset. In Web of Science we used the default “Relevance” sort, which prioritizes records according to the occurrence and weighting of the query terms in the title, abstract, and author keywords. In Scopus we likewise kept the default relevance-based ordering, internally computed from the frequency of the search terms and their weighting by fields (title/abstract/keywords) for the query used. From each database we exported the first 250 records in the relevance-sorted list (500 records in total before de-duplication), treating this step as a transparent and reproducible heuristic: any researcher can replicate it by applying the same search string, temporal and document-type filters, and relevance ordering before exporting the top-ranked 250 records from each database.
After merging both sets, duplicates were removed based on DOI and title, yielding 444 unique documents. Throughout the screening process, records were discarded at several stages, as summarized in Figure 2. First, after applying the exclusion criteria (document type, 2020–2024 date range, subject areas, and language), 2529 Web of Science records and 2993 Scopus records were excluded. Second, at the relevance-selection stage—in which only the 250 highest-ranked records in each database were retained—an additional 4130 Web of Science and 4294 Scopus documents were discarded. Full-text screening of the 444 remaining items led to the exclusion of 206 papers that either did not sufficiently describe an Industry 4.0/5.0 maturity model or were not available in full text, leaving 238 documents for detailed assessment. Finally, during the refinement of the bibliography and data extraction, 163 documents were excluded because they did not provide information that could be operationalized for the present analysis, yielding a final corpus of 75 primary articles, each defining at least one maturity model.
Because the relevance functions of commercial databases are proprietary and may under-represent certain topics (for example, niche sectors or alternative terminologies), we took three steps to reduce potential selection bias. First, we combined relevance-ranked results from two independent databases (Web of Science and Scopus), whose coverage and ranking algorithms differ. Second, we applied broad subject-area filters (engineering, computer science, management, and related domains; see Table 2) instead of very narrow subfields, so as not to exclude a priori peripheral but substantively relevant contributions. Third, during full-text screening we complemented database searches with backward reference chasing: any maturity model cited as a primary instrument in an included article but missing from the initial set of 444 records was checked manually and, if it met the 2020–2024 inclusion criteria, added to the corpus. We also verified that Industry 4.0 maturity models considered canonical in prior reviews were retrieved by this strategy. These steps do not completely eliminate the risk of relevance-based bias but make the resulting set of 75 models more robust and transparent.

2.2. Sensitivity Analysis for Selection Bias

To assess the potential selection bias introduced by relevance-based truncation of database search results, a sensitivity analysis was conducted using an additional sample of maturity models drawn from ranks 251–500 of the original Web of Science and Scopus result lists.
Records ranked 251–500 were exported from Web of Science ( n = 250 ) and Scopus ( n = 249 ), yielding a combined universe of 499 unique references after removal of one duplicate. From this universe, a high-sensitivity candidate list was generated through automated screening of titles, abstracts, and keywords for maturity- or readiness-model terminology (e.g., “maturity model”, “readiness model”, “capability maturity”, “maturity assessment”). This step was intentionally designed to maximize recall and resulted in 122 candidate records, accepting the presence of false positives.
Random selection was then implemented in Python using uniform sampling without replacement (random.sample), with a fixed random seed (seed = 20,260,109) to ensure full reproducibility. An initial random draw of 20 candidate records was manually screened against the eligibility criterion (i.e., the paper proposes, adapts, or applies a maturity or readiness model). Ineligible records were discarded and iteratively replaced through additional random draws from the remaining candidate pool until 20 eligible studies were identified.
In total, 60 candidate records were manually screened to obtain 20 eligible maturity-model articles, resulting in 40 exclusions. The same coding scheme, variable definitions, and boundary rules used for the main corpus were applied without modification to this sensitivity sample. Findings from this analysis are used exclusively as a robustness check and are not pooled with the primary dataset of 75 models. Full procedural traceability of the sampling process, including universe definition, candidate preselection, randomization settings, and exclusion counts, is documented in Appendix A.

2.3. Operationalization of Variables and Categories

After identifying the maturity models under study, a set of variables was defined whose analysis supports each of the formulated research questions.

2.3.1. Type of Authorship

To analyze how the model’s origin influences whether authors discuss its applicability in developing countries, the variable Type of Authorship was defined. For the purposes of this study, it corresponds to the classification of the authors’ institutional affiliations.
  • Single-country: all authors are affiliated with institutions from the same country.
  • Collaboration (≥2 countries): affiliations from two or more countries are present.
  • Global: the model is declared as having no national affiliation or belonging to a supranational consortium.
Boundary rules: co-leadership between authors from different countries ⇒ Collaboration; declaration of global scope or supranational consortium ⇒ Global.

2.3.2. Development Level of the Country of Origin

To assess how the development level of the model’s country of origin influences the likelihood of explicitly considering developing countries, the variable Development Level of the Country of Origin was defined, based on the World Bank classification.
  • Developed: country classified as developed.
  • Developing: country not classified as developed.
  • Mixed: co-leadership involving countries with different development levels.
  • Global: no national affiliation (consortium/supranational).
Boundary rule: co-leadership between countries with different levels ⇒ Mixed.

2.3.3. Validation Method

The variable Validation Method was defined and analyzed to determine the extent to which existing models have been empirically validated. This variable indicates the type of empirical verification used to test the model:
  • No validation: no empirical verification performed.
  • Simulation/Demonstration: validation through simulation or demonstration without observed data.
  • Survey: evidence obtained through questionnaires.
  • Single case: validation conducted with a single case/company.
  • Multiple cases: validation conducted with more than one case/company.
Boundary rule: when multiple methods coexist, report the most empirical one (Multiple cases > Single case > Survey > Simulation > No validation).

2.3.4. Sectorization

The Sectorization variable was defined to identify how maturity models have been developed and validated across different sectors. It indicates whether the model specifies the sector(s) where it applies:
  • Specific: a single sector is explicitly declared (e.g., manufacturing).
  • Multisector: two or more sectors are explicitly mentioned.
  • Undefined: no sectoral boundaries are specified.
Boundary rule: if transversal activities (e.g., logistics, maintenance, quality) are described without a defined sector, classify as Undefined.

2.3.5. Worked Examples for Borderline Sector Classification

To ensure transparency and reproducibility in sector coding, explicit examples are provided for borderline cases frequently encountered in the corpus, such as automotive manufacturing, textiles/apparel, and cross-cutting logistics.
Automotive manufacturing. When a model explicitly targets the automotive industry (e.g., car assembly, OEM-specific production systems, or validation within automotive plants), it is coded as Specific, with the sector label Automotive manufacturing. For inferential analyses requiring normalized categories, this label is mapped to the broader Manufacturing category. By contrast, if automotive is mentioned only as one illustrative example within a general manufacturing framework, no automotive-specific sector code is assigned.
Textiles and apparel. Models are coded as Specific with the sector label Textiles/Apparel when their structure, indicators, or empirical validation are explicitly grounded in textile or garment production contexts. If textiles or apparel appear only as one example among several manufacturing sectors, the model is coded according to its dominant declared scope rather than as textile-specific.
Logistics as a cross-cutting function. When logistics (e.g., warehousing, transport, intralogistics) constitutes the primary object of assessment and is treated independently of the firm’s industrial sector, the model is coded as Specific with the sector label Logistics/Supply Chain. If logistics is addressed only as a transversal activity embedded within a broader manufacturing or supply-chain model, the sector classification follows the dominant production context rather than logistics as a standalone sector.
These examples illustrate how sector classification prioritized the declared scope, primary object of analysis, and validation context of each model, rather than incidental mentions of sectors or functions. For full transparency, the normalization rules and the mapping table used to harmonize synonymous sector labels (e.g., “Apparel”, “Clothing”, “Textile & Confections”) into analytic categories are provided in Appendix B.

2.3.6. Identified Gaps/Limitations

To identify the research gaps recognized or highlighted by the authors of the maturity models, the variable Identified Gaps/Limitations was defined. Boundary rule: if a record fits multiple themes, assign all relevant tags and document the decision.

2.4. Data Extraction, Normalization, and Coding (RQ1–RQ5)

The data extraction and coding process (see Figure 3) began with document collection, incorporating full-text articles (PDFs) and associated metadata into a document management system. This was followed by GPT-assisted pre-extraction using controlled, stepwise prompts designed to identify key elements, including title, year, author affiliations, sectoral focus, applicability mentions, validation approaches, and reported limitations or research gaps. The GPT agent was used strictly as a primary extractor to support scalability and consistency; it did not perform final coding or interpretation.
To ensure transparency and reproducibility of the AI-assisted extraction process used for RQ1–RQ4, the full interaction protocol and step-by-step instructions provided to the GPT-based agent are documented in Appendix C. This protocol specifies the rules, scope, and sequential tasks guiding the identification of maturity models, their classification (Industry 4.0/5.0), sectoral scope, validation evidence, applicability to developing-country contexts, and inclusion or exclusion decisions. The agent was used exclusively as a primary extractor to support scalability and consistency, while all final coding, interpretation, and analytical decisions were performed by the authors through systematic human verification.
All automatically extracted outputs were subsequently subjected to complete human verification. Reviewers checked the accuracy, conceptual validity, and traceability of each extracted item against the source article. For research gaps (RQ5), GPT-assisted pre-extraction was conducted using a dedicated prompt specifically designed to identify explicit statements of prior gaps in the literature, limitations of the current study, and future research opportunities. The full prompt and detailed extraction instructions are documented in Appendix D. During verification, reviewers explicitly distinguished between support at the level of the cited excerpt (local textual support) and support at the level of the full document. In a small number of cases, gaps were conceptually present elsewhere in the article but not fully supported by the specific excerpt initially selected by the agent; these instances were treated as local citation misalignments rather than false-positive identifications and were corrected during verification.
Institution names and country affiliations were then standardized, after which the development level of each model was assigned. This variable was operationalized using the World Bank country income classification applied to the affiliation country of the lead author. In addition, a Global category was used for models that explicitly self-identify as global in scope or are developed by multinational consortia without a clearly dominant national anchoring. This approach provides a transparent and reproducible proxy for comparative analysis, but it does not capture within-country heterogeneity, transnational research trajectories, diaspora effects, or potential decoupling between the institutional origin of a model and the contexts in which it is validated or applied. Accordingly, development level is interpreted in this study as an indicator of institutional origin rather than as a comprehensive descriptor of contextual grounding.
Using the cleaned and normalized dataset, the core analytical variables (type_authorship, development_level, applicability_mention, validation_method, and sectorization) were coded according to predefined rules, with explicit boundary conditions applied where necessary to ensure consistency and reproducibility. For RQ5, thematic analysis of reported research gaps was conducted through inductive coding, iterative grouping, and triangulation, resulting in the variable gap_theme. To assess the robustness of the AI-assisted extraction and standardization process, a stratified audit covering 10% of the extracted gaps was conducted, balanced across gap types. The audit evaluated excerpt-level textual support, fidelity of standardization, and correctness of gap-type classification, and informed minor corrections to the dataset.
The first three phases of the workflow (review design, document retrieval, and data extraction/coding) are shared with our previous metatypology study [17]. However, the extraction template and coding scheme were extended in the present study to capture additional variables related to authorship configuration, development level, empirical validation, sectoral focus, and research gaps, in direct alignment with the research questions addressed here.
All extracted data used to generate the figures and results are available in an open repository (version v1; DOI: 10.5281/zenodo.17454112) under the title “Industrial Digital Maturity (I4→I5)—75 Models, 2020–2024”. The repository includes the full dataset as well as a BibTeX file listing all primary sources used in the analysis and can be accessed at https://zenodo.org/records/17454113, accessed on 27 October 2025.

2.5. Analytical Procedures

This section describes how the data were processed and analyzed for each research question (RQ). Building on the five-phase workflow shown in Figure 1, the analytical procedures are grouped into four main families: (i) descriptive statistics, (ii) inferential statistics, (iii) text mining and qualitative analysis, and (iv) visualization and bibliometrics. A separate subsection summarizes the software environment used for data processing and analysis.

2.5.1. Descriptive Statistics (RQ1–RQ5)

Descriptive statistics were used to summarize and present the distributions of all variables defined in Section 2.2 (Operationalization of Variables and Categories). For each variable, absolute (N) and relative (%) frequencies were computed, as well as cross-tabulations where relevant. These descriptive results provide the empirical basis for the figures and tables reported in Section 3 and serve as inputs for subsequent inferential analyses. For RQ1, the distributions by country of origin and authorship type informed the geographic and collaboration visualizations. For RQ2 and RQ3, contingency tables were constructed for Authorship × Applicability Mention and Development Level × Applicability Mention. For RQ4, validation method and sectorization were summarized both individually and in combination. For RQ5, frequencies and percentages of gap themes derived from the text-mining pipeline were calculated.

2.5.2. Inferential Statistics (RQ2, RQ3, RQ4)

Associations between categorical variables were examined using contingency-table analyses. Pearson’s chi-square tests of independence ( χ 2 ) with a significance level of α = 0.05 were considered when expected cell frequencies were adequate for asymptotic inference. For each table, observed and expected frequencies, χ 2 values, degrees of freedom ( d f = ( r 1 ) ( c 1 ) ), and p-values were computed, and standardized residuals were inspected when informative.
For cross-tabulations related to RQ2 and RQ3, small subgroup sizes produced sparse tables with low expected cell frequencies. In these cases, inferential analysis relied on exact tests. Specifically, Fisher’s exact test was used for 2 × 2 tables, and its generalization to r × c tables (Fisher–Freeman–Halton test) was applied using Monte Carlo approximation (20,000 simulations). All Monte Carlo simulations were conducted with a fixed random seed (seed = 20,260,113) to ensure full reproducibility of the results.
The variable Applicability to developing-country contexts was originally coded into three categories (explicit mention, implicit or partial mention, and no mention). Implicit applicability was coded when the study did not explicitly claim relevance for developing or emerging economies, but contextual features of the empirical setting strongly suggested such applicability, for example when validation was conducted in organizations located in developing regions, focused on SMEs operating under resource-constrained conditions, or embedded in infrastructural or institutional contexts typically associated with emerging economies. This full categorization was retained for descriptive reporting. For inferential analyses in both RQ2 and RQ3, the variable was operationalized as a binary indicator (any mention vs. no mention) to improve analytical tractability under sparse data conditions while preserving interpretability.
Effect sizes were reported using Cramér’s V to complement significance testing, particularly given limited statistical power in small categories. All inferential analyses were conducted in R using base functions for contingency-table testing, including Fisher’s exact test with Monte Carlo simulation. Specific results for the cross-tabulations Authorship × Applicability Mention (RQ2), Development Level × Applicability Mention (RQ3), and Sectorization × Validation Method (RQ4) are presented in Section 3.
In addition, to assess whether the distribution of models across target sectors departed from a neutral reference pattern (RQ4), a chi-square goodness-of-fit test was conducted on the frequency counts of Sector_Objetivo. As a benchmark, a uniform distribution across the observed sector categories was used (i.e., equal expected counts per sector), serving as a transparent reference for detecting over- and under-representation. Because several sectors occur with low frequencies, p-values were obtained via Monte Carlo simulation (10,000 replicates; fixed random seed = 20,260,113). To support interpretation beyond statistical significance, effect size was reported using Cohen’s w (equivalently, χ 2 / N ), and standardized residuals were inspected to identify the sector categories contributing most to deviations from the benchmark. Prior to the goodness-of-fit analysis, sector labels were harmonized into a reduced set of analytically meaningful macro-sectors to avoid artificial fragmentation due to synonymous or highly specific sector descriptions. For the sectoral analysis (RQ4), target sectors reported in the original studies were harmonized into a set of normalized sector categories prior to statistical testing. This normalization step was required to reduce terminological heterogeneity across studies and to enable meaningful goodness-of-fit testing under highly asymmetric distributions. The normalization criteria, including mapping rules and decision logic for ambiguous or compound sector labels, are fully documented in Appendix B.

2.5.3. Text Mining and Qualitative Analysis (RQ5)

For RQ5, 562 text fragments related to gaps and limitations were extracted from the 75 maturity-model articles and incorporated into the dataset as individual records. A Natural Language Processing (NLP) pipeline was then applied in four steps: (i) preprocessing, (ii) inductive coding, (iii) thematic grouping, and (iv) quantitative–qualitative triangulation. Preprocessing included tokenization, stop-word removal, lemmatization, and part-of-speech tagging, implemented with NLTK [18] and spaCy [19]. Based on the preprocessed text, inductive coding was performed manually by the authors to identify recurrent gap statements, which were then grouped into broader themes (e.g., empirical validation, sectoral coverage, SME orientation, sustainability, implementation tools). Finally, the thematic structure was triangulated with quantitative frequencies to obtain the distribution of gap types reported in Section 3.

2.5.4. Visualization and Bibliometrics (RQ1)

To represent the geographical origin of the models and the patterns of international collaboration underpinning knowledge production (RQ1), a set of visualizations and bibliometric analyses was produced. The distribution of maturity models by authorship configuration (single-country, international collaboration, and global) was first visualized using a pie chart created in Microsoft Excel (see Figure 4), based on descriptive frequency counts derived from the curated dataset. Choropleth maps of geographical origin were then generated using the Power Map add-in for Microsoft Excel [20], shading countries by the number of maturity models identified (see Figure 5). International co-authorship networks were constructed with the Bibliometrix package for R [21], based on a country–country collaboration matrix in which each row represents at least one co-authored article between two countries (e.g., “GERMANY–BRAZIL–1”). From this matrix, collaboration networks were visualized to highlight regional clusters and cross-regional ties (see Figure 6).

2.5.5. Software Environment

All data processing and statistical analyses were conducted in Python 3.11. Tabular data were managed and transformed using the pandas library [22], which was also used to construct the contingency tables and descriptive summaries reported in Section 3. Chi-square tests of independence were implemented with the chi2_contingency function from scipy.stats in SciPy [23]. For the NLP pipeline described above, NLTK [18] and spaCy [19] were used for tokenization, lemmatization, stop-word removal, and part-of-speech tagging. Bibliometric analyses (including co-authorship networks) were carried out with Bibliometrix in R [21], and choropleth maps of geographical origin were developed using the Power Map add-in for Microsoft Excel [20].

2.6. Traceability Matrix: RQs–Method–Result–Evidence

To support methodological transparency and reproducibility, Table 3 links each research question (RQ) to the corresponding variables, methods, evidence, and the main results discussed in Section 3.
Table 3. Traceability  matrix linking research questions (RQ) to variables, methods, evidence, and results.
Table 3. Traceability  matrix linking research questions (RQ) to variables, methods, evidence, and results.
RQVariable(s)Procedure (Methods)Evidence/ArtifactLinked Result
RQ1—Origin of KnowledgeV1; V2Section 2.4
Section 2.5.1
Section 2.5.4
Section 2.5.5
DS1—dataset of 75 maturity models
Figure 4—origin by authorship type
Figure 5—geographical distribution by country
Figure 6—origin and collaboration networks
R1.1 Concentration of Industry 4.0/5.0 maturity models in a small group of developed countries, especially in Europe.
R1.2 International collaboration networks dominated by institutions from developed economies, with limited South–South ties.
RQ2—Authorship and ApplicabilityV1Section 2.4
Section 2.5.1
Section 2.5.2
Section 2.5.5
DS1—dataset of 75 maturity models
Table 4—authorship × applicability to developing-country contexts
R2.1 Evidence is insufficient to detect systematic differences in applicability to developing-country contexts by authorship type.
R2.2 Small subgroup sizes limit statistical power, suggesting caution in interpreting non-significant differences.
RQ3—Development Level and ApplicabilityV2Section 2.4
Section 2.5.1
Section 2.5.2
Section 2.5.5
DS1—dataset of 75 maturity models
Table 5—development level × applicability to developing-country contexts
R3.1 An exploratory association is observed between development level of institutional origin and the likelihood of mentioning applicability to developing-country contexts.
R3.2 Models originating in developing economies show higher rates of explicit or implicit applicability mentions, within the limits imposed by sparse categories.
RQ4—Validation and SectorsV3; V4Section 2.4
Section 2.5.1
Section 2.5.2
Section 2.5.5
DS1—dataset of 75 maturity models
Table 6—validation status distribution
Table 7—sectoral focus distribution
Figure 7—target sectors (specific)
Table 8—sectoral focus × validation
Table 9—validation by primary sector
R4.1 Most models report some empirical validation, dominated by surveys and case studies of limited scale.
R4.2 Sectoral coverage is highly skewed toward manufacturing, with statistically significant under-representation of other sectors.
R4.3 Multisectoral scope is often declared but less frequently demonstrated through cross-sector validation.
RQ5—Research GapsV5Section 2.4
Section 2.5.1
Section 2.5.3
Section 2.5.5
DS1—dataset of 75 maturity models
Figure 8—distribution of gap types
R5.1 Recurrent gaps concern limited empirical validation, narrow sectoral scope, and lack of longitudinal or cross-country evidence.
R5.2 Many gaps highlight insufficient attention to developing-country contexts, SMEs, and resource-constrained environments.
R5.3 Explicit Industry 5.0 dimensions are infrequently operationalized and empirically validated.
R5.4 Open tools, datasets, and implementation resources remain rare, limiting replication and adoption.
Table 4. Cross-distribution of Authorship Type × Applicability to Developing-Country Contexts (RQ2).
Table 4. Cross-distribution of Authorship Type × Applicability to Developing-Country Contexts (RQ2).
OriginExplicit MentionImplicit or Partial MentionNo Mention at AllTotal
Collaborations1067
Global10910
Single-country564758
Total766275
Table 5. Cross-distribution of Development Level × Applicability to Developing-Country Contexts (RQ3).
Table 5. Cross-distribution of Development Level × Applicability to Developing-Country Contexts (RQ3).
Development LevelExplicit MentionImplicit or Partial MentionNo Mention at AllTotal
Developed013536
Developing651627
Global10910
Mixed0022
Total766275
Table 6. Distribution of maturity models by validation status ( n = 75 ).
Table 6. Distribution of maturity models by validation status ( n = 75 ).
Validation StatusModels% of Total
Not validated (no evidence)1114.7%
Simulation/Demo (laboratory environment)45.3%
Survey (questionnaire data)1824.0%
Single case study (one case study)1722.7%
Multiple case studies (≥2 case studies)2533.3%
Total75100%
Table 7. Distribution of maturity models by level of sectoral focus ( n = 75 ).
Table 7. Distribution of maturity models by level of sectoral focus ( n = 75 ).
Level of Sectoral FocusModels% of Total
Specific (defined sector)5776.0%
Multisector (multiple sectors)1418.7%
Undefined (general sector)45.3%
Total75100%
Table 8. Relationship between sectoral scope and validation status. Number of maturity models by combination of sectoral focus and validation type ( n = 75 ).
Table 8. Relationship between sectoral scope and validation status. Number of maturity models by combination of sectoral focus and validation type ( n = 75 ).
Sectoral Focus∖ValidationNot ValidatedSimulation/DemoSurveySingle CaseMultiple Cases
Specific73131717
Multisector11408
Undefined30100
Table 9. Validation methods by primary sector of the models. Distribution of validation types used across the five most frequent target sectors.
Table 9. Validation methods by primary sector of the models. Distribution of validation types used across the five most frequent target sectors.
Primary SectorNot ValidatedSimulation/DemoSurveySingle CaseMultiple Cases
Manufacturing329812
Manufacturing (SMEs)10221
Automotive00003
Multisector (NA)01100
Textile00101

2.7. Use of AI Tools

As part of the data extraction and synthesis process, we used controlled, task-specific prompts with a large language model to assist in the systematic extraction of research gaps. This AI assistance was strictly limited to the **initial identification of candidate text segments** according to predefined criteria. All extracted content was subsequently subjected to **complete human verification** and manual coding by the authors; no part of the manuscript was generated by AI without human oversight.
Detailed procedures for AI-assisted extraction, including verification and auditing protocols, are documented in Section 2.4.

3. Results

This section presents the empirical results structured around the research questions. Geographical patterns of model origin and collaboration are reported in Figure 4, Figure 5 and Figure 6, and their relationship with explicit applicability to developing-country contexts is examined in Table 4 (RQ2) and Table 5 (RQ3). Empirical validation strategies and sectoral concentration are analysed in Table 6, Table 7 and Table 8, with subsector-level detail provided in Figure 7 (RQ4). Finally, the synthesis of research gaps and limitations identified by primary studies is presented in Section 3.6, including the distribution of gap types (Figure 8) and their dominant thematic patterns (RQ5).

3.1. Applicability and Geographical Origin

A relevant analysis derived from the maturity model dataset concerns their applicability in the context of developing countries and how the model’s origin influences such applicability. Based on the available data, it is possible to identify key patterns and trends that help to understand both the opportunities and challenges related to the adoption of these models across different economic and geographic contexts.
To this end, the distribution of the 75 maturity models published between 2020 and 2024 was examined, classified into three categories according to the country of origin and type of collaboration: Single-Country, Collaborative, and Global. This classification followed the criteria defined in the methodology:
  • Single-Country: models developed and published entirely from one country;
  • Collaborative: those resulting from joint efforts between two or more nations (e.g., identified as “Collaboration (Italy–Poland)”); and
  • Global: models that are not explicitly linked to any specific national context in their documentation, suggesting a universal scope or orientation.
The frequency analysis (see Figure 4) revealed that 58 models, representing approximately 77.3% of the total, fall into the Single-Country category; 7 models (9.3%) correspond to the Collaborative category; and 10 models (13.3%) were classified as Global.
These results show that the majority of maturity models are created within a national framework (Single-Country), followed by a smaller share developed through international collaborations and globally oriented approaches. This distribution provides an initial indication of prevailing trends in knowledge production within this field, suggesting that the literature base is largely rooted in local initiatives. It also raises important questions about how the origin, whether individual, collaborative, or global, might influence the applicability of these models in developing countries contexts.
On the other hand, the geographical distribution is highly uneven (see Figure 5). Nearly half (48%) of the models originate from developed countries, and just over one-third (36%) from developing economies; the remainder come from models declared as “global” with no specific country of reference or from mixed collaborations with limited scope. Europe overwhelmingly dominates model production, accounting for around 60% of the total, led by Germany and supported by strong contributions from other countries such as Portugal, Italy, and Spain. Asia ranks second with approximately 27% of the models, particularly from Turkey, Indonesia, and other industrialized economies in the region, while Latin America accounts for only 13%, almost exclusively due to Brazil and Mexico. Africa, North America (beyond a single U.S. study), and Oceania appear only marginally, each represented by one model.
These distributions indicate a strong concentration of model production in developed economies, with relatively fewer contributions from emerging regions.
In addition to analyzing the individual distribution of models, the network of international collaborations was also examined (see Figure 6). For this analysis, a matrix was created in which each row represents a collaboration between two countries (e.g., “GERMANY–BRAZIL–1” indicates that at least one article was co-authored by authors from Germany and Brazil).
A large share of collaborations occurred among European countries, including partnerships between Finland and Norway, Italy and Poland, Italy and Spain, and Portugal with the Netherlands and Romania, as well as various interactions among the United Kingdom, Sweden, Italy, and Spain. This highlights the high density and cohesion of the European bloc in knowledge production. Connections between Europe and other regions were also identified. For example, collaborations between Germany and Brazil (Europe–Latin America), Germany and Indonesia (Europe–Asia), as well as Mexico with Spain (Latin America–Europe) and Spain with Colombia (Europe–Latin America) indicate that Europe stands out not only for its internal output but also for its capacity for international cooperation. Similarly, collaborations between Austria and Thailand (Europe–Asia), and between Mexico and Australia (Latin America–Oceania), demonstrate the global dimension of these research networks. In Latin America, the collaboration between Brazil and Mexico is particularly noteworthy, with a frequency of 2, suggesting a strong connection within the region.

3.2. Origin vs. Applicability to Developing-Country Contexts (RQ2)

Descriptive row-wise analysis indicates that explicit references to applicability to developing-country contexts are uncommon across all origin categories. Within the Collaborations category, 14.3% of models include an explicit mention, while 85.7% do not address applicability. Among Global models, 10.0% include an explicit mention and 90.0% provide no reference. In the Single-country category, 8.6% of models include an explicit mention and 10.3% provide an implicit or partial reference, whereas 81.0% do not address applicability in developing-country contexts.
For inferential analysis, explicit and implicit mentions were collapsed into a single binary category (any mention vs. no mention), as described in Section 2. Given the small size of several subgroups and the presence of low expected cell frequencies, asymptotic chi-square tests were not applied. Instead, the association between model origin and applicability to developing-country contexts was evaluated using a Fisher–Freeman–Halton exact test with Monte Carlo approximation (10,000 simulations).
The exact test indicates a statistically significant association between model origin and applicability mention (Monte Carlo p < 0.01 ). The corresponding effect size, measured using Cramér’s V, was approximately 0.47, indicating a moderate association in the analyzed dataset. Given the exploratory nature of the analysis and the limited size of some origin categories, these results are reported as evidence of association within the studied corpus and do not imply causal relationships. However, this association should be interpreted with caution. Several origin categories, particularly Collaborative and Global models, are represented by small cell counts (see Table 4), which limits the stability of estimated associations. Although exact tests and effect sizes were employed to mitigate violations of asymptotic assumptions, the results should be understood as indicative patterns within the analyzed corpus rather than as definitive or generalizable evidence.

3.3. Development Level vs. Applicability to Developing-Country Contexts (RQ3)

To analyze the possible influence of the development level (classified as Developed, Developing, Global, and Mixed) on the mention of applicability in developing-country contexts, the following 4 × 3 contingency table ( n = 75 ) was created. As in the previous section, inferential statistical methods were applied (see Materials and Methods). Table 5 shows the cross-distribution.
The results indicate that, within the Developed group, the vast majority of models (35 out of 36, 97.2%) make no reference to applicability in developing-country environments, and only one model (2.8%) includes an implicit mention. In contrast, within the Developing group, there is greater sensitivity to the topic: 22.2% of the models (6 out of 27) make an explicit reference and 18.5% (5 out of 27) make an implicit one, totaling 40.7% of models that address applicability in such contexts. Conversely, in the Global and Mixed categories, the trend mirrors that of the Developed group, with omission predominating (90% and 100%, respectively).
To test whether the observed differences were statistically significant, a chi-square test of independence was performed. With 6 degrees of freedom and a 5% significance level, the critical chi-square value is approximately 12.59. Since the obtained test statistic ( χ 2 = 17.13 ) exceeds this threshold, the result indicates p < 0.05 .
The analysis indicates a statistically detectable association between the development level of the country (or countries) of origin and the way applicability in developing-country contexts is mentioned. This result should be interpreted as exploratory, given the presence of sparsely populated categories—most notably the Mixed and Global groups (Table 5)—which constrain statistical power and the stability of effect estimates. Accordingly, the observed association is indicative of patterns within the reviewed literature rather than conclusive evidence of systematic differences. Specifically, models originating from Developing environments are more likely to include explicit or implicit mentions, in contrast to those developed in Developed countries, where such mentions are almost nonexistent.
At the same time, the limited participation of entire regions (e.g., Sub-Saharan Africa, Southeast Asia beyond a few economies, or North America beyond the United States) indicates geographical gaps where these models have had little penetration, or at least limited visibility, in the analyzed literature. This geographic pattern likely influences the applicability of maturity models in developing contexts. Since most models were created with industrialized environments in mind, they may incorporate implicit assumptions (advanced technological infrastructure, mature 4.0 ecosystems, high baseline digitalization) that do not hold true in many emerging economies. The fact that 83% of the studies do not discuss applicability in developing countries reinforces this interpretation: it implies that authors typically did not explicitly consider contextual differences in industries outside the developed world when formulating their models. Consequently, such models may not fit well in settings where manufacturing SMEs face capital, talent, or connectivity constraints that differ substantially from those in European or North American contexts.
Conversely, the few studies that do address applicability in emerging economies tend to come from authors based in those same contexts or from global analyses concerned with transferring Industry 4.0 practices to less developed regions. This points to an emerging awareness among part of the academic community regarding the need to adapt maturity models to diverse realities. The explicit mentions identified (in about 9% of the works) reflect deliberate efforts to consider local factors, such as technological gaps, government incentive policies, or cultural and organizational differences, that could affect the attainable 4.0 maturity level in developing countries. However, since these cases remain scarce, it is plausible that many model proposals still suffer from an origin bias, being calibrated primarily according to the experiences of advanced economies and thus limiting their universal validity.
Overall, the geographical analysis of Industry 4.0/5.0 maturity models reveals an uneven distribution with important implications for their global adaptability. The concentration of contributions in the developed world suggests that the characteristics and criteria of these models predominantly reflect the conditions of highly digitalized industries and favorable socioeconomic environments. This geographical focus may limit the capacity of these models to effectively adapt to developing contexts, where companies are often at very different stages of digitalization and face distinct structural challenges. In other words, the relevance and utility of many existing maturity models in developing countries remain questionable unless differences in resources, scale, and priorities are adequately considered.
Nonetheless, the landscape also reveals opportunities. The presence of several proposals originating from emerging economies (and some global initiatives) demonstrates that it is possible to reconceptualize maturity models by incorporating the perspective of developing countries. Although still in the minority, these emerging models could serve as starting points to enrich or reorient existing methodologies, emphasizing flexibility and contextualization. Ultimately, the findings suggest a pressing need to update and validate Industry 4.0/5.0 maturity models across a broader range of geographical environments. Only through greater geographical inclusiveness, fostering more balanced international collaborations and field studies in underrepresented regions, can the maturity assessment tools become truly universal and applicable across the full spectrum of industrial realities, from the most advanced economies to developing nations.

3.4. Validation Status and Sectoral Analysis

Among the main challenges in the development of maturity models are those related to empirical validation and clarity regarding their intended sectors of application. Through a systematic analysis of the selected maturity models, this section explores how the models have been validated, ranging from single case studies to surveys and simulations, and the diversity of industrial sectors they address. By identifying validation patterns, levels of sectoral focus, and their evolution over time, this section provides a critical perspective on the rigor and practical applicability of these models, as well as their sectoral orientation. This approach not only reveals trends and gaps in the literature but also lays the groundwork for the design of more robust and contextually relevant models, especially for developing-country settings.

3.4.1. Validation Status of Maturity Models (RQ4)

The validation quality of each model was analyzed based on the variable Validation Method. Table 6 summarizes the distribution of the 75 maturity models across the different categories of this variable.
It can be observed that nearly two-thirds of the models (55 out of 75) include validations through case studies: 17 models (23%) performed a single, in-depth case study (e.g., applied in one company), while 25 models (33%) conducted multiple case studies across different organizations or contexts. Another 24% (18 models) used survey-based validation, generally to measure perceived maturity levels across multiple firms simultaneously. In contrast, only 4 models (5%) were validated through simulation or demonstration in controlled environments (laboratory or pilot), making this the least frequent category. Finally, 11 models (15%) were published without any empirical validation, relying instead on theoretical foundations or the authors’ expert judgment.
Regarding the intensity of validation, notable differences exist across categories. Models with multiple case studies included on average ~6 real cases (median = 3 ), ranging from a minimum of 2 to a maximum of 32 in one exceptional model. By definition, single case study models each employed one case. For survey-based studies, the validation scale tends to be broader: excluding two cases where the sample size was not reported (recorded as 0), the surveys collected between 16 and 323 responses, with a median of approximately 70 participants. This indicates that many studies surveyed several dozen companies, and one case reached over 300 respondents (MM.12). In contrast, simulation/demo validations are typically not counted as empirical “cases”: three out of four simulation models reported no cases ( N _ Cases = 0 ), and only one model described three simulated scenarios as its form of validation. Overall, the majority of maturity models seek some level of empirical support, either through case studies or surveys, although the depth and robustness of the evidence vary widely, ranging from conceptual simulations or perception-based surveys to practical implementations in dozens of factories.

3.4.2. Level of Sectoral Focus of Maturity Models (RQ4)

To describe how the models are distributed according to their sectoral focus (Specific, Multisector, Undefined), Table 7 summarizes the distribution by category.
As shown in Table 7, an overwhelming majority of the analyzed maturity models, 57 out of 75 (76.0%), were developed with a specific sector in mind. This indicates that Industry 4.0/5.0 maturity research is largely dominated by sector-tailored approaches. In contrast, only 14 models (18.7%) claim a multisectoral scope, and a very small subset of 4 models (5.3%) do not define any particular sector, presenting themselves as general-purpose or horizontal frameworks.
When examining the target sectors of the models classified as Specific, manufacturing and industrial production contexts clearly dominate the design space. Manufacturing-oriented models represent the largest share within this category, followed at a considerable distance by models targeting manufacturing SMEs and, to a lesser extent, the automotive sector. Beyond these cases, other sectors appear only sporadically, typically represented by one model each. Overall, sectoral coverage can therefore be characterized as broad but shallow: a relatively wide variety of sectors is mentioned across the corpus, but most are supported by very limited empirical and conceptual depth.
To formally assess whether this observed concentration reflects a statistically meaningful pattern rather than random variation, target sectors were harmonized into normalized sector categories following explicit criteria (documented in Appendix B) and tested against a uniform benchmark using a goodness-of-fit test. Given the strong asymmetry in the distribution and the presence of small cell counts, inference relied on a Monte Carlo approximation. The results indicate that the observed sectoral distribution differs significantly from a uniform distribution ( p < 0.001 ), with a very large effect size (Cohen’s w = 1.25 ), confirming a highly uneven representation of sectors in the analyzed maturity models.
Among the 14 multisector models, nearly all declare applicability to more than one sector. However, the degree of specification varies substantially. Some models explicitly list multiple distinct sectors, while others label themselves as multisectoral without detailing how sectoral differences are accounted for in the model structure or assessment criteria. Only a small number of multisector models clearly articulate differentiated sectoral scopes, which constrains meaningful comparative analysis between sector-specific and genuinely cross-sector frameworks.
Finally, the four models classified as having an undefined sectoral focus correspond to general-purpose approaches, often emphasizing horizontal dimensions such as organizational culture, workforce capabilities, or human–technology interaction rather than industry-specific characteristics. These models frequently align with Industry 5.0 perspectives that prioritize holistic and human-centric considerations over sectoral specialization.
Figure 7 provides a descriptive subsector-level view of the most frequently targeted sectors among models classified as Specific. This visualization is intended to illustrate patterns of concentration within manufacturing and related domains. Statistical inference regarding sectoral concentration, however, is conducted on normalized sector categories, as reported above, to ensure analytical robustness under highly asymmetric distributions.

3.5. Relationship Between Sectoral Focus and Validation (RQ4)

Given that differences exist both in sectoral scope and in validation methods, it is pertinent to analyze how these two dimensions intersect. The following contingency table (Table 8) cross-tabulates the level of sectoral focus (rows) with the validation status (columns) of each model. Table 9 then provides a more detailed breakdown, showing how the five most frequent sectors employed each type of validation.
Several trends can be drawn from this table. First, Undefined models (general-purpose frameworks) show the highest proportion of non-validation: 3 out of 4 (75%) were not validated. This suggests that the more theoretical or universal frameworks, often associated with Industry 5.0 visions, remain largely at the conceptual level.
In contrast, Multisector models are usually validated: 13 out of 14 (93%) include some type of empirical evidence. Notably, none of the multisector models relied on a single case study (0 single cases), likely because a model designed for several industries cannot be sufficiently demonstrated through a single case. Instead, multisector validations are mainly divided between surveys (4 models) and, especially, multiple case studies (8 models). This indicates a preference for gathering broad evidence, either through cross-sectoral surveys or by applying the model in multiple organizations from different industries, to support multisector applicability. Only one multisector model lacks validation (a 2023 case, possibly an emerging conceptual framework).
Specific models, on the other hand, exhibit a more balanced distribution of validation methods. Being the largest category (57 models), they also dominate in absolute numbers across all validation types. However, it is noteworthy that even among the specific models, 7 (12%) lack validation, reflecting that not all sectoral proposals reached the testing phase. Within partial validations, 3 specific models employed simulations. Surveys were used in 13 specific models, typically focusing on manufacturing SMEs, where it is feasible to survey a large number of firms. Case studies are almost evenly split: 17 specific models used single case studies (e.g., one automotive company), while another 17 employed multiple case studies within the same sector (e.g., several manufacturing plants). This indicates that both validation strategies are common in sector-specific models, some emphasizing in-depth exploration of a representative organization, and others seeking generalization within a sector through multiple implementations.
The table above provides a deeper sector-level perspective. Manufacturing (general) accounts for 34 models employing a variety of validation methods, with multiple case studies (12 models) and surveys (9 models) standing out as the most frequent, reflecting the abundance of empirical research in manufacturing environments. Only 3 manufacturing models lack validation, a relatively low percentage given the sector’s prominence.
For Manufacturing SMEs, out of six total models, two were either unvalidated or minimally validated, while two conducted survey-based studies (e.g., national surveys of industrial SMEs) and three performed one or more case studies.
Interestingly, Automotive models (3 total) were all validated through multiple case studies, meaning each automotive model was tested in more than one setting (for example, across several plants or divisions within an automotive company).
In contrast, in the Textile/Apparel sector, the two available models used distinct validation strategies: one employed a survey (a perception study across textile firms) and the other multiple case studies (applied across several textile factories), showing methodological diversification even with limited samples.
The Multisector (NA) row in the table corresponds to models whose “primary sector” was undefined (marked as Multisector in the dataset). Among these two models, one used a survey and the other a simulation, aligning with the previously noted tendency of broad-scope models to rely on survey-based or pilot-simulation evidence rather than single-case demonstrations.

Illustrative Cases of Exceptional Validations and Atypical Sectors in the Sample

One model (see Appendix E, MM.12) took quantitative validation to the highest level by conducting an extensive survey among hundreds of manufacturing companies to assess their maturity level. With 323 responses collected, this study presents the largest N Cases in the entire sample, demonstrating the feasibility of obtaining maturity metrics at a national scale. This approach stands in stark contrast to the majority of models, which rarely exceed a few dozen observations.
A recent multi-sector model (Appendix B, MM.34) employed an unusual hybrid validation strategy, combining real-world testing with simulations. Specifically, components of the model were tested across different sectors (e.g., aerospace component manufacturing and engine maintenance), and its use was also simulated in hypothetical scenarios (e.g., urban flooding in a smart city). This mixed validation demonstrates an exceptional level of verification, covering both real and virtual environments to assess the model’s robustness.
In 2024, a model emerged focusing on a rarely covered sector in the literature: the food industry, specifically seafood processing. This model (Appendix E, MM.37) adapted the maturity framework to an Industry 5.0 context for small-scale fishing enterprises, an area previously unexplored. Its validation, carried out through a survey (42 responses), represents a pioneering application within the food value chain.
The existence of only one model in this domain highlights current sectoral gaps and the pressing need to expand maturity model research into industries such as food, agriculture, healthcare, and others that remain largely overlooked. Short summaries of these exemplary models are also provided in Appendix F.

3.6. Gaps and Limitations (RQ5)

To address RQ5, evidence on research gaps reported across 75 Industry 4.0 and Industry 5.0 maturity models was synthesized, based on 562 gap records extracted from the analyzed articles. The analytical framework integrates two complementary perspectives: (i) a quantitative analysis describing the distribution of gaps by type (prior gap, intrinsic limitation, future opportunity), and (ii) a qualitative thematic analysis identifying recurrent content patterns (e.g., empirical validation, sustainability, SME orientation, assessment tools). This section reports the relative frequencies of gap types and the dominant thematic patterns observed over the 2020–2024 period, highlighting persistent structural voids and emerging research directions in maturity-model development.
Figure 8 summarizes the distribution of research gaps identified across the analyzed maturity models, distinguishing between pre-existing gaps inherited from prior literature, intrinsic limitations acknowledged by the authors, and future research opportunities proposed by the studies.
Gap extraction was supported by a GPT-based agent using controlled, task-specific prompts, strictly as a primary extractor to enhance scalability and consistency. All extracted gap records were subsequently subjected to complete human verification prior to analysis. To explicitly assess the reliability of the AI-assisted extraction and standardization process, a stratified audit covering 10% of the gap records ( n = 56 ) was conducted, balanced across the three gap types.
The audit evaluated four aspects: (i) textual support of the extracted gap at the level of the cited excerpt (hallucination control), (ii) fidelity of gap standardization, (iii) correctness of gap-type classification, and (iv) the corrective action required after review. The results indicate high semantic and procedural reliability. Of the audited gaps, 52 out of 56 (92.9%) were directly supported by the cited text. The remaining cases reflected local citation misalignments rather than false-positive gap identification: the gap was present elsewhere in the article but not optimally anchored in the initially selected excerpt. No conceptually invalid gaps were identified.
Standardization fidelity was assessed as accurate in 51 cases (approximately 91%), with one instance of over-generalization detected and corrected during verification. Similarly, gap-type classification (prior gap, intrinsic limitation, future opportunity) was correct in 51 audited cases (approximately 91%), with a single misclassification corrected. Overall, 49 gap records required no modification, while 7 cases (12.5%) were edited, reformulated, reclassified, or discarded. The discard rate was low (3 cases; 5.4%), indicating effective control without systematic inflation or suppression of gaps. These results demonstrate that the AI-assisted pipeline, when combined with systematic human verification, provides a transparent and reliable basis for gap synthesis.
Among the 562 gap records, prior gaps represent the most frequent category, accounting for 42.6% (239 cases). These refer to deficiencies inherited from earlier studies or existing models, such as missing dimensions or analytical perspectives not addressed in previous frameworks. Future opportunities constitute 31.0% of the records (174 cases) and focus on proposing new research directions or extensions, including the integration of emerging technologies or underexplored organizational perspectives. Self-acknowledged limitations account for the remaining 26.4% (148 cases) and capture constraints explicitly recognized by the authors, such as limited empirical validation or restricted sectoral scope.
This distribution suggests that recent literature places slightly greater emphasis on identifying inherited gaps and future research avenues than on critically elaborating the limitations of individual studies. Nevertheless, all three categories are substantively represented, indicating a balanced pattern of retrospective critique, self-reflection, and forward-looking orientation.

Common Thematic Patterns and Persistent Gaps

The qualitative thematic analysis of the standardized gap entries revealed several recurring patterns across the 2020–2024 period. The most prominent theme concerns insufficient empirical validation, frequently manifested through references to limited case studies, absence of large-scale testing, or lack of longitudinal evidence. This issue appears consistently throughout the period, indicating a persistent structural weakness in maturity-model research.
A second dominant theme relates to the need for new or expanded maturity models that incorporate dimensions previously underrepresented in the literature. These include sustainability, human-centric and social factors, and the integration of advanced digital technologies such as Artificial Intelligence, Internet of Things, and data analytics. Such gaps are particularly prominent in years with higher publication volume.
Sectoral limitations also emerge as a recurrent concern. Many gap statements highlight the narrow focus of existing models on manufacturing, with limited applicability to sectors such as logistics, healthcare, services, or agri-industry. Closely related is the frequent lack of explicit orientation toward small and medium-sized enterprises (SMEs), which are often characterized by resource constraints and informal organizational structures.
Additional themes include weak integration of sustainability and social responsibility criteria, limited standardization across maturity models, and the absence of practical implementation support such as assessment tools, digital platforms, or structured roadmaps. Finally, a smaller number of gap statements point to underexplored issues such as cybersecurity and the lack of dynamic or longitudinal perspectives capable of capturing maturity evolution over time.

4. Discussion

The results obtained across the five research questions offer a multi-layered picture of how Industry 4.0/5.0 maturity models are conceived, validated, and framed for use in developing-country contexts. Taken together with the conceptual metatypology developed in our previous study [17], they highlight not only where maturity-model research is being produced and how it is distributed across sectors and regions, but also how unevenly empirical validation and contextual sensitivity are incorporated. In this section we discuss the implications of these patterns for knowledge production, model design, and policy in developing economies, structuring the discussion by research question.

4.1. RQ1 Where Is Knowledge on Industry 4.0/5.0 Maturity Models Being Produced, and How Well Represented Are Developing Countries?

RQ1 examined where Industry 4.0/5.0 maturity models are being produced and how strongly developing countries are represented in the existing knowledge base. The descriptive results reveal a clear geographical concentration of model development. Of the 75 maturity models analysed, 77.3% are authored within a single country; almost half originate in developed economies (48%), just over one third in developing economies (36%), and the remainder are framed as global or emerge from mixed-country collaborations. In regional terms, Europe accounts for approximately 60% of all models, followed by Asia (27%) and Latin America (13%), while Africa, North America beyond a single U.S. study, and Oceania are only marginally represented. Taken together, these figures point to a geographically asymmetric and uneven map of knowledge production.
The analysis of co-authorship networks reinforces this picture. Collaborative ties are dense within Europe and along Europe–Latin America and Europe–Asia corridors, whereas South–South collaborations are rare and occur at very low frequency. This structure suggests that the industrial conditions, research agendas, and policy priorities of highly digitalised and innovation-intensive ecosystems have played a disproportionate role in shaping the design space of Industry 4.0/5.0 maturity models. Although this concentration does not imply that models originating in such contexts are intrinsically superior, it does indicate that many frameworks are likely to embed implicit assumptions regarding infrastructure availability, workforce skills, governance capacity, and data readiness that may not hold in less endowed environments.
An additional confounding factor in interpreting the geographical concentration observed in RQ1 is language bias. The review was deliberately restricted to English-language publications indexed in Web of Science and Scopus. This decision was motivated by the objective of ensuring comparability, traceability, and methodological consistency across studies, given that English-language outlets constitute the dominant medium for internationally visible, peer-reviewed research and provide standardized metadata, citation practices, and indexing criteria required for systematic analysis.
At the same time, this choice inevitably privileges research outputs that are visible within English-dominant academic indexing systems. As a consequence, locally developed maturity models published in other languages—such as Spanish or Portuguese in Latin America, or Chinese and Japanese in parts of Asia—are likely to be underrepresented or entirely absent from the analyzed corpus. Accordingly, the observed geographical distribution should not be interpreted as a comprehensive map of global maturity-model development, but rather as a representation of where such models are most visible within international, English-indexed academic outlets.
The apparent dominance of Europe and other developed economies may therefore partially reflect differential indexation and language visibility rather than substantive differences in model-building activity. While the sensitivity analysis mitigates relevance-ranking bias within the selected databases, it does not eliminate language-related visibility constraints. These constraints should be explicitly borne in mind when generalizing the findings of RQ1 beyond the scope of English-indexed academic literature.
From a design perspective, these findings highlight the importance of making contextual assumptions explicit rather than treating advanced industrial settings as the default benchmark. Dimensions, indicators, and maturity pathways should be specified in ways that allow adaptation to different levels of technological intensity, firm size, and institutional maturity. Examples include tiered or phased maturity paths, alternative indicators suitable for low-data environments, or variants explicitly calibrated for resource-constrained contexts. For firms and SMEs in developing economies, the strong origin bias observed in the literature suggests that existing maturity models should be applied with caution: they may function as aspirational reference points, but often require substantial local adaptation and prioritisation before being translated into actionable implementation plans. For policy makers, the concentration of model production in a relatively small group of countries underscores the need to support locally grounded research and co-designed maturity frameworks that reflect national and regional industrial structures, including the prevalence of SMEs, informal sectors, and hybrid production systems.
The empirical basis for RQ1 is primarily descriptive and bibliometric. Country counts and collaboration networks capture patterns of volume and connectivity, but they do not establish causal relationships between geographical origin and model effectiveness. Moreover, the temporal window analysed (2020–2024), the reliance on Web of Science and Scopus, and the use of relevance-ranked subsets introduce potential selection and visibility biases, particularly for contributions from emerging economies, non-English outlets, or less indexed venues. These limitations do not invalidate the observed asymmetries, but they warrant caution when generalising the findings beyond the analysed corpus.
To further assess the robustness of the observed geographical concentration, a sensitivity analysis was conducted using an additional sample of maturity models drawn from ranks 251–500 of the original database search results. While models originating in developed economies remain predominant, this supplementary sample exhibits a higher proportion of models developed in emerging or mixed-economy contexts than the main corpus. This pattern suggests that relevance-based ranking mechanisms may partially amplify the apparent dominance of developed-economy contributions. Accordingly, geographical concentration should be interpreted as relative and conditioned by database ranking and indexation practices, rather than as an absolute reflection of global maturity-model development.
Overall, the evidence indicates that enhancing the global relevance of Industry 4.0/5.0 maturity models requires a more geographically balanced research effort. This includes fostering cross-regional consortia with shared leadership between institutions from developed and developing economies, designing explicit adaptation protocols that distinguish universal from context-dependent components, and systematically documenting applications in low-digital-intensity environments. Such steps would help shift maturity models from tools largely shaped by industrialised contexts toward instruments capable of meaningfully supporting digital transformation across the full spectrum of industrial realities.

4.2. RQ2 Does the Type of Model Origin (Single-Country, Collaboration, or Global Approach) Influence Whether Authors Discuss Applicability to Developing-Country Contexts?

RQ2 examined whether the origin of maturity models, operationalized through authorship configuration, is associated with whether authors explicitly or implicitly discuss applicability in developing-country contexts. The descriptive cross-tabulation between model origin and applicability mention (Table 4) indicates that, across single-country, collaborative, and global models, explicit references to developing-country applicability are generally uncommon, and the majority of studies do not address this issue in the published text.
When explicit and implicit mentions were combined for inferential analysis, the application of an exact Fisher–Freeman–Halton test with Monte Carlo approximation indicated a statistically detectable association between model origin and applicability mention. This result suggests that, within the analyzed corpus, the likelihood that applicability to developing-country contexts is mentioned varies across origin categories. However, given the exploratory nature of the analysis and the limited size of some subgroups, this association should be interpreted cautiously and not as evidence of a causal relationship.
Importantly, the observed association does not imply that collaborative or global models systematically embed stronger contextual sensitivity by design. Rather, the results indicate heterogeneity in reporting practices across origins, with no origin category consistently addressing applicability in a comprehensive or systematic manner. Even among collaborative and global models, references to developing-country contexts remain sparse, suggesting that broader authorship configurations do not automatically translate into explicit engagement with contextual constraints, infrastructural limitations, or institutional conditions characteristic of developing economies.
From a research design perspective, these findings highlight that collaboration per se is not sufficient to ensure contextualization. What appears to matter is how contextual issues are framed and operationalized within the study, and whether researchers with direct experience of developing-country settings play a substantive role in model design, validation, and interpretation. Authorship configuration alone is therefore an imperfect proxy for contextual sensitivity.
The analysis of RQ2 is subject to several limitations. Subgroup sizes for collaborative and global models are relatively small, which increases uncertainty and limits the stability of inferential results. In addition, authorship origin was operationalized using affiliation information and declared scope, which does not capture asymmetries of influence within research teams or the extent to which partners from different regions shape conceptual and methodological choices. Finally, the coding of applicability mention is based exclusively on what is reported in published articles and cannot account for contextual considerations that may have informed the research process but were not explicitly documented. Consequently, the results should be understood as reflecting patterns of reporting within the academic literature, rather than the full range of considerations that may guide maturity-model development in practice.

4.3. RQ3 Does the Development Level of the Country of Origin Influence the Likelihood That the Model Explicitly Considers Developing Countries?

RQ3 examined whether the development level of the country or countries in which a maturity model is developed is associated with the likelihood that the model explicitly or implicitly addresses its applicability in developing-country contexts. The descriptive results reveal clearly differentiated patterns across development-level categories. Among models originating in developed economies, almost all make no reference to applicability in developing-country settings, whereas models originating in developing economies show a substantially higher incidence of both explicit and implicit mentions. Models classified as Global or Mixed follow a pattern closer to that of developed economies, with applicability considerations being largely absent.
The inferential analysis provides exploratory statistical support for this pattern. Importantly, this support emerges from contingency tables that include modest cell sizes in several categories, particularly for Global and Mixed models. While exact tests were used to address sparse data conditions, the resulting associations should be interpreted as suggestive rather than definitive, reflecting tendencies within the analyzed corpus rather than stable population-level effects. Using an exact test appropriate for sparse contingency tables, a statistically detectable association was identified between development level and applicability mention. While this result does not imply causality, it indicates that the likelihood of explicitly addressing developing-country applicability is not randomly distributed across development-level categories within the analyzed corpus.
Any interpretation of this association must, however, clearly distinguish empirical evidence from explanatory inference. One plausible interpretation is a form of contextual proximity: research teams embedded in developing economies may be more directly exposed to constraints related to infrastructure, financing, skills availability, and institutional support, making contextual fit a more salient concern during model design. In contrast, models developed in advanced industrial settings may implicitly assume relatively mature digital ecosystems and higher baseline capabilities, reducing the perceived need to explicitly discuss transferability. This interpretation should be regarded as exploratory rather than definitive.
At the same time, alternative explanations must be explicitly considered. Publication bias may play a role, as authors based in developing-country contexts may face stronger expectations to justify relevance, contextualization, or applicability when publishing in international, English-language journals. Conversely, authors from developed economies may be less compelled to articulate contextual boundaries that are implicitly assumed by editors and reviewers. In addition, indexation and visibility effects associated with Web of Science and Scopus may amplify certain writing conventions and epistemic norms, thereby shaping what is made explicit in published articles without necessarily reflecting differences in underlying design intent. Importantly, the absence of an explicit applicability statement does not imply that contextual considerations were absent from the research process, only that they were not foregrounded in the published text.
This association should therefore not be interpreted in essentialist terms. The operationalization of development level relies on country-level classifications applied to author affiliations, which abstract from substantial heterogeneity within countries, transnational research trajectories, and possible decoupling between the institutional origin of a model and the contexts in which it is validated or applied. In particular, the Global category aggregates heterogeneous configurations under a single label and does not guarantee that models were empirically grounded across diverse development environments.
These findings refine and complement the results of RQ1 and RQ2. RQ1 demonstrated that the production of Industry 4.0/5.0 maturity models is heavily concentrated in developed economies, shaping the dominant academic design space. RQ2 showed that authorship configuration (single-country, collaborative, or global) is not systematically associated with applicability mentions. Taken together, RQ3 suggests that the development context may be associated with whether developing-country applicability is explicitly articulated, but it does not establish the mechanism underlying this association.
From a design and research-policy perspective, the results underscore the importance of explicitly stating contextual assumptions in maturity-model research, regardless of country of origin. Models intended for broad use would benefit from clearly specifying boundary conditions, including minimum infrastructural, organizational, and institutional requirements, as well as potential limitations when applied in resource-constrained settings. At the same time, disentangling contextual proximity effects from publication and indexation biases would require complementary qualitative, comparative, or field-based research designs that go beyond the scope of the present study.
Several limitations of the RQ3 analysis should be acknowledged. The number of models classified as Global or Mixed remains small, which limits the stability of estimates for these categories. Development level was operationalized using World Bank classifications applied to author affiliations, which necessarily simplifies heterogeneous national and subnational realities. Finally, the coding of applicability mentions relies exclusively on what is documented in published articles and cannot capture undocumented adaptations or informal uses of models in practice. These constraints suggest that the observed association should be interpreted as indicative rather than conclusive.

4.4. RQ4 To What Extent Are Existing Models Empirically Validated and in Which Sectors, and How Does This Limit Their Transferability to Emerging Economies?

RQ4 examined how patterns of empirical validation and sectoral focus shape the transferability of Industry 4.0/5.0 maturity models to emerging economies. Rather than focusing solely on whether validation is reported, the analysis highlights how the hlcontext and sectoral anchoring of validation constrain the external validity of existing frameworks.
A central finding is the strong concentration of validation efforts in manufacturing and closely related industrial domains. This concentration systematically privileges environments characterized by relatively high capital intensity, standardized processes, and more mature digital infrastructures. As a result, the empirical foundations of many maturity models implicitly reflect the operational logics and resource assumptions of industrial production settings, while sectors that are strategically important for developing economies—such as agri-food systems, logistics services, healthcare, public administration, or infrastructure-related activities—remain weakly represented in validation designs.
This imbalance has direct implications for transferability. Models validated primarily in advanced manufacturing contexts often embed assumptions about data availability, workforce skills, organizational maturity, and investment capacity that may not hold in other sectors or in resource-constrained environments. When applied beyond their original validation settings, such models risk producing distorted maturity assessments or unrealistic improvement pathways. In this sense, the critical issue is not whether a model has been validated per se, but whether its validation context meaningfully resembles the environments in which it is later deployed.
A key distinction emerging from the analysis concerns the difference between declared multisectoral scope and empirically demonstrated cross-sector applicability. Several models explicitly describe themselves as multisectoral, yet their validation evidence does not substantiate this claim. For example, the Modular Maturity Model for Industry 4.0 (MM.15; [24]) presents itself as broadly applicable across industries; however, its empirical validation is limited to manufacturing-related cases, with no documented application in clearly distinct sectors such as services, logistics, or public organizations. In such cases, multisectoral scope remains largely aspirational, supported by conceptual generality rather than by sectorally differentiated empirical evidence.
By contrast, a small number of models do demonstrate empirically grounded multisectoral validation. The Digital Twin Maturity Model (MM.34; [25]) provides a notable example: its validation strategy combines applications in different industrial contexts (e.g., aerospace component manufacturing and maintenance operations) with simulation-based assessments in non-industrial scenarios, such as smart-city infrastructure use cases. Although still limited in scale, this approach illustrates how maturity constructs can be tested across heterogeneous sectoral logics, strengthening claims of cross-sector applicability.
From an Industry 5.0 perspective, the limitation observed in most models is not simply the absence of human-centric or sustainability-oriented dimensions, but the lack of empirical designs capable of substantiating them. Human-centricity and sustainability are frequently articulated at a conceptual level, yet rarely operationalized through measurable indicators or validated against socio-technical outcomes. The near absence of predictive or longitudinal validation further exacerbates this mismatch, as it prevents assessing whether improvements along such dimensions translate into durable organizational, social, or environmental benefits over time.
Finally, the sectoral patterns identified must be interpreted in light of the deliberate exclusion of practitioner-developed maturity models produced by consultancies, industry associations, and standardization bodies. Many of these frameworks are widely applied in service sectors, logistics, energy, or public organizations, but they fall outside the scope of this review due to the focus on peer-reviewed academic literature. Consequently, the findings of RQ4 characterize the priorities and validation practices of academically developed maturity models rather than the full ecosystem of maturity assessment tools used in practice.
Taken together, the results underscore the need for validation strategies that are both sector-sensitive and context-aware. Strengthening the external validity of maturity models requires clearer differentiation between declared scope and demonstrated applicability, systematic empirical work in underrepresented sectors, and explicit reporting of sector-specific constraints encountered during validation. These steps are particularly critical if maturity models are to function as credible decision-support instruments for firms and policy-makers in emerging-economy contexts.

4.5. RQ5 What Research Gaps Do Authors Themselves Identify as Obstacles to Industry 4.0/5.0 Adoption in Developing Countries?

RQ5 synthesized the research gaps and limitations explicitly identified by primary studies as barriers to Industry 4.0/5.0 adoption, with particular relevance for developing-country contexts. Rather than reiterating their frequency, the analysis focuses on the structural patterns that emerge across these self-reported gaps and their implications for the evolution of maturity-model research.
The most persistent concern across the corpus relates to empirical robustness. Authors repeatedly acknowledge that their models rely on limited forms of evidence—small samples, single-case studies, cross-sectional surveys, or simulations—rather than on systematic multi-site or longitudinal validation. This recurrent admission signals a shared recognition that many maturity models remain weakly tested against real organizational trajectories or performance outcomes. As such, confidence in their predictive capacity and suitability for guiding high-stakes investment or policy decisions remains constrained.
A second cluster of gaps highlights limited sectoral coverage and organizational scope. Numerous studies explicitly note that their models have been developed for, or validated within, narrowly defined industrial contexts, most often manufacturing. At the same time, authors frequently point to the insufficient adaptation of existing frameworks to small and medium-sized enterprises, despite SMEs constituting the dominant organizational form in many economies. These acknowledgements reinforce the observation that much of the maturity-model evidence base is anchored in relatively formalized, capital-intensive settings, which may not reflect the operational realities of firms in developing countries.
A third group of gaps concerns the integration of sustainability, resilience, and human-centric dimensions. While these aspects are increasingly invoked in the discourse surrounding Industry 5.0, authors commonly frame them as future extensions rather than as empirically grounded components of current models. The lack of operational indicators and validation strategies for socio-technical dimensions creates a disconnect between normative ambitions and available measurement tools, particularly in contexts where social and environmental vulnerabilities are pronounced.
Beyond conceptual limitations, the literature also reveals significant operational shortcomings. Many studies recognize the absence of standardized taxonomies, shared metrics, and openly accessible assessment instruments, which hampers replication, cross-model comparison, and cumulative synthesis. Practical tools such as questionnaires, digital platforms, or reference datasets are rarely provided in reusable formats, and longitudinal or cross-country evidence remains scarce. These deficits limit the capacity of maturity models to function as scalable and adaptable instruments beyond their original study settings.
Taken together, the gaps identified by primary studies point to an implementability deficit that is especially salient for emerging economies. From the perspective of firms and policy-makers, maturity models often appear as promising but weakly substantiated artefacts—designed in contexts that differ from their constraints and lacking the operational infrastructure needed for low-cost deployment. Importantly, these gaps do not merely enumerate shortcomings; they delineate a coherent research agenda. Addressing them requires stronger empirical validation across diverse sectors and contexts, explicit attention to SME realities, systematic integration of sustainability and human-centric indicators, and a commitment to openness through shared data, instruments, and codebooks. In this sense, the self-identified gaps in the literature function as a roadmap for reorienting maturity-model development toward the practical needs of developing economies.

5. Conclusions

This article has examined how Industry 4.0/5.0 maturity models are being validated, scoped, and framed for use in developing-country contexts, based on an academic literature analysis of 75 models published between 2020 and 2024. By combining descriptive profiling, inferential statistics, and text-based synthesis of 562 gap statements, we mapped the geographical origin of these models, their empirical grounding, their sectoral focus, and the main limitations that authors themselves recognise as barriers to adoption in emerging economies.

5.1. Contributions and Main Findings

The study makes four main contributions. First, it documents the geography of knowledge production on Industry 4.0/5.0 maturity models. Most frameworks are proposed in single-country studies from developed economies, with Europe as the dominant hub and limited participation from underrepresented regions. Although more than one-third of models originate in developing or emerging economies, the overall landscape remains strongly asymmetric.
Second, the analysis shows that the type of authorship (single-country, collaborative, global) is not significantly associated with whether a study mentions applicability to developing countries, whereas the development level of the country of origin is. Models originating from, or co-led with, developing economies are substantially more likely to include explicit or implicit statements about their use in resource-constrained contexts, while models from developed countries largely omit this discussion.
Third, we provide a systematic overview of empirical validation strategies and sectoral scope. Most models report some empirical evidence, but validation is often limited to small samples, cross-sectional surveys, or single-case studies; only a minority provide more robust multiple-case or large-sample designs, and around 15% have no reported validation at all. Sectoral coverage is heavily skewed toward manufacturing and manufacturing SMEs, with services, agri-food, healthcare, and public-sector applications appearing only sporadically. Multisector models are rare but tend to rely on stronger empirical designs, especially multiple-case applications.
Fourth, by synthesizing 562 gap and limitation statements, we outline a research agenda grounded in the concerns of the primary studies themselves. Recurrent themes include insufficient empirical rigor, narrow sectoral and SME coverage, limited predictive or longitudinal validation, and weak operationalization of sustainability and human-centric dimensions associated with Industry 5.0. While such dimensions are increasingly invoked at a conceptual level, they are rarely anchored in measurable indicators or tested through validation designs capable of assessing their substantive impact over time. Together, these issues point to a persistent implementability deficit that constrains the usefulness of maturity models for guiding digital transformation in developing-country settings.

5.2. Implications for Research and Practice

For designers of maturity models and researchers, the findings underscore the importance of treating contextual applicability and empirical validation as first-order design requirements rather than ancillary considerations. Models intended for use in emerging economies should (i) make assumptions about infrastructure, skills, and organizational capabilities explicit; (ii) offer tiered and resource-aware maturity pathways suitable for SMEs; and (iii) operationalize Industry 5.0 principles—such as human-centricity and sustainability—through measurable indicators linked to predictive or longitudinal validation designs. Strengthening empirical rigor will often require mixed-method approaches that combine multi-site case studies, surveys, and longitudinal follow-ups capable of relating maturity trajectories to organizational, social, or environmental outcomes.
For policy-makers and practitioners in developing countries, the results caution against the uncritical transfer of maturity models developed in highly digitalised industrial environments. When selecting or adapting a model, decision-makers should interrogate its origin, validation track record, sectoral scope, and relevance to local constraints. Public programmes that promote Industry 4.0/5.0 adoption can play a catalytic role by supporting co-design and validation of models with local firms, funding open assessment tools and reference datasets, and encouraging the inclusion of sustainability, resilience, and inclusion targets alongside purely technological metrics. In this way, maturity models can evolve from static diagnostic checklists into dynamic instruments for planning, monitoring, and governing digital transformation.

5.3. Limitations

Several limitations should be considered when interpreting the findings of this study. First, the corpus is restricted to peer-reviewed publications indexed in Web of Science and Scopus during the 2020–2024 period. Although these databases provide broad coverage of high-impact academic literature, they do not capture the full universe of Industry 4.0/5.0 maturity models. In particular, practitioner-oriented models developed by consultancies, industrial associations, and standardization bodies, as well as grey literature and proprietary frameworks, fall outside the scope of this review. As a result, the conclusions of this article apply specifically to maturity models developed and disseminated within the academic literature and should not be generalized to the wider ecosystem of industrial assessment tools used in practice.
Second, the reliance on relevance-based ranking to truncate the initial search results (top 250 records per database) introduces a potential selection and visibility bias. Although this heuristic was applied transparently and consistently across both databases, relevance algorithms are proprietary and may amplify the prominence of contributions from well-established research communities, journals, and English-language outlets. To assess the robustness of the observed geographical patterns, a sensitivity analysis was conducted using an additional sample of 20 maturity models drawn from ranks 251–500. This supplementary analysis indicates a higher relative presence of models originating in developing or mixed-economy contexts than in the main corpus, suggesting that relevance-based ordering may partially overstate the dominance of developed-economy contributions. Nevertheless, the results should be interpreted as relative patterns conditioned by database ranking and indexation practices, rather than as an exhaustive representation of global model development.
Third, the corpus is limited to English-language publications. This language restriction may confound interpretations of geographical concentration, as maturity models developed and published in other languages (e.g., Spanish, Portuguese, Chinese, or Japanese) are likely underrepresented. Consequently, locally grounded models addressing developing-country contexts may remain invisible to the present analysis. While this limitation does not invalidate the patterns identified within the English-indexed academic literature, it constrains the extent to which the findings can be extrapolated to global knowledge production as a whole.
Fourth, the operationalization of development level constitutes an additional limitation. Assigning models to development categories based on the affiliation country of the lead author and World Bank country classifications necessarily simplifies complex realities. This approach does not capture substantial heterogeneity within countries, transnational research trajectories and diaspora patterns, or cases in which the institutional origin of a model differs from the context in which it is validated or applied. In addition, the Global category aggregates heterogeneous configurations under a single label and should not be interpreted as evidence that models were empirically grounded across diverse development environments. Alternative or complementary proxies, such as the geographical location of empirical validation, the primary application setting, or multi-level coding schemes distinguishing origin, validation, and use contexts, could provide a more fine-grained representation of contextual grounding and represent a promising direction for future research.
Fifth, all variables were coded exclusively from information explicitly reported in the published articles. Under-reporting of validation procedures, sectoral scope, or applicability conditions may therefore bias the results. In particular, the absence of an explicit applicability statement does not necessarily imply that a model cannot be used in developing-country contexts, but rather that such considerations were not documented by the authors. Similarly, the degree of validation or contextual adaptation applied in practice may exceed what is reported in the academic publications.
Sixth, although research gaps were identified using a hybrid AI–human workflow with complete human verification, the AI-assisted extraction process operates at the level of local textual excerpts. As evidenced by the stratified audit, a small number of gaps were conceptually valid at the document level but were not fully supported by the specific excerpt initially selected by the automated extractor. These cases required manual correction and highlight a limitation related to excerpt-level anchoring rather than to the validity of the identified gaps themselves. While systematic human verification mitigates the risk of hallucination and misclassification, the quality of excerpt selection remains sensitive to document structure and reporting practices. Future work could strengthen traceability by aggregating evidence from multiple excerpts or by incorporating explicit document-level referencing during automated extraction. In addition, gap verification was conducted by a single expert reviewer; although appropriate for an audit-oriented validation, future studies could employ dual coding and inter-rater reliability measures to further enhance robustness.
Seventh, the quantitative analyses are exploratory and rely on relatively small subgroups for certain categories (e.g., collaborative and global models). To address violations of asymptotic assumptions, exact tests with Monte Carlo approximation were employed where appropriate. While this approach improves robustness under sparse data conditions, the resulting p-values remain sensitive to recoding decisions and small changes in cell counts. In addition, the operationalization of applicability as a binary variable for inferential purposes entails a loss of granularity, which may obscure distinctions between explicit and implicit forms of contextual consideration. Accordingly, statistically detectable associations should be interpreted cautiously and not as evidence of stable or generalizable effects. In addition, effect sizes were reported to support substantive interpretation of associations, but these estimates remain sensitive to small cell counts and should be interpreted as corpus-specific rather than population-level parameters.
Finally, and most importantly, this review does not assess the real-world effectiveness or impact of the analyzed maturity models. The study maps the academic design space of Industry 4.0/5.0 maturity models as conceptual and methodological artefacts, focusing on their origin, validation claims, sectoral scope, and stated applicability. It does not evaluate whether the application of these models leads to successful digital transformation outcomes, performance improvements, or sustained capability development in practice. This constitutes a substantive boundary of the study and highlights the need for future research linking maturity assessments to longitudinal evidence on organizational, sectoral, and policy outcomes, particularly in developing-country settings.

5.4. Directions for Future Research

Future work could extend and deepen this analysis in several directions. Longitudinal and cross-country studies are needed to track how industrial digital maturity evolves over time and to assess whether maturity scores derived from existing models predict changes in performance, resilience, or sustainability. Sectoral blind spots identified in this review—including agri-food, healthcare, logistics, and public services—warrant targeted efforts to design and validate models that reflect their specific constraints and opportunities, especially in developing-country contexts. There is also a need for greater standardisation and interoperability across frameworks, including shared taxonomies, benchmark datasets, and open-source assessment tools that enable replication and comparative meta-analyses. Finally, as Industry 5.0 matures, future maturity models should more systematically incorporate human-centric, social, and environmental dimensions, and empirically test how these interact with technological and economic indicators.
Taken together with our previous metatypology of Industry 4.0/5.0 maturity models [17], which focused on the evolution of model scope, level structures, dimensions, and enabling technologies, this article completes a two-part research programme. The earlier contribution clarifies how industrial digital maturity is conceptualised at the interface between Industry 4.0 and Industry 5.0, while the present study examines how those models are empirically validated, sectorally scoped, and framed for use in developing-country contexts. Both strands point toward the same conclusion: a qualitative leap in validation practices, contextual sensitivity, and openness is required if maturity models are to function as effective instruments for a fair and sustainable digital transformation that narrows, rather than widens, the gap between developed and developing economies. Finally, as Industry 5.0 matures, future maturity models should more systematically incorporate human-centric, social, and environmental dimensions, operationalize them through observable indicators, and empirically test their effects using longitudinal or predictive validation designs.

Author Contributions

Conceptualization, D.R.D. and A.L.P.; methodology, D.R.D.; software, A.L.P.; validation, D.R.D., A.L.P. and M.B.I.A.; formal analysis, D.R.D.; investigation, D.R.D., A.L.P. and M.B.I.A.; resources, A.L.P.; data curation, M.B.I.A.; writing—original draft preparation, D.R.D.; writing—review and editing, D.R.D., A.L.P. and M.B.I.A.; visualization, D.R.D.; supervision, A.L.P. and M.B.I.A.; project administration, A.L.P.; funding acquisition, A.L.P. All authors have read and agreed to the published version of the manuscript.

Funding

This research was supported primarily by the doctoral training program of the Transilvania University of Brasov. The funding organization had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript, or in the decision to publish the results.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

All data used to generate the figures and results in this study, which were extracted from the 75 primary literature sources included in the main corpus, are available in an open repository (version v1) under the title Industrial Digital Maturity (I4→I5)—75 Models, 2020–2024, https://zenodo.org/records/17454113, accessed on 27 October 2025. The repository includes the full extracted dataset and a BibTeX file listing all resources used in the research, accessible via: https://zenodo.org/records/17454113, accessed on 27 October 2025 (DOI: 10.5281/zenodo.17454112). In addition, the dataset used for the sensitivity analysis addressing potential selection bias introduced by relevance-based database truncation is publicly available in a separate Zenodo record under the title Sensitivity Analysis Dataset for Industry 4.0/5.0 Maturity Models (WoS/Scopus Ranks 251–500), https://zenodo.org/records/18215243, accessed on 27 October 2025. This supplementary dataset contains the 20 maturity-model candidates analysed as part of the robustness check, including eligibility decisions and coded variables, and is used exclusively for sensitivity analysis purposes. It is not pooled with the primary dataset.

Acknowledgments

The authors gratefully acknowledge the teams and services that supported this study. We thank the Web of Science and Scopus platforms for providing the bibliographic sources used to construct the dataset. We are indebted to the developers and maintainers of the tools employed in our workflow, including Bibliometrix for bibliometric analyses, Python (pandas) and SciPy for statistical processing, and spaCy for text preprocessing, as well as Power Map for geographic visualizations. We also appreciate the constructive comments from colleagues at the Department of Manufacturing Engineering, Transilvania University of Brasov, and the Technological University of Havana (CUJAE), which improved earlier versions of the manuscript, and we thank the assistants who supported data extraction and curation. Any remaining errors are the sole responsibility of the authors.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
4IRFourth Industrial Revolution
AIArtificial Intelligence
BibTEXReference format/standard for LATEX bibliographies
dfDegrees of freedom
DOIDigital Object Identifier
DTDigital Transformation
GPTGenerative Pre-trained Transformer (family of AI models)
I4.0Industry 4.0
I5.0Industry 5.0
IoTInternet of Things
KPI(s)Key Performance Indicator(s)
MMMaturity Model
MMsMaturity Models
MR MapTraceability map Research Question—Method—Result—Evidence
N, nSample size (N: total; n: subsample)
NLPNatural Language Processing
pp-value (statistical significance)
R&DResearch & Development
RQResearch Question(s)
SME(s)Small and Medium-sized Enterprise(s)
α Significance level
χ 2 Chi-square test

Appendix A. Sampling Traceability for the Sensitivity Analysis

This appendix documents the complete, transparent, and reproducible sampling procedure used to construct the sensitivity-analysis sample addressing potential selection bias induced by relevance-based truncation of database search results.

Appendix A.1. Universe Definition

The sampling universe consisted of all records ranked 251–500 in the original database queries. A total of 250 records were exported from Web of Science and 249 from Scopus. After removal of one duplicated reference, the final universe comprised 499 unique records. All records were exported in EndNote XML format and converted to a spreadsheet for subsequent processing.

Appendix A.2. High-Sensitivity Candidate Preselection

To reduce the number of false positives prior to manual screening, an automated preselection step was applied to the full universe of 499 records. Titles, abstracts, and keywords were screened for the presence of maturity- or readiness-model terminology, including expressions such as “maturity model”, “readiness model”, “capability maturity”, and “maturity assessment”. This preselection was intentionally designed to maximise recall rather than precision, accepting the presence of non-eligible records. The resulting high-sensitivity candidate list comprised 122 records.

Appendix A.3. Random Sampling and Reproducibility

Random sampling was implemented in Python using uniform sampling without replacement via the random.sample function. To ensure full reproducibility, a fixed random seed was specified (seed = 20,260,109). Sampling was performed exclusively on the candidate list of 122 records generated in the preselection step.

Appendix A.4. Iterative Manual Screening and Replacement

An initial random draw of 20 candidate records was manually screened against the eligibility criterion used throughout the review, namely that the article explicitly proposes, adapts, or applies a maturity or readiness model (or an equivalent evaluative framework). Records failing to meet this criterion were excluded. To maintain a final sample size of 20 eligible studies, excluded records were iteratively replaced through additional random draws from the remaining candidate pool, without replacement. This process was repeated until 20 eligible maturity-model articles were identified.

Appendix A.5. Final Counts and Consistency with the Main Corpus

Across all iterations, a total of 60 unique candidate records were manually reviewed, resulting in 20 eligible and 40 excluded studies. Table A1 summarises the complete sampling flow. All eligible records were coded using the same codebook, variable definitions, and boundary rules applied to the main corpus of 75 models. The sensitivity-analysis sample was used exclusively as a robustness check and was not pooled with the primary dataset.
Table A1. Sampling traceability for the sensitivity analysis (R1–SEL–BIAS).
Table A1. Sampling traceability for the sensitivity analysis (R1–SEL–BIAS).
StageN
Universe (WoS + Scopus, ranks 251–500)499
Candidate records after term-based preselection122
Manually screened records60
Eligible maturity-model articles20
Excluded records40

Appendix A.6. Reproducible Sampling Snippet

The following Python code illustrates the reproducible random sampling procedure used in the sensitivity analysis:
  • import random
  • random.seed (20260109)
  • sample = random.sample (candidate_ids, 20)  # uniform
  •   sampling without replacement

Appendix B. Sector Normalization Criteria for RQ4

To enable a statistically robust sectoral analysis in RQ4, target sectors reported in the primary studies were normalized into a reduced set of analytically meaningful categories. This step was necessary because sector labels across the literature exhibit substantial heterogeneity in terminology, granularity, and scope, ranging from broad industrial domains (e.g., “manufacturing”) to highly specific or compound descriptions (e.g., “manufacturing SMEs in the automotive supply chain”).

Appendix B.1. Normalization Principles

The normalization followed four guiding principles:
1.
Conceptual coherence: Sector labels referring to the same industrial domain were grouped under a single normalized category, even when expressed at different levels of specificity.
2.
Minimal aggregation: Sectors were aggregated only to the extent required to ensure analytical tractability and avoid excessive fragmentation.
3.
Transparency and reproducibility: All normalization decisions were rule-based and consistently applied across the dataset.
4.
Preservation of multisector intent: Models explicitly claiming applicability to multiple distinct sectors were retained as a separate Multi-sector category.

Appendix B.2. Operational Mapping Rules

The following rules were applied:
  • Sector labels containing terms such as manufacturing, industrial production, or closely related variants were mapped to the normalized category Manufacturing.
  • Labels explicitly referring to manufacturing SMEs (e.g., “manufacturing SMEs”, “industrial SMEs”) were retained as manufacturing-related and included in Manufacturing, but were separately visualized at the subsector level for descriptive purposes.
  • Sector labels referring to logistics, supply chains, or distribution activities were mapped to Logistics & Supply Chain.
  • Labels referring to textiles, apparel, or garment production were mapped to Textile & Apparel.
  • Models explicitly declaring applicability to two or more distinct sectors were classified as Multi-sector, regardless of whether the individual sectors were listed exhaustively.
  • Models that did not specify any target sector and positioned themselves as applicable to “any organization” or “any industry” were classified as Undefined.
  • Highly specific or niche sectors that did not recur across studies (e.g., seafood processing, construction) were grouped under Other/Niche sectors for inferential analysis.

Appendix B.3. Worked Examples (Borderline Cases)

To enhance transparency and facilitate reproducibility of the sector normalization procedure, this subsection provides worked examples illustrating how borderline or ambiguous cases were coded and normalized in practice. These examples reflect recurrent situations encountered in the corpus and demonstrate the application of the mapping rules defined above.

Appendix B.3.1. Example A—Automotive (Manufacturing)

Alcácer et al. apply an Industry 4.0 readiness/maturity assessment in “three departments … of an automotive company” within the Portuguese automotive industry. Although the empirical setting is clearly automotive, the model is designed and validated within a manufacturing production context. Coding decision: Sectorization = Specific; Target_Sector = Automotive (manufacturing); normalized category = Manufacturing.

Appendix B.3.2. Example B—Textiles & Apparel (Manufacturing)

The study To diagnose Industry 4.0 by maturity model … explicitly targets “Moroccan-based clothing enterprises” operating in the apparel domain. The assessment framework, indicators, and empirical validation are grounded in manufacturing activities related to garment production. Coding decision: Sectorization = Specific; Target_Sector = Textiles & Apparel; normalized category = Manufacturing.

Appendix B.3.3. Example C—Cross-Cutting Logistics (Function-First Rule)

Facchini et al. develop “a maturity model for Logistics 4.0” focused on logistics processes and subsystems (e.g., purchasing, production logistics, distribution, and after-sales logistics). Although one empirical application involves an automotive manufacturer, the object of assessment is the logistics function rather than an industry vertical. Following a function-first coding rule, the model is classified independently of the host industry. Coding decision: Sectorization = Specific; Target_Sector = Logistics/Transport/Storage; normalized category = Logistics & Supply Chain.

Appendix B.4. Relation to Statistical Analysis

Normalized sector categories were used exclusively for inferential testing (goodness-of-fit analysis) to ensure adequate expected cell frequencies and interpretability. Descriptive figures additionally report subsector-level detail to illustrate within-category concentration patterns, particularly within manufacturing. This two-level representation allows both statistical robustness and substantive transparency regarding sectoral diversity.

Appendix C. GPT Agent Instructions for Maturity Model Extraction

Systems 14 00134 i009aSystems 14 00134 i009bSystems 14 00134 i009c

Appendix D. GPT Prompt for Research Gap Extraction

This appendix documents the prompt used to guide the GPT-based agent in the pre-extraction of research gaps and limitations (RQ5). The agent was employed as a primary extractor to identify candidate gaps from the full text of each article, while all extracted outputs were subsequently verified and, where necessary, corrected by human reviewers.

Prompt Instructions

You are part of a research team conducting a systematic review of Industry 4.0 and Industry 5.0 maturity models. Your task is to identify and extract research gaps, limitations, and future research opportunities explicitly stated in the article.
Proceed step by step and follow the instructions below carefully. Do not infer or speculate beyond what is stated in the text.
Step 1: Identify relevant statements Scan the full article and locate sentences or paragraphs where the authors:
  • describe limitations of their own study (e.g., methodological, empirical, contextual);
  • identify gaps or shortcomings in the existing literature or body of knowledge;
  • explicitly propose directions or opportunities for future research.
Step 2: Extract verbatim evidence For each identified gap, copy the exact sentence(s) from the article that support it. Preserve the original wording and do not paraphrase at this stage.
Step 3: Classify the type of gap Assign exactly one of the following categories:
  • Prior gap in the literature: statements indicating a lack, scarcity, or limitation in existing research.
  • Limitation of the current study: statements describing constraints, weaknesses, or boundaries of the study being reported.
  • Future research opportunity: statements explicitly proposing or recommending future work.
Step 4: Provide a standardized summary Produce a concise, standardized description of the gap that captures its core meaning while remaining faithful to the extracted text. Do not generalize beyond the scope of the original statement.
Step 5: Output format Present the results in a table with the following columns:
  • Article title;
  • Full reference;
  • Exact quoted text;
  • Gap type (one of the three categories);
  • Standardized and summarized gap description.
If no research gaps, limitations, or future research opportunities are explicitly stated in the article, indicate this clearly.
This prompt was applied consistently across all articles in the corpus. The resulting outputs served as candidate extractions and were not treated as final codings without human verification.

Appendix E. Corpus of Analyzed Maturity Models (MM.01–MM.75)

This appendix documents the complete corpus of the 75 maturity models (MM.01–MM.75) that constitute the core of the Academic Literature Analysis conducted in this study. Each model is assigned an internal identifier (MM.xx), which is used consistently throughout the article to ensure traceability between the narrative, figures, and tables.
Table A2 summarizes the identification, industrial scope, and bibliographic information of each maturity model, including the model code (MM.xx), model name, industrial generation (Industry 4.0, hybrid 4.0/5.0, or Industry 5.0), year of publication, and reference in APA format.
Table A2. Overview of the 75 maturity models included in the corpus (MM.01–MM.75).
Table A2. Overview of the 75 maturity models included in the corpus (MM.01–MM.75).
IDMaturity ModelScopeReferenceYear
MM.01RA–RE–RI–RO Maturity Model (Production Management as-a-Service)Industry 4.0Abner et al. [26]2020
MM.02IMPULSIndustry 4.0 Alcácer et al. [27]2022
MM.03Maturity Framework for SMEs in Industry 4.0Industry 4.0 Amaral and Peças [28]2021
MM.04I4.0 MMIndustry 4.0 Angreani et al. [29]2024
MM.05Fuzzy Maturity ModelIndustry 5.0Bajic et al. [30]2023
MM.06Agca et al. Maturity Model applied to inventory managementIndustry 4.0 Barbalho and Dantas [31]2021
MM.07Warehouse 4.0 Maturity Model for SMEsIndustry 4.0Benmimoun et al. [32]2024
MM.08Industry 4.0 with Industry 5.0 aspects (sustainability)Industry 4.0 & 5.0 Bernhard and Zaeh [33]2023
MM.09Maturity Model for SMEs in Industry 4.0Industry 4.0 Bohorquez and Gil-Herrera [34]2022
MM.10ECO Maturity ModelIndustry 4.0Bretz et al. [35]2022
MM.11Fuzzy-logic-based Maturity Model for OSCMIndustry 4.0 Caiado et al. [36]2021
MM.12Industry 4.0 Maturity Model (Portugal)Industry 4.0 Castelo-Branco et al. [37]2022
MM.13Smart Logistics Maturity Model for SMEsIndustry 4.0 Chaopaisarn and Woschank [38]2021
MM.14Maturity Model for Industry 4.0 adoption in Passenger Railway CompaniesIndustry 4.0 Chaves Franz et al. [39]2024
MM.15Modular Maturity Model for Industry 4.0Industry 4.0 Çinar et al. [24]2021
MM.16IPM (Industry 4.0 Perception Maturity)Industry 4.0 Ciravegna-Martins-da fonseca et al. [40]2024
MM.17Maturity Model for General Contractors in Industry 4.0Industry 4.0 Das et al. [41]2024
MM.18S3RM Maturity Model for smart and sustainable supply chainsIndustry 4.0 & 5.0 Demir et al. [42]2023
MM.19Maturity Model based on TQMIndustry 4.0 Elibal and Özceylan [43]2024
MM.20Diagnostics of Opportunities—Maturity Model for Digital TransformationIndustry 4.0Ericson Öberg et al. [44]2024
MM.21Maturity Model for Logistics 4.0Industry 4.0 Facchini et al. [45]2020
MM.223D-CUBE Readiness Model for Industry 4.0Industry 4.0 Felippes et al. [46]2022
MM.23FITradeoff Industry 4.0 Maturity ModelIndustry 4.0 Ferreira et al. [47]2024
MM.24Data Science Maturity Model (DSMM)Industry 4.0 Gökalp et al. [48]2021
MM.25Industry X.0 Fuzzy Inference EngineIndustry 4.0 Gomes and Basilio [49]2024
MM.26Maturity Model for sPSSIndustry 4.0Heinz et al. [50]2022
MM.27Deloitte-based Digital Maturity ModelIndustry 4.0 Herceg et al. [51]2020
MM.28Singapore Smart Industry Readiness Index (SIRI) adapted to the Moroccan textile industryIndustry 4.0 Jamouli et al. [52]2023
MM.29Upper Austria Industry 4.0 Maturity ModelIndustry 4.0 Kieroth et al. [53]2022
MM.30Maturity model for digital transformation in the manufacturing industryIndustry 4.0 Kırmızı and Kocaoglu [54]2022
MM.31Maturity model to improve company performance through Industry 4.0Industry 4.0 Koldewey et al. [55]2022
MM.32Capability-based maturity model for smart manufacturingIndustry 4.0 Lin et al. [56]2020
MM.33Singapore Smart Industry Readiness Index (SIRI)Industry 4.0 Lin et al. [57]2020
MM.34Digital Twin Maturity Model (DTMM)Industry 4.0 Liu et al. [25]2024
MM.35Innovative Capability Maturity ModelIndustry 4.0 Lookman et al. [58]2022
MM.36COMMA 4.0 (Comprehensive I4.0 Maturity Assessment Model)Industry 4.0Lukhmanov et al. [59]2022
MM.37Maturity Framework for Readiness toward Industry 5.0Industry 5.0 Madhavan et al. [60]2024
MM.38Maturity model for MSMEs in Industry 4.0Industry 4.0Magdalena et al. [61]2021
MM.39Industry 4.0 Maturity IndexIndustry 4.0 Magnus [62]2023
MM.40Lean Smart Maintenance Maturity Model (LSM MM)Industry 4.0 Maier et al. [63]2020
MM.41I4.0 Competency Maturity Model (I4.0CMM)Industry 4.0 Maisiri et al. [64]2021
MM.42Maturity Model for Industry 4.0 integration in any companyIndustry 4.0Melnik et al. [65]2020
MM.43Maturity Model for the Autonomy of Manufacturing SystemsIndustry 4.0 Mo et al. [66]2023
MM.44CUDIE Model (Capability to Utilize Data in Industrial Enterprises)Industry 4.0Nausch et al. [67]2020
MM.45CCMS 2.0 (Company CoMpaSs) with an AI focusIndustry 4.0 Nick et al. [68]2022
MM.46Company Compass (CCMS) 2.0Industry 4.0 Nick et al. [69]2021
MM.47CCMSIndustry 4.0 Nick et al. [70]2020
MM.48CCMS2.0e (with AI)Industry 4.0 Nick et al. [71]2024
MM.49TOE-based Digital Maturity ModelIndustry 4.0 P. Senna et al. [72]2023
MM.50RAISE 4.0Industry 4.0Pan Nogueras et al. [73]2022
MM.51VPi4 (Industry 4.0 Index)Industry 4.0 Pech and Vrchota [74]2020
MM.52Process Model for the Implementation of Industry 4.0 Use Cases in SMEsIndustry 4.0 Peukert et al. [75]2020
MM.53Maturity Model for Machine Tool CompaniesIndustry 4.0 Rafael et al. [76]2020
MM.54Maturity Model for Smart Manufacturing in SMEs in MalaysiaIndustry 4.0 Rahamaddulla et al. [77]2021
MM.55SSTRA (Smart SME Technology Readiness Assessment)Industry 4.0 Saad et al. [78]2021
MM.56Maturity Model for Urban Smart FactoriesIndustry 4.0 Sajadieh and Noh [11]2024
MM.57LM4I4.0 Maturity Model for manufacturing SMEs in developing countriesIndustry 4.0 Sajjad et al. [79]2024
MM.58Industry 4.0 Maturity Model ProposalIndustry 4.0 Santos and Martinho [80]2020
MM.59Maturity model for Digital Twins in battery cellsIndustry 4.0 Schabany et al. [81]2023
MM.60Pay-Per-X Maturity Model (PPX)Industry 4.0Schroderus et al. [82]2021
MM.61Maturity model to assess the impact of Industry 4.0 on SMEsIndustry 4.0 Semeraro et al. [83]2023
MM.62I4MMSME Maturity Model for manufacturing SMEsIndustry 4.0Simetinger and Basl [84]2022
MM.63DigiCoMIndustry 4.0Steinlechner et al. [85]2021
MM.64Readiness Assessment of SMEs in Transitional EconomiesIndustry 4.0Suleiman et al. [86]2021
MM.65Maturity Model for Worker 4.0 adoptionIndustry 4.0Treviño-Elizondo and García-Reyes [87]2021
MM.66ECDMM4.0–Employee Competency Development Maturity Model for Industry 4.0Industry 4.0 Treviño-Elizondo and García-Reyes [88]2023
MM.67Maturity Model to Become a Smart Organization based on Lean and I4.0Industry 4.0 Treviño-Elizondo et al. [89]2023
MM.68Maturity model for digital twinsIndustry 4.0 Uhlenkamp et al. [90]2022
MM.69SANOL Industry 4.0 Maturity ModelIndustry 4.0 Ünal et al. [91]2022
MM.70Industry 4.0 Maturity Model for Manufacturing in IndiaIndustry 4.0 Wagire et al. [92]2021
MM.71Maturity model for technological integration in industrial companiesIndustry 4.0Widmer et al. [93]2022
MM.72Logistics 4.0 Maturity ModelIndustry 4.0 Zoubek and Simon [94]2021
MM.73Environmental Maturity Model for Industry 4.0Industry 4.0 Zoubek et al. [95]2021
MM.74Maturity model to assess the automation of production processesIndustry 5.0 Hetmanczyk [7]2024
MM.75Maturity model for implementing Logistics 5.0 based on decision-support systemsIndustry 5.0 Trstenjak et al. [96]2022

Appendix F. Illustrative Examples of Notable Models

  • MM.12—Multi-sector model (Portugal, 2022) validated through a large-scale survey ( n = 323 ), representing the broadest company coverage among maturity model validations.
  • MM.34—Model with hybrid validation: applied in multiple real-world sectors (e.g., aerospace, energy) and simulated scenarios, exemplifying the combination of demo and real-case approaches.
  • MM.37—Sector-specific model for seafood processing (2024), introducing Industry 5.0 to an underrepresented sector; validated through a survey of n = 42 small fishing enterprises (SMEs).

References

  1. Matthess, M.; Kunkel, S. Structural change and digitalization in developing countries: Conceptually linking the two transformations. Technol. Soc. 2020, 63, 101428. [Google Scholar] [CrossRef] [Scilit]
  2. Peerally, J.A.; Santiago, F.; De Fuentes, C.; Moghavvemi, S. Towards a firm-level technological capability framework to endorse and actualize the Fourth Industrial Revolution in developing countries. Res. Policy 2022, 51, 104563. [Google Scholar] [CrossRef] [Scilit]
  3. Manda, M.I.; Ben Dhaou, S. Responding to the challenges and opportunities in the 4th Industrial revolution in developing countries. In Proceedings of the 12th International Conference on Theory and Practice of Electronic Governance, Melbourne, Australia, 3–5 April 2019; pp. 244–253. [Google Scholar]
  4. Aly, H. Digital transformation, development and productivity in developing countries: Is artificial intelligence a curse or a blessing? Rev. Econ. Political Sci. 2022, 7, 238–256. [Google Scholar] [CrossRef] [Scilit]
  5. Rodríguez-Espíndola, O.; Chowdhury, S.; Dey, P.K.; Albores, P.; Emrouznejad, A. Analysis of the adoption of emergent technologies for risk management in the era of digital manufacturing. Technol. Forecast. Soc. Change 2022, 178, 121562. [Google Scholar] [CrossRef] [Scilit]
  6. Chirumalla, K.; Oghazi, P.; Nnewuku, R.E.; Tuncay, H.; Yahyapour, N. Critical factors affecting digital transformation in manufacturing companies. Int. Entrep. Manag. J. 2025, 21, 1–52. [Google Scholar] [CrossRef] [Scilit]
  7. Hetmanczyk, M.P. A Method for Evaluating the Maturity Level of Production Process Automation in the Context of Digital Transformation-Polish Case Study. Appl. Sci. 2024, 14, 4380. [Google Scholar] [CrossRef] [Scilit]
  8. Issa, A.; Hatiboglu, B.; Bildstein, A.; Bauernhansl, T. Industrie 4.0 roadmap: Framework for digital transformation based on the concepts of capability maturity and alignment. Procedia Cirp 2018, 72, 973–978. [Google Scholar] [CrossRef] [Scilit]
  9. Latino, M.E. A maturity model for assessing the implementation of Industry 5.0 in manufacturing SMEs: Learning from theory and practice. Technol. Forecast. Soc. Change 2025, 214, 124045. [Google Scholar] [CrossRef] [Scilit]
  10. Teichert, R. Digital transformation maturity: A systematic review of literature. Acta Univ. Agric. Et Silvic. Mendel. Brun. 2019, 67, 1673–1687. [Google Scholar] [CrossRef] [Scilit]
  11. Sajadieh, S.M.M.; Noh, S.D. Towards Sustainable Manufacturing: A Maturity Assessment for Urban Smart Factory. Int. J. Precis. Eng. Manuf.—Green Technol. 2024, 11, 909–937. [Google Scholar] [CrossRef] [Scilit]
  12. Fortier, J.; Gamache, S.; Fonrouge, C. Integrating Sustainable Performance into the Digital Maturity Models for SMEs in Manufacturing. Appl. Sci. 2025, 15, 4041. [Google Scholar] [CrossRef] [Scilit]
  13. Hizam-Hanafiah, M.; Soomro, M.A.; Abdullah, N.L. Industry 4.0 Readiness Models: A Systematic Literature Review of Model Dimensions. Information 2020, 11, 13. [Google Scholar] [CrossRef] [Scilit]
  14. Bertolini, M.; Esposito, G.; Neroni, M.; Romagnoli, G. Maturity Models in Industrial Internet: A Review. In Proceedings of the 25th International Conference on Production Research Manufacturing Innovation (ICPR)—Cyber Physical Manufacturing, Amsterdam, The Netherlands, 9–14 August 2019; Volume 39, pp. 1854–1863. [Google Scholar] [CrossRef] [Scilit]
  15. Hein-Pensel, F.; Winkler, H.; Brückner, A.; Wölke, M.; Jabs, I.; Mayan, I.J.; Kirschenbaum, A.; Friedrich, J.; Zinke-Wehlmann, C. Maturity assessment for Industry 5.0: A review of existing maturity models. J. Manuf. Syst. 2023, 66, 200–210. [Google Scholar] [CrossRef] [Scilit]
  16. Onyeme, C.; Liyanage, K. A Critical Review of Smart Manufacturing & Industry 4.0 Maturity Models: Applicability in the O&G Upstream Industry. In Proceedings of the 18th International Conference on Manufacturing Research (ICMR)/35th National Conference on Manufacturing Research (NCMR), Amsterdam, The Netherlands, 7–10 September 2021; Volume 15, pp. 347–354. [Google Scholar] [CrossRef] [Scilit]
  17. Reyes Domínguez, D.; Infante Abreu, M.B.; Parv, A.L. Evolution and Key Differences in Maturity Models for Industrial Digital Transformation: Focus on Industry 4.0 and 5.0. Sustainability 2025, 17, 11042. [Google Scholar] [CrossRef] [Scilit]
  18. Bird, S.; Klein, E.; Loper, E. Natural Language Processing with Python; O’Reilly Media: Sebastopol, CA, USA, 2009. [Google Scholar]
  19. Explosion. spaCy: Industrial-Strength Natural Language Processing (NLP), Spanish Medium-Sized Model (es_core_news_md); Open-Source Software for Advanced NLP in Python and Cython. Available online: https://spacy.io/ (accessed on 24 October 2025).
  20. Microsoft. Power Map for Microsoft Excel: A 3D Geospatial Data Visualization Tool Integrated in Excel as Part of Microsoft 365 and Business Intelligence Features. Available online: https://support.microsoft.com/office/get-started-with-power-map-88a28df6-8258-40aa-b5cc-577873fb0f4a (accessed on 24 October 2025).
  21. Aria, M.; Cuccurullo, C. bibliometrix: An R-tool for comprehensive science mapping analysis. J. Inf. 2017, 11, 959–975. [Google Scholar] [CrossRef] [Scilit]
  22. The Pandas Development Team. Pandas: Powerful Python Library for Data Structures and Data Analysis (Version 2.x); Open-Source Software Under the BSD License. Available online: https://pandas.pydata.org/docs/ (accessed on 24 October 2025).
  23. The SciPy Community. SciPy: Open-Source Python Library for Scientific Computing, Including Statistical Functions (e.g., scipy.stats.chi2_contingency). Available online: https://docs.scipy.org/doc/scipy/reference/ (accessed on 24 October 2025).
  24. Çinar, Z.M.; Zeeshan, Q.; Korhan, O. A Framework for Industry 4.0 Readiness and Maturity of Smart Manufacturing Enterprises: A Case Study. Sustainability 2021, 13, 32. [Google Scholar] [CrossRef] [Scilit]
  25. Liu, Y.; Feng, J.; Lu, J.; Zhou, S. A review of digital twin capabilities, technologies, and applications based on the maturity model. Adv. Eng. Inform. 2024, 62, 102592. [Google Scholar] [CrossRef] [Scilit]
  26. Abner, B.; Rabelo, R.J.; Zambiasi, S.P.; Romero, D. Production Management as-a-Service: A Softbot Approach. In IFIP Advances in Information and Communication Technology; Lalic, B., Marjanovic, U., Majstorovic, V., von Cieminski, G., Romero, D., Eds.; Springer: Berlin/Heidelberg, Germany; Volume 592 IFIP, pp. 19–30. [CrossRef] [Scilit]
  27. Alcácer, V.; Rodrigues, J.; Carvalho, H.; Cruz-Machado, V. Industry 4.0 maturity follow-up inside an internal value chain: A case study. Int. J. Adv. Manuf. Technol. 2022, 119, 5035–5046. [Google Scholar] [CrossRef] [Scilit]
  28. Amaral, A.; Peças, P. A Framework for Assessing Manufacturing SMEs Industry 4.0 Maturity. Appl. Sci. 2021, 11, 17. [Google Scholar] [CrossRef] [Scilit]
  29. Angreani, L.S.; Vijaya, A.; Wicaksono, H. Enhancing strategy for Industry 4.0 implementation through maturity models and standard reference architectures alignment. J. Manuf. Technol. Manag. 2024, 35, 848–873. [Google Scholar] [CrossRef] [Scilit]
  30. Bajic, B.; Moraca, S.; Rikalovic, A. Fuzzy maturity model for Smart Manufacturing Readiness: Industry 5.0 perspective. In 2023 IEEE Zooming Innovation in Consumer Technologies Conference, ZINC 2023, Novi Sad, Serbia, 29–31 May 2023; Institute of Electrical and Electronics Engineers Inc.: New York, NY, USA, 2023; pp. 142–147. [Google Scholar] [CrossRef] [Scilit]
  31. Barbalho, S.C.M.; Dantas, R.F. The effect of islands of improvement on the maturity models for industry 4.0: The implementation of an inventory management system in a beverage factory1,2. Braz. J. Oper. Prod. Manag. 2021, 18, 1–17. [Google Scholar] [CrossRef] [Scilit]
  32. Benmimoun, R.; El Kihel, Y.; Embarki, S.; El Kihel, B. Towards a Warehouse 4.0: Proposal of a Maturity Model for SMEs. In 2024 4th International Conference on Innovative Research in Applied Science, Engineering and Technology, IRASET 2024, Fez, Morocco, 16–17 May 2024; Benhala, B., Raihani, A., Qbadou, M., Eds.; Institute of Electrical and Electronics Engineers Inc.: New York, NY, USA, 2024. [Google Scholar] [CrossRef] [Scilit]
  33. Bernhard, O.; Zaeh, M.F. A Concept For The Development Of A Maturity Model For The Holistic Assessment Of Lean, Digital, And Sustainable Production Systems. In Proceedings of the Conference on Production Systems and Logistics; Herberger, D., Hubner, M., Eds.; Publish-Ing in cooperation with TIB—Leibniz Information Centre for Science and Technology University Library: Hanover, Germany, 2023; pp. 837–847. [Google Scholar] [CrossRef]
  34. Bohorquez, J.H.A.; Gil-Herrera, R.D. Proposal and Validation of an Industry 4.0 Maturity Model for SMEs. J. Ind. Eng.-Manag. Jiem 2022, 15, 433–454. [Google Scholar] [CrossRef] [Scilit]
  35. Bretz, L.; Klinkner, F.; Kandler, M.; Shun, Y.; Lanza, G. The ECO Maturity Model - A human-centered Industry 4.0 maturity model. Procedia CIRP 2022, 106, 90–95. [Google Scholar] [CrossRef] [Scilit]
  36. Caiado, R.G.G.; Scavarda, L.F.; Gavião, L.O.; Ivson, P.; Nascimento, D.L.D.M.; Garza-Reyes, J.A. A fuzzy rule-based industry 4.0 maturity model for operations and supply chain management. Int. J. Prod. Econ. 2021, 231. [Google Scholar] [CrossRef] [Scilit]
  37. Castelo-Branco, I.; Oliveira, T.; Simoes-Coelho, P.; Portugal, J.; Filipe, I. Measuring the fourth industrial revolution through the Industry 4.0 lens: The relevance of resources, capabilities and the value chain. Comput. Ind. 2022, 138, 16. [Google Scholar] [CrossRef] [Scilit]
  38. Chaopaisarn, P.; Woschank, M. Maturity Model Assessment of SMART Logistics for SMEs. Chiang Mai Univ. J. Nat. Sci. 2021, 20, 1–8. [Google Scholar] [CrossRef] [Scilit]
  39. Chaves Franz, M.L.; Ayala, N.F.; Larranaga, A.M. Industry 4.0 for passenger railway companies: A maturity model proposal for technology management. J. Rail Transp. Plan. Manag. 2024, 32. [Google Scholar] [CrossRef] [Scilit]
  40. Ciravegna-Martins-da fonseca, L.M.; Pereira, T.; Oliveira, M.; Ferreira, F.; Busu, M. Manufacturing Companies Industry 4.0 Maturity Perception Level: A Multivariate Analysis. J. Ind. Eng. Manag. 2024, 17, 196–216. [Google Scholar] [CrossRef] [Scilit]
  41. Das, P.; Perera, S.; Senaratne, S.; Osei-Kyei, R. Industry 4.0 Maturity of General Contractors: An In-Depth Case Study Analysis. Buildings 2024, 14, 18. [Google Scholar] [CrossRef] [Scilit]
  42. Demir, S.; Gunduz, M.A.; Kayikci, Y.; Paksoy, T. Readiness and Maturity of Smart and Sustainable Supply Chains: A Model Proposal. EMJ—Eng. Manag. J. 2023, 35, 181–206. [Google Scholar] [CrossRef] [Scilit]
  43. Elibal, K.; Özceylan, E. An Industry 4.0 Maturity Model Proposal Based on Total Quality Management Principles: An Application to an Automotive Parts Manufacturer. IEEE Trans. Eng. Manag. 2024, 71, 10815–10832. [Google Scholar] [CrossRef] [Scilit]
  44. Ericson Öberg, A.; Goncalves Machado, C.; Stålberg, L. Diagnostics of Opportunities—A Dialogue Tool for Addressing Digital Factory Maturity. In Advances in Transdisciplinary Engineering; Andersson, J., Joshi, S., Malmskold, L., Hanning, F., Eds.; IOS Press BV: Amsterdam, The Netherlands, 2024; Volume 52, pp. 395–406. [Google Scholar] [CrossRef] [Scilit]
  45. Facchini, F.; Olesków-Szlapka, J.; Ranieri, L.; Urbinati, A. A Maturity Model for Logistics 4.0: An Empirical Analysis and a Roadmap for Future Research. Sustainability 2020, 12, 18. [Google Scholar] [CrossRef] [Scilit]
  46. Felippes, B.; da Silva, I.; Barbalho, S.; Adam, T.; Heine, I.; Schmitt, R. 3D-CUBE readiness model for industry 4.0: Technological, organizational, and process maturity enablers. Prod. Manuf. Res. 2022, 10, 875–937. [Google Scholar] [CrossRef] [Scilit]
  47. Ferreira, D.V.; de Gusmão, A.P.H.; de Almeida, J.A. A multicriteria model for assessing maturity in industry 4.0 context. J. Ind. Inf. Integr. 2024, 38. [Google Scholar] [CrossRef] [Scilit]
  48. Gökalp, M.O.; Gökalp, E.; Kayabay, K.; Koçyiğit, A.; Eren, P.E. Data-driven manufacturing: An assessment model for data science maturity. J. Manuf. Syst. 2021, 60, 527–546. [Google Scholar] [CrossRef] [Scilit]
  49. Gomes, A.D.O.; Basilio, J.C. A Fuzzy Inference Model to Identify the Current Industry Maturity Stage in the Transformation Process to Industry 4.0. IEEE Trans. Autom. Sci. Eng. 2024, 21, 1607–1622. [Google Scholar] [CrossRef] [Scilit]
  50. Heinz, D.; Benz, C.; Silbernagel, R.; Molins, B.; Satzger, G.; Lanza, G. A Maturity Model for Smart Product-Service Systems. Procedia CIRP 2022, 107, 113–118. [Google Scholar] [CrossRef] [Scilit]
  51. Herceg, I.V.; Kuc, V.; Mijuskovic, V.M.; Herceg, T. Challenges and Driving Forces for Industry 4.0 Implementation. Sustainability 2020, 12, 22. [Google Scholar] [CrossRef] [Scilit]
  52. Jamouli, Y.; Tetouani, S.; Cherkaoui, O.; Soulhi, A. To diagnose industry 4.0 by maturity model: The case of moroccan clothing industry. Data Metadata 2023, 2. [Google Scholar] [CrossRef] [Scilit]
  53. Kieroth, A.; Brunner, M.; Bachmann, N.; Jodlbauer, H.; Kurz, W. Investigation on the acceptance of an Industry 4.0 maturity model and improvement possibilities. Procedia Comput. Sci. 2022, 200, 428–437. [Google Scholar] [CrossRef] [Scilit]
  54. Kırmızı, M.; Kocaoglu, B. Digital transformation maturity model development framework based on design science: Case studies in manufacturing industry. J. Manuf. Technol. Manag. 2022, 33, 1319–1346. [Google Scholar] [CrossRef] [Scilit]
  55. Koldewey, C.; Hobscheidt, D.; Pierenkemper, C.; Kühn, A.; Dumitrescu, R. Increasing Firm Performance through Industry 4.0—A Method to Define and Reach Meaningful Goals. Sci 2022, 4, 39. [Google Scholar] [CrossRef] [Scilit]
  56. Lin, T.C.; Sheng, M.L.; Wang, K.J. Dynamic capabilities for smart manufacturing transformation by manufacturing enterprises. Asian J. Technol. Innov. 2020, 28, 403–426. [Google Scholar] [CrossRef] [Scilit]
  57. Lin, T.C.; Wang, K.J.; Sheng, M.L. To assess smart manufacturing readiness by maturity model: A case study on Taiwan enterprises. Int. J. Comput. Integr. Manuf. 2020, 33, 102–115. [Google Scholar] [CrossRef] [Scilit]
  58. Lookman, K.; Pujawan, N.; Nadlifatin, R. Measuring innovative capability maturity model of trucking companies in Indonesia. Cogent Bus. Manag. 2022, 9, 2094854. [Google Scholar] [CrossRef] [Scilit]
  59. Lukhmanov, Y.; Dikhanbayeva, D.; Yertayev, B.; Shehab, E.; Turkyilmaz, A. An advisory system to support Industry 4.0 readiness improvement. Procedia CIRP 2022, 107, 1361–1366. [Google Scholar] [CrossRef] [Scilit]
  60. Madhavan, M.; Sharafuddin, M.A.; Wangtueai, S. Measuring the Industry 5.0-Readiness Level of SMEs Using Industry 1.0-5.0 Practices: The Case of the Seafood Processing Industry. Sustainability 2024, 16, 19. [Google Scholar] [CrossRef] [Scilit]
  61. Magdalena, L.; Isnanto, R.R.; Wibowo, A.; Warsito, B. Decision Making To Assess The Maturity Dimensions of MSME Using A Data Analysis Approach. In 2021 5th International Conference on Informatics and Computational Sciences (ICICoS), Semarang, Indonesia, 24–25 November 2021; Institute of Electrical and Electronics Engineers Inc.: New York, NY, USA, 2021; pp. 6–11. [Google Scholar] [CrossRef] [Scilit]
  62. Magnus, C.S. Smart factory mapping and design: Methodological approaches. Prod. Eng. 2023, 17, 753–762. [Google Scholar] [CrossRef] [Scilit]
  63. Maier, H.T.; Schmiedbauer, O.; Biedermann, H. Validation of a Lean Smart Maintenance Maturity Model. Teh. Glas.-Tech. J. 2020, 14, 296–302. [Google Scholar] [CrossRef] [Scilit]
  64. Maisiri, W.; van Dyk, L.; Coetzee, R. Development of an Industry 4.0 Competency Maturity Model. Saiee Afr. Res. J. 2021, 112, 189–197. [Google Scholar]
  65. Melnik, S.; Magnotti, M.; Butts, C.; Putman, C.; Aqlan, F. Developing a maturity model and an implementation plan for industry 4.0 integration. In International Conference on Industrial Engineering and Operations Management; IEOM Society: Southfield, MI, USA, 2024; Volume 59, pp. 1695–1707. [Google Scholar]
  66. Mo, F.; Monetti, F.M.; Torayev, A.; Rehman, H.U.; Mulet Alberola, J.A.; Rea Minango, N.; Nguyen, H.N.; Maffei, A.; Chaplin, J.C. A maturity model for the autonomy of manufacturing systems. Int. J. Adv. Manuf. Technol. 2023, 126, 405–428. [Google Scholar] [CrossRef] [Scilit]
  67. Nausch, M.; Schumacher, A.; Sihn, W. Assessment of Organizational Capability for Data Utilization—A Readiness Model in the Context of Industry 4.0. In Lecture Notes in Mechanical Engineering; Durakbasa, N.M., Osman Zahid, M.N., Abd. Aziz, R., Yusoff, A.R., Mat Yahya, N., Abdul Aziz, F., Yazid Abu, M., Gençyilmaz, M.G., Eds.; Springer Science and Business Media Deutschland GmbH: Berlin/Heidelberg, Germany, 2019; pp. 243–252. [Google Scholar] [CrossRef] [Scilit]
  68. Nick, G.; Ko, A.; Szaller, Á.; Zeleny, K.; Kádár, B.; Kovács, T. Extension of the CCMS 2.0 maturity model towards Artificial Intelligence. IFAC-PapersOnLine 2022, 55, 293–298. [Google Scholar] [CrossRef] [Scilit]
  69. Nick, G.; Kovács, T.; Ko, A.; Kádár, B. Industry 4.0 readiness in manufacturing: Company Compass 2.0, a renewed framework and solution for Industry 4.0 maturity assessment. In Proceedings of the 10th CIRP Conference on Digital Enterprise Technologies (DET)—Digital Technologies as Enablers of Industrial Competitiveness and Sustainability, Amsterdam, The Netherlands, 11–13 October 2021; Volume 54, pp. 39–44. [Google Scholar] [CrossRef] [Scilit]
  70. Nick, G.; Szaller, Á.; Várgedo, T. CCMS Model: A novel approach to digitalization level assessment for manufacturing companies. In Proceedings of the 16th European Conference on Management Leadership and Governance, ECMLG 2020, Academic Conferences International, Online, 26–27 October 2020; Griffiths, P., Ed.; pp. 195–203. [Google Scholar] [CrossRef] [Scilit]
  71. Nick, G.; Zeleny, K.; Kovács, T.; Járvás, T.; Pocsarovszky, K.; Ko, A. Artificial intelligence enriched industry 4.0 readiness in manufacturing: The extended CCMS2.0e maturity model. Prod. Manuf. Res. 2024, 12, 23. [Google Scholar] [CrossRef] [Scilit]
  72. P. Senna, P.; Barros, A.C.; Bonnin Roca, J.; Azevedo, A. Development of a digital maturity model for Industry 4.0 based on the technology-organization-environment framework. Comput. Ind. Eng. 2023, 185. [Google Scholar] [CrossRef] [Scilit]
  73. Pan Nogueras, M.L.; Perea Muñoz, L.; Cosentino, J.P.; Suarez Anzorena, D. RAISE 4.0: A Readiness Assessment Instrument Aimed at Raising SMEs to Industry 4.0 Starting Levels—An Empirical Field Study. In Lecture Notes in Mechanical Engineering; Andersen, A.L., Andersen, R., Brunoe, T.D., Larsen, M.S.S., Nielsen, K., Napoleone, A., Kjeldgaard, S., Eds.; Springer Science and Business Media Deutschland GmbH: Berlin/Heidelberg, Germany, 2021; pp. 713–720. [Google Scholar] [CrossRef] [Scilit]
  74. Pech, M.; Vrchota, J. Classification of small-and medium-sized enterprises based on the level of industry 4.0 implementation. Appl. Sci. 2020, 10. [Google Scholar] [CrossRef] [Scilit]
  75. Peukert, S.; Treber, S.; Balz, S.; Haefner, B.; Lanza, G. Process model for the successful implementation and demonstration of SME-based industry 4.0 showcases in global production networks. Prod. Eng. 2020, 14, 275–288. [Google Scholar] [CrossRef] [Scilit]
  76. Rafael, L.D.; Jaione, G.E.; Cristina, L.; Ibon, S.L. An Industry 4.0 maturity model for machine tool companies. Technol. Forecast. Soc. Change 2020, 159, 13. [Google Scholar] [CrossRef] [Scilit]
  77. Rahamaddulla, S.R.B.; Leman, Z.; Baharudin, B.T.H.T.B.; Ahmad, S.A. Conceptualizing smart manufacturing readiness-maturity model for small and medium enterprise (Sme) in malaysia. Sustainability 2021, 13, 9793. [Google Scholar] [CrossRef] [Scilit]
  78. Saad, S.M.; Bahadori, R.; Jafarnejad, H. The smart SME technology readiness assessment methodology in the context of industry 4.0. J. Manuf. Technol. Manag. 2021, 32, 1037–1065. [Google Scholar] [CrossRef] [Scilit]
  79. Sajjad, A.; Ahmad, W.; Hussain, S.; Chuddher, B.A.; Sajid, M.; Jahanjaib, M.; Ali, M.K.; Jawad, M. Assessment by Lean Modified Manufacturing Maturity Model for Industry 4.0: A Case Study of Pakistan’s Manufacturing Sector. IEEE Trans. Eng. Manag. 2024, 71, 6420–6434. [Google Scholar] [CrossRef] [Scilit]
  80. Santos, R.C.; Martinho, J.L. An Industry 4.0 maturity model proposal. J. Manuf. Technol. Manag. 2020, 31, 1023–1043. [Google Scholar] [CrossRef] [Scilit]
  81. Schabany, D.; Hülsmann, T.H.; Schmetz, A. Development of a Maturity Assessment Model for Digital Twins in Battery Cell Industry. Procedia CIRP 2023, 120, 946–951. [Google Scholar] [CrossRef] [Scilit]
  82. Schroderus, J.; Lasrado, L.A.; Menon, K.; Kärkkäinen, H. Towards a Pay-Per-X Maturity Model for Equipment Manufacturing Companies. Procedia Comput. Sci. 2021, 196, 226–234. [Google Scholar] [CrossRef] [Scilit]
  83. Semeraro, C.; Alyousuf, N.; Kedir, N.I.; Lail, E.A. A maturity model for evaluating the impact of Industry 4.0 technologies and principles in SMEs. Manuf. Lett. 2023, 37, 61–65. [Google Scholar] [CrossRef] [Scilit]
  84. Simetinger, F.; Basl, J. A pilot study: An assessment of manufacturing SMEs using a new Industry 4.0 Maturity Model for Manufacturing Small- and Middle-sized Enterprises (I4MMSME). Procedia Comput. Sci. 2022, 200, 1068–1077. [Google Scholar] [CrossRef] [Scilit]
  85. Steinlechner, M.; Schumacher, A.; Fuchs, B.; Reichsthaler, L.; Schlund, S. A maturity model to assess digital employee competencies in industrial enterprises. Procedia CIRP 2021, 104, 1185–1190. [Google Scholar] [CrossRef] [Scilit]
  86. Suleiman, Z.; Dikhanbayeva, D.; Shaikholla, S.; Turkyilmaz, A. Readiness Assessment of SMEs in Transitional Economies: Introduction of Industry 4.0. In Proceedings of the ACM International Conference Proceeding Series, Barcelona, Spain, 8–11 January 2021; Association for Computing Machinery: New York, NY, USA, 2021; pp. 8–13. [Google Scholar] [CrossRef] [Scilit]
  87. Treviño-Elizondo, B.L.; García-Reyes, H. The challenge of becoming a worker 4.0—A human-centered maturity model for industry 4.0 adoption. IISE Annu. Conf. Expo 2021 2021, 2021, 584–589. [Google Scholar]
  88. Treviño-Elizondo, B.L.; García-Reyes, H. An Employee Competency Development Maturity Model for Industry 4.0 Adoption. Sustainability 2023, 15, 29. [Google Scholar] [CrossRef] [Scilit]
  89. Treviño-Elizondo, B.L.; García-Reyes, H.; Peimbert-García, R.E. A Maturity Model to Become a Smart Organization Based on Lean and Industry 4.0 Synergy. Sustainability 2023, 15, 24. [Google Scholar] [CrossRef] [Scilit]
  90. Uhlenkamp, J.F.; Hauge, J.B.; Broda, E.; Lutjen, M.; Freitag, M.; Thoben, K.D. Digital Twins: A Maturity Model for Their Classification and Evaluation. IEEE Access 2022, 10, 69605–69635. [Google Scholar] [CrossRef] [Scilit]
  91. Ünal, C.; Sungur, C.; Yildirim, H. Application of the Maturity Model in Industrial Corporations. Sustainability 2022, 14, 25. [Google Scholar] [CrossRef] [Scilit]
  92. Wagire, A.A.; Joshi, R.; Rathore, A.P.S.; Jain, R. Development of maturity model for assessing the implementation of Industry 4.0: Learning from theory and practice. Prod. Plan. Control 2021, 32, 603–622. [Google Scholar] [CrossRef] [Scilit]
  93. Widmer, N.; Hassan, A.; Monticolo, D. Assessment Model to Support the Technological Integration within Industrial Companies in the Context of Industry 4.0. In 2022 IEEE 28th International Conference on Engineering, Technology and Innovation, ICE/ITMC 2022 and 31st International Association for Management of Technology, IAMOT 2022 Joint Conference—Proceedings, Nancy, France, 19–23 June 2022; Institute of Electrical and Electronics Engineers Inc.: New York, NY, USA, 2022. [Google Scholar] [CrossRef] [Scilit]
  94. Zoubek, M.; Simon, M. A Framework for a Logistics 4.0 maturity model with a specification for internal logistics. Mm Sci. J. 2021, 2021, 4264–4274. [Google Scholar] [CrossRef] [Scilit]
  95. Zoubek, M.; Poor, P.; Broum, T.; Basl, J.; Simon, M. Industry 4.0 Maturity Model Assessing Environmental Attributes of Manufacturing Company. Appl. Sci. 2021, 11, 24. [Google Scholar] [CrossRef] [Scilit]
  96. Trstenjak, M.; Opetuk, T.; Dukic, G.; Cajner, H. Logistics 5.0 Implementation Model Based on Decision Support Systems. Sustainability 2022, 14, 19. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Overview of the five-phase review method. Each rounded rectangle represents a collapsed sub-process; the plus symbol indicates that more detailed activities are described in the corresponding subsection.
Figure 1. Overview of the five-phase review method. Each rounded rectangle represents a collapsed sub-process; the plus symbol indicates that more detailed activities are described in the corresponding subsection.
Systems 14 00134 g001
Figure 2. Document Retrieval Process. Solid boxes represent the methodological steps, while dotted boxes indicate the number of records excluded or retained at each stage [17].
Figure 2. Document Retrieval Process. Solid boxes represent the methodological steps, while dotted boxes indicate the number of records excluded or retained at each stage [17].
Systems 14 00134 g002
Figure 3. Data Extraction, Normalization, and Coding workflow. Solid arrows represent the main sequential workflow, while dotted arrows indicate feedback loops in which the master dataset is updated or consulted during intermediate and final stages [17].
Figure 3. Data Extraction, Normalization, and Coding workflow. Solid arrows represent the main sequential workflow, while dotted arrows indicate feedback loops in which the master dataset is updated or consulted during intermediate and final stages [17].
Systems 14 00134 g003
Figure 4. Geographic origin of maturity models: single-country studies, international collaborations, and global frameworks.
Figure 4. Geographic origin of maturity models: single-country studies, international collaborations, and global frameworks.
Systems 14 00134 g004
Figure 5. Geographical distribution of the origin of the analysed maturity models (frequency by country).
Figure 5. Geographical distribution of the origin of the analysed maturity models (frequency by country).
Systems 14 00134 g005
Figure 6. Global map of origin and collaboration networks in maturity-model development: countries shaded by model count and lines indicating international co-authorships.
Figure 6. Global map of origin and collaboration networks in maturity-model development: countries shaded by model count and lines indicating international co-authorships.
Systems 14 00134 g006
Figure 7. Most frequent target sectors among sector-specific maturity models. Subsector-level breakdown for descriptive purposes only; inferential analysis is conducted on normalized sector categories.
Figure 7. Most frequent target sectors among sector-specific maturity models. Subsector-level breakdown for descriptive purposes only; inferential analysis is conducted on normalized sector categories.
Systems 14 00134 g007
Figure 8. Distribution of Types of Research Gaps Identified in Maturity Models (2020–2024).
Figure 8. Distribution of Types of Research Gaps Identified in Maturity Models (2020–2024).
Systems 14 00134 g008
Table 1. Research Questions.
Table 1. Research Questions.
RQResearch Question Addressed
RQ1Where is knowledge on 4.0/5.0 maturity models being produced, and how well are developing countries represented?
RQ2Does the type of model origin (single-country, collaborative, or global) influence whether authors consider its applicability in developing countries?
RQ3Does the development level of the country of origin affect the likelihood that the model explicitly addresses developing countries?
RQ4To what extent are existing models empirically validated, in which sectors, and how does this limit their transferability to emerging economies?
RQ5What research gaps do authors themselves identify as barriers to 4.0/5.0 adoption in developing countries?
Table 2. Selected themes.
Table 2. Selected themes.
Web of ScienceScopus
Engineering Industrial; Engineering Manufacturing; Management; Engineering Electrical Electronic; Computer Science Information Systems; Computer Science Interdisciplinary Applications; Operations Research Management Science; Green Sustainable Science Technology; Computer Science Theory Methods; Environmental Sciences; Computer Science Artificial Intelligence; Business; Telecommunications; Engineering Multidisciplinary; Environmental Studies; Automation Control Systems; Computer Science Software Engineering; Robotics; Economics; Ergonomics; Industrial Relations LaborEngineering; Computer Science; Environmental Science; Energy; Economics, Econometrics and Finance; Multidisciplinary
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Reyes Domínguez, D.; Infante Abreu, M.B.; Parv, A.L. Industry 4.0/5.0 Maturity Models: Empirical Validation, Sectoral Scope, and Applicability to Emerging Economies. Systems 2026, 14, 134. https://doi.org/10.3390/systems14020134

AMA Style

Reyes Domínguez D, Infante Abreu MB, Parv AL. Industry 4.0/5.0 Maturity Models: Empirical Validation, Sectoral Scope, and Applicability to Emerging Economies. Systems. 2026; 14(2):134. https://doi.org/10.3390/systems14020134

Chicago/Turabian Style

Reyes Domínguez, Dayron, Marta Beatriz Infante Abreu, and Aurica Luminita Parv. 2026. "Industry 4.0/5.0 Maturity Models: Empirical Validation, Sectoral Scope, and Applicability to Emerging Economies" Systems 14, no. 2: 134. https://doi.org/10.3390/systems14020134

APA Style

Reyes Domínguez, D., Infante Abreu, M. B., & Parv, A. L. (2026). Industry 4.0/5.0 Maturity Models: Empirical Validation, Sectoral Scope, and Applicability to Emerging Economies. Systems, 14(2), 134. https://doi.org/10.3390/systems14020134

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop