Skip to Content
  • Article
  • Open Access

18 April 2026

MIDA—Method for Industrial Data Analysis Based on CRISP-DM

and
1
Coimbra Institute of Engineering, Polytechnic University of Coimbra, Rua Pedro Nunes, 3030-199 Coimbra, Portugal
2
RCM2+ Research Centre for Asset Management and Systems Engineering, Rua Pedro Nunes, 3030-199 Coimbra, Portugal
*
Author to whom correspondence should be addressed.
This article belongs to the Section Data

Abstract

As modern computers became increasingly more popular and larger amounts of digital data were available, different methodologies were proposed to extract information from data. CRISP-DM methodology quickly spread and is currently one of the most popular approaches used for data analysis. However, it has some shortcomings, such as being too general or business-centered. Different authors have proposed variations more suitable to specific fields in order to overcome those limitations. The present paper reviews CRISP-DM, some variations and similar methodologies, and proposes a Methodology for Industrial Data Analysis (MIDA)—a methodology conceived and improved over time, based on previous experience in industrial engineering processes. MIDA consists of eight steps and partially overlaps with CRISP-DM. It has been successfully applied in several previous projects.

1. Introduction

CRoss-Industry Standard Process for Data Mining (CRISP-DM) is one of the most popular methodologies used in data mining and analysis processes. The first references to the methodology date back to 1996. It was developed as part of European CORDIS grants 24959 (https://cordis.europa.eu/project/id/24959)/25959 (https://cordis.europa.eu/project/id/25959) (accessed on 31 January 2026), from 1997-07-01 to 1998-12-31 (dates on ISO format YYYY-MM-DD). CRISP-DM was officially presented in 1999. The official document is publicly available online (https://public.dhe.ibm.com/software/analytics/spss/documentation/modeler/14.2/es/CRISP-DM.pdf) (accessed on 10 January 2026) [1].
CRISP-DM is a six-step methodology. The steps are also referred in the literature as phases. It was conceived to be used in projects where the aim is to extract information from data, and produce models which can be deployed to solve industrial problems. Due to its simple yet powerful structure, CRISP-DM became quickly popular among data scientists, analysts, engineers and other professionals and researchers involved in data analysis projects.
CRISP-DM is a general framework, designed to be applied across different industries. Its level of abstraction is simultaneously its power and its main drawback. It is its power because the methodology can and has been successfully used in projects in countless different areas. As long as the steps are correctly followed, the chances of success of the project are very high. But it is also its main drawback because it loses the specificity required for particular projects and fields. In some cases this shortcoming is so important that different authors have proposed variations of the methodology, more adequate to comply with the requirements of particular areas. A comprehensive number of variations are reviewed in Section 3.4.
Physical Asset Management (PAM) is one key area for large industries and institutions, important for optimizing resources and production. Its main goal is to minimize costs of purchasing and ownership of the assets, while maximizing their production and availability. This is a very specific field, involving management, financial data, and equipment monitoring and maintenance. Decision-making is increasingly more data-driven, as in many other areas of industry. Parts of the data used are finance-, management- and business-related. Most of the times those data can be modeled as time series with very low sampling rate. Other data are normally sensory inputs that come from the assets themselves, reporting their condition, usage and other functioning aspects. Those data are normally time series of variable length and sampling rates.
Whilst the original CRISP-DM methodology can and has been used in PAM projects, its abstract and general structure do cause significant embarrassments in practice, namely when it is necessary to perform all the processes necessary for data-driven decisions, from requirement analysis to model deployment and maintenance.
This paper proposes a Methodology for Industrial Data Analysis (MIDA), a variation of CRISP-DM that aims to include the most frequent steps necessary for the data collection and analysis explicitly in the project, along with other minor variations that make it more suitable for asset management projects. Nonetheless, the methodology is general enough to be useful in other fields too.

2. Original CRISP-DM

2.1. The 6-Step Methodology

Figure 1 shows the six steps of the original CRISP-DM methodology. They are:
Figure 1. Diagram of the steps of the original CRISP-DM methodology, as proposed in [1].
  • Business Understanding—Typically seen as the step where the problem is studied, the objectives of the project are defined and a tentative strategy is outlined.
  • Data understanding—This is the step where data is collected, if not already available, and explored in order to get the first insights both on the data quality and the information it contains.
  • Data preparation—Data are cleaned, transformed and prepared for the modeling phase.
  • Modeling—This is often the step that adds more value to the project, even though sometimes it only lasts a small fraction of the project duration. During this step, statistical as well as classical or modern machine learning models are tested and trained to model the data to get deeper insights or create models that could be deployed in practice.
  • Evaluation—Sometimes also referred to as “Results,” during this step there is an assessment of the quality of the model and how it suits the objectives defined in the Business Understanding step.
  • Deployment—This is the last step, when the models are deployed to production. The outputs of the models can be reports, dashboards, or other tools useful for decision-making.

2.2. Limitations of the Original Methodology

Even though the CRISP-DM methodology has been widely successful, it has quite a few shortcomings that hinder its application in different fields. Numerous gaps have been identified by different authors, as reviewed in Section 3. A quick summary of the main limitations is presented below.

2.2.1. Strong Focus on the Business Aspects

The first step is known as “Business Understanding.” It is where goals are defined, and that is done from a business perspective. In many projects the business perspective is not necessarily relevant during this first stage, and stressing this aspect may even cause a loss of focus of the project. For example, in a predictive maintenance project, the focus is usually on reducing equipment downtime and increasing availability. The business perspective is implicit. This aspect is very important in many different fields, as demonstrated by the number of methodology variations proposed that change the name of the step (summary provided in Table 1).
Table 1. Comparison of Methodologies. Acronyms are as follows: DM = Data Mining, DS = Data Science, BU = Business Understanding, DU = Data Understanding, DP = Data Preparation, MOD = Modeling, EVAL = Evaluation, DEP = Deployment.

2.2.2. Planning Step Is Implicit

No step of the traditional CRISP-DM methodology explicitly refers to the project planning; it is a task of the broader “Business Understanding” step. However, planning is crucial in a project of predictive maintenance or data analysis for data-driven physical asset management, just as in other fields.

2.2.3. Data Collection and Management Are Implicit

The original CRISP-DM methodology is also oblivious to the data collection steps in the six main steps. Data collection appears as a task of the broader “Data Understanding” step. Many authors point out that a more explicit reference in the methodology can underline and evidence the importance of this step in the whole project.
Additionally, nowadays the amounts of data generated can easily grow exponentially, posing important challenges to storage and processing systems. Hence, data governance and life cycle management strategies must be defined during the project.

2.2.4. Evaluation Step

Naming the analysis of the results simply “Evaluation” can be confusing. In fact, this step is normally when the results are obtained, analysed, and compared to the state of the art and the requirements. Then there is a decision whether to get back to one of the previous steps or to proceed ahead to the deployment and/or conclusion of the project. The evaluation is also focused on the results of the models and not a general balance of the project. In some alternative methodologies, this step is also called “Interpretation.”

2.2.5. Legal and Regulatory Issues Are Not Addressed

CRISP-DM, as well as other similar methodologies, were conceived in a time when there was still little concern about legal issues when accessing and using the data. In the meantime, however, the regulatory framework evolved in Europe and other continents. In the European Union, Directive 95/46/EC (https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=celex%3A31995L0046, accessed on 31 January 2026) of the European Parliament and of the Council of 24 October 1995 focuses on the protection of individuals with regard to the processing of personal data and on the free movement of such data. This directive was later replaced by Regulation (EU) 2016/679 (https://eur-lex.europa.eu/legal-content/en/TXT/?uri=CELEX%3A32016R0679, accessed on 31 January 2026), popularly known as the General Data Protection Regulation (GDPR). This regulation aims to protect personal data and must be taken into account when data is harvested, collected, stored and used for data analysis. More recently, the AI Act, Regulation (EU) 2024/1689 (https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai, accessed on 31 January 2026), poses rules and constraints on the processing of massive amounts of data using artificial intelligence methods. Those regulations and subsequent local and regional laws and rules must be taken into account in modern data analysis projects.

3. Literature Review

As soon as technology evolved and electronic data was available, after initial research on different knowledge discovery methodologies, some methodologies were proposed. The most notable ones are reviewed below.

3.1. CRISP-DM and Alternative Methodologies

Knowledge Discovery in Databases (KDD) methodology was formulated as early as 1989 [2], even though the landmark paper usually cited as clarifying the steps of the methodology was published in 1996 by Fayyad et al. [3]. Compared to CRISP-DM, KDD is oblivious to the business-side of the project and the deployment issues. Nonetheless, it is often cited as one important landmark in the history of the methodologies and to this day it is still one important alternative to CRISP-DM. Figure 2 illustrates the steps of the KDD methodology.
Figure 2. Diagram of the KDD methodology, as proposed in [3].
Sample, Explore, Modify, Model, Assess (SEMMA) was proposed as early as 1997 [4]. Compared to CRISP-DM, this methodology ignores the business and the deployment steps, just as KDD does. Figure 3 illustrates the SEMMA methodology. It is arguably one of the three most popular methodologies, along with KDD and CRISP-DM.
Figure 3. Diagram of the SEMMA methodology, as proposed in [4].
Following in the footsteps of KDD and SEMMA, CRISP-DM was originally proposed in a 1999 document [1] and quickly became the preferred methodology among key players in academia and industry. Wirth and Hipp present one of the first case studies of the successful application of the methodology [5]. Azevedo and Santos compare KDD, SEMMA and CRISP-DM methodologies, highlighting the differences and similarities between the three approaches [6].
Another methodology that gained some popularity was OSEMN—Obtain, Scrub, Explore, Model, iNterpret. According to online records (https://introdatasci.dlilab.com/pdf/A_Taxonomy_of_Data_Science.pdf) (accessed on 31 January 2026), it first appeared in 2010, but has received much less attention from the academic community, when compared to the previously mentioned three proposals. Shameti and Cico [7] compare OSEMN to CRISP-DM and conclude on the superiority of the latter. Nonetheless, OSEMN has been successfully in different data science projects. Dineva and Atasanova [8] used OSEMN in a data science project to process IoT data from beehives. Kumari et al. [9] also report the use of the methodology in a sentiment analysis project.
In 2021, Martínez-Plumed et al. [10] affirm that CRISP-DM is still the de facto standard methodology for data mining and data science projects. This is observed, even though the methods and algorithms evolved over time. Additionally, the term “data science” slowly replaced the old designation of “data mining,” more prevalent when CRISP-DM was proposed, for information and knowledge extraction projects. The authors also affirm that when data science projects become more exploratory, a more flexible model is preferred. In the same year, a study by Shröer et al. reviews the application of CRISP-DM in data science projects and comes to similar conclusions [11].
In 2024, Shimaoka et al. [12] published another survey where they affirm that 82% of the teams do not use any process model—however, CRISP-DM is still the preferred method due to its flexibility, robustness and adaptability to different domains. The authors also compare 16 variations of the methodology and analyze their differences and similarities.

3.2. Comparison of CRISP-DM and Alternatives

Zavaleta-Sánchez et al. [13] compare the use of KDD and CRISP-DM in a data mining project for phishing identification. They consider that CRISP-DM is more detailed, but KDD is more intuitive and can be followed more naturally, without awareness of following a formal methodology. The conclusions are understandable, considering that KDD steps are very specific, even though some common steps are missing (e.g., requirement analysis, planning objectives, exploratory data analysis). CRISP-DM entails those tasks, but under a very high level formulation, where problem understanding, requirement analysis and planning fall into the step of “Business Understanding” and exploratory data analysis falls into the step of “Data Understanding.”
Palacios et al. [14] compare the use of SEMMA and CRISP-DM on a project to construct a repository for studies of land use and cover change. They conclude that SEMMA is more adequate for projects when specific tools are used, while CRISP-DM provides a more general framework. The successful application of the latter, however, may require adherence to tasks/substeps of the main six steps.
Rosander [15] compares the use of CRISP-DM, SEMMA and KDD for signal processing projects. In the particular case, the signals were from radar warning systems. The conclusion was that CRISP-DM was the best methodology, because of its ease of implementation in an agile Scrum project.

3.3. Applications

CRISP-DM has been used in different areas, including industrial, engineering projects [16], teaching [17,18], paediatrics [19], and agriculture [20].
Bosnjak et al. [21] report on the use of the CRISP-DM methodology to analyse data from small and medium enterprises. They conclude that many of the difficulties of the process could have been overcome if the data collection process had been supervised by the data analysts.
Kannengiesser and Gero [22] compare CRISP-DM to models of design in different engineering fields and service design, using a Function–Behaviour–Structure framework. They conclude that CRISP-DM is the most solution-oriented methodology, rather than problem-oriented. Therefore, CRISP-DM could be improved by introducing some more attention to the problem in the methodology.
Saltz [23] analyses the use of CRISP-DM in data science projects and concludes that, while it is a robust and powerful method, it misses some key aspects of the project life cycle. Hence, the suggestion is to combine CRISP-DM with other popular project management approaches, such as Scrum.
Plotnikova et al. [24] use CRISP-DM in a financial data analysis project. They analyse the strengths and weaknesses of the methodology and identify a total of 18 shortcomings, grouped into three main categories: (i) interdependencies between the six steps; (ii) requirement analysis and validation; and (iii) universality of some of the steps.

3.4. Variations

Because of its flexibility and robustness, CRISP-DM became widely popular. Because of its limitations, some authors have proposed improvements or adaptations of the method for particular fields. Shimaoka et al. [12] present a thorough review of the data science methodologies, as well as other agile project management methodologies (namely Scrup and Kanban).

DMME

Huber et al. [25] propose Data Mining for Engineering Applications (DMME), as an extension of CRISP-DM to engineering applications. The authors also identify that CRISP-DM does not explicitly foresee a data acquisition phase within production scenarios. According to the authors, DMME provides a communication and planning foundation for data analytics. The key differences of DMME when compared to CRISP-DM are: (i) there are two steps added between Business Understanding and Data Understanding. Those steps are “Technical Understanding” and “Technical Realization;” and (ii) there is one more step between Evaluation and Deployment. That step is “Technical Implementation.” Those new steps may mitigate some of the shortcomings of the original methodology, but it is evident that their names are still quite general.
Cazacu and Titan [26] redefine some tasks of the CRISP-DM methodology for better fit in data science projects. Namely, Business Understanding and Data Understanding are the steps that suffer more changes. This is in line with observations from other authors, who also point out that the first steps of CRISP-DM are the ones that exhibit more gaps.
Venter et al. [27] propose a variation of CRISP-DM for evidence mining in forensic data. The variation is called CRISP-EM, where EM stands for “Evidence Mining.” The key differences to the original methodology are: (i) Business Understanding is called “Case Understanding;” (ii) Modeling is called “Evidence Modeling;” (iii) “Evaluation” is also called “Evidence Extraction;” and (iv) there is “Evidence Reporting” in place of Deployment.
Nagashima and Kato [28] propose APREP-DM as a framework for automatic pre-processing of sensory data based on CRISP-DM. The authors evaluate KDD, SEMMA and CRISP-DM, and propose a variation of the latter where data preparation includes specific preparation and cleaning steps, performed automatically.
Ayele [29] proposes CRISP-IM, an adaptation of CRISP-DM for idea mining. The original method was modified, so that instead of Business Understanding there is technology need assessment. Data Understanding is called “Data collection and understanding,” Modeling is renamed “Modeling for Idea Extraction,” Evaluation is called “Evaluation and Idea Extraction,” and Deployment is replaced by “Reporting Innovative Ideas.”
Bokrantz et al. [30] focus on the application of AI tools in manufacturing and propose an enhanced version of CRISP-DM, with an additional step of “Operations and Maintenance.” This new step is placed after Deployment, and therefore stands between the end of a project (or iteration) and the Business Understanding step at the beginning of a new project (or iteration).
Kristoffersen et al. [31] use the CRISP-DM methodology in the context of circular economy data analysis. The method is enhanced with an additional phase of “Data Validation” and integrates the concept of “analytic profiles.” The Data Validation step is introduced after the typical Data Preparation. The Analytic Profiles are used between Business Understanding and Data Understanding, in order to improve communication of knowledge/insights.
Acuña-Cid et al. [32] propose the integration of Network Analysis methods with CRISP-DM methodology. The resulting method is called CRISP-NET. The steps are the same as the original method. However, the subtasks of each step are strictly aligned with the network analysis methodology, aiming at better integration of the relationships within the systems.
Catley et al. [33] propose CRISP-TDM, a variation of the original method that aims at incorporating a temporal dimension into the process, specifically tailored for the case of medical data. Steps 1, 2, 4 and 6 of CRISP-DM are enriched with tasks such as clinical reporting.
Asamoah and Sharda [34] propose CRISP-eSNeP—Cross Industry Standard Process for Electronic Social Network Platforms. This variation is specially tailored to process social networks’ big data. The steps are very different from the original method: Cluster Development, Data Acquisition, Data Cleaning, Data Formatting, Data Validation, Data Analysis, and Deployment.
Ramos et al. [35] propose CRISP-EDM, a variation adapted to projects of data analysis in the educational domain. The steps are the same, even though renamed and with the tasks slightly adapted to the educational domain.
Niaksu [36] proposes an extension of CRISP-DM for the medical field. The modified method is called CRISP-MED-DM. In this proposal, the six steps of the original methodology are maintained; only the tasks are adapted to the particularities of the medical field.
Shafer et al. [37] propose QM-CRISP-DM, a variation that aims to adapt the methodology for quality management. The steps of the original CRISP-DM methodology are maintained; only the tasks are redefined to integrate quality management tools.

4. MIDA

4.1. Methodology Proposed

Figure 4 gives an overview of the proposed method. It shows the eight steps of the methodology and also the tasks of each step. Hence, it aims at being a reference guide of the MIDA methodology.
Figure 4. Diagram of the proposed methodology. Steps different from the original CRISP-DM are highlighted in color: yellow steps are significantly changed, green are new.
MIDA consists of eight steps, with a variable number of tasks. Some of the steps coincide with those of CRISP-DM, and others are variations or proposed as new. As in other methodologies, the steps and tasks are optional and the methodology is agile, flexible and iterative.
The steps and tasks are described in more detail below.
  • Problem Understanding—Inspired by the Business Understanding step of CRISP-DM, this is the first step of the MIDA methodology. Its name, however, aims to give the method a more problem-oriented focus, which is possibly more adequate in engineering and industrial settings. The tasks in this step are described below.
    1.1
    Background and Planning Outline—Most industrial projects start with an idea or need that is identified. The first task should be a quick check of the background need or idea that led to the project, in order to get an overview of the needs and feasibility of the project. Initial planning of the project should be done.
    1.2
    User Requirement Analysis—If the first background check is successful, the project should proceed to the second task, which is a more detailed analysis of the requirements. Adequate formal methodologies, such as use cases, can and must be used if needed.
    1.3
    Check of Legal Requirements—Legal implications should be checked in this step. Those include, for example, restrictions on the access to the data, imposed by copyright protections or regulations such as GDPR or the AI Act.
    1.4
    State of the Art—Once the user requirements are understood and the legal aspects are verified, a survey of the state of the art is recommended. This is important to verify the latest solutions available, from the technical and scientific points of view. This is fundamental to help establish realistic objectives and lay the ground for the Planning step.
    1.5
    Objectives and Success Criteria—Once the previous tasks are completed, clear objectives and success criteria should be defined.
    1.6
    Risks and Contingencies—A risk analysis and solutions to mitigate those risks should also be done, before proceeding to the next steps.
    Deliverable: Report—The output of this step should be a report, detailing the main findings of the Problem Understanding step, along with the conclusions, and proposing the roadmap for the next steps.
  • Planning—After Problem Understanding, detailed planning should be performed. Adjustments and refinements to the initial planning should be performed as needed.
    2.1
    Define Tasks and Deliverables—The work should be divided into adequate tasks. Deliverables should be defined.
    2.2
    Schedule—Tasks should be scheduled and milestones defined. Adequate tools, such as Gantt diagrams, should be used as needed.
    2.3
    Tools and Techniques—The methods and strategy to be used to achieve the goals should be defined. Materials necessary to the project should be listed and chosen. This includes data analysis software and other materials, if needed, for data collection, storage or deployment.
    Deliverable: Report—The output of this step should also be a report, which includes the planning as well as a list of methods and tools necessary for the successful completion of the project.
  • Data Collection—If the data necessary for the project is not immediately available, it should be collected in this step. The MIDA tasks to perform during this step are as follows.
    3.1
    Select Variables, Sensors and Logging Tools—Data collection could be achieved through different methods, depending on the type of problem. Sensors and support equipment, such as data loggers, may be necessary.
    3.2
    Harvest or Collect Data—The necessary, or possible, amounts of data should be harvested or collected in order to build the dataset necessary for the following steps.
    Deliverables: Dataset, Report, Documentation—The output of this step is a dataset, along with a report detailing the collection process, variable specifications and other details relevant for the analysis process. Other materials can be included, such as equipment manuals, datasheets or other technical documents.
  • Data Understanding—Once the dataset is available, the next step is its exploration, using adequate tools. In this step the first conclusions about the data and the processes in their origin are made. The tasks in this step are as follows.
    4.1
    Assess Data Quality—The first task is assessment of the quality of the data. Missing values, discrepant values and artifacts should be identified.
    4.2
    Exploratory Data Analysis—This task involves normally descriptive statistics, different visualizations and manipulations that lead to a better understanding of the equipment and/or processes that generated the data.
    4.3
    Data Validation—From the other tasks, conclusions can be taken about which parts of the dataset are relevant for the remainder of the project, which ones may need transformation, and those that must be discarded.
    Deliverables: Report, Visualizations—The output of this step should be a report with the statistical results, visualizations and other descriptions and conclusions about the quality of the data and the nature of the underlying processes.
  • Data Preparation—Once the data and the underlying processes are understood, data should be prepared for the modeling step. The MIDA tasks are the following.
    5.1
    Clean Data—Invalid data, such as discrepant samples, or irrelevant data, such as categories with insufficient data, should be removed from further processing. Variables can also be eliminated in this task based on different criteria.
    5.2
    Select and/or Augment Data—The important variables, or data chunks, should be selected for further steps. Data augmentation techniques could be applied, if needed, for balancing skewed datasets.
    5.3
    Transform Data for Modeling—Data should be transformed as needed for the models to be developed in the next step. This could involve normalization, standardization, application of sliding windows, filters and other transformations. Encoding of qualitative variables is also performed in this step.
    Deliverables: Report, Transformed Dataset—The output of this step should be a dataset prepared for the modeling step, along with an explanatory report.
  • Modeling—This step is the one where normally more value is added to the data analysis process. However, it is often one of the shortest ones. The MIDA tasks are described below.
    6.1
    Create Models—Most models are nowadays classical machine learning or deep learning models. Supervised and unsupervised machine learning models are often used. Statistical, numerical and other approaches can also be used in many processes.
    6.2
    Establish Evaluation Criteria—The models should be evaluated with adequate performance metrics. The metrics should be chosen according to the type of variables and models applied, as well as the objectives and acceptance criteria defined in the first step.
    6.3
    Experiment and Optimize Models—Models should be optimized according to the results of the evaluation criteria.
    Deliverables: Report, Model(s)—The output of this step should include the optimized models as well as an explanatory report.
  • Analysis of Results—This step includes thorough analysis of the best results, physical interpretation, scope, advantages and limitations of the approach. The MIDA tasks should be as follows.
    7.1
    Interpret the Results—The results must be interpreted to uncover the physical, processual and other possible implications. This often requires deep understanding of the tools, equipment and processes involved.
    7.2
    Compare Results to Objectives and State Of The Art—The results must be compared to the objectives of the project established in the first step, as well as the state of the art, for better framing of their quality, as well as advantages and limitations of the approach followed.
    7.3
    List possible actions and improvements—After interpretation of the results in a broad context, lines of action should be defined, namely aiming at the deployment step if the results are acceptable, or one of the previous steps of the methodology.
    Deliverable: Report—The output of this step is normally a report. This report can contain the conclusion of the project if there is no deployment step.
  • Deployment—Last step of the project. The implementation depends not only on the objectives of the project but also on the quality of the results obtained. The MIDA tasks are listed below.
    8.1
    Plan Deployment—Deployment of the models and/or results in production often requires specific and careful planning. In industrial settings, for example, it can impact the production lines and pose safety risks or economic risks.
    8.2
    Export Models—It is desirable that models are deployed in standard, or at least platform-independent, formats. This often requires the use of formats and platforms such as ONNX [38] for machine learning models.
    8.3
    Integrate with Existing Systems—Proper deployment of the new models often requires installation of software, as well as creation of pipelines for data inputs and channeling of model outputs.
    8.4
    System and Data Maintenance Plans—Models deployed in industrial settings, for example, require adequate monitoring and maintenance, to guarantee good working conditions and minimize downtime or chances of misbehaviour. Hence, careful planning of the monitoring and maintenance tasks required to maintain the system working well should be outlined before completion of the project. Additionally, the data generated may require life cycle planning and governance strategies, which should be outlined for future maintenance and use of the system.
    Deliverables: Report, Deployed Models—The output of this step should be a technical report on the deployment environment, including monitoring and maintenance plans.

4.2. Comparison of Methodologies

Table 1 presents a comparison of all the methodologies analysed and compares them with MIDA. The last column of the table stresses the main differences of each methodology compared to the others. MIDA stands out in the introduction of legal issues into the process, as well as the need for maintenance and data governance plans. Namely, it clarifies requirement analysis and planning phases in industrial settings. It also introduces legal analysis into the process, as well as maintenance and data governance plans. These aspects are implicit or neglected in other methodologies, but due to the latest advances of technology and regulations, they need to be incorporated in the data analysis projects. MIDA is, therefore, the only methodology that covers those needs.

5. Discussion

The MIDA methodology has been successfully applied in previous projects led by the authors’ team.

5.1. Previous Applications

Rodrigues et al. [39] used the methodology in a project where the goal was to predict bus motor oil condition using artificial neural networks. The data were collected from two different bus companies. The models were not deployed. The datasets in both companies consisted of tens of results of laboratory analysis.
Cruz et al. [40] also followed the methodology in a project where the goal was to automatically detect problems in heat-sealed bottles in a chemical industry. In this case, the methodology was followed, from Problem Understanding to the successful deployment of the models in the shop floor.
Costa et al. [41] applied the full methodology to determine the working states of a plastic injection machine. The model was also deployed in a large plastic industry. Contrary to the previously mentioned projects, where the data collection process was also designed during the project, in this case, the dataset consisted of several years of records already available. Nonetheless, the methodology was successfully applied from Problem Understanding to Deployment.
Silva et al. [42] also used the MIDA methodology in a project where the aim was to improve automatic help for the encoding of Electronic Health Records in ICD-10 taxonomy. This project did not aim at developing a model ready for deployment. However, it was particularly important because of the sensitivity of the data used. Access to EHR requires particular attention to the ethical and legal implications of accessing, storing and processing such data, and was one of the main motivators to include task 1.3 in the methodology.
During the projects cited above in this section, the weaknesses revealed by the original CRISP-DM, its variations and competing methodologies lead to the refinement of MIDA over time. In the end, the methodology proposed was revealed to be necessary and sufficient for the successful completion of all the projects. The methodology is also being used in other ongoing projects, whose results will be published soon.

5.2. Advantages

In those projects, the main advantages of MIDA over older methodologies became clear: (i) Problem understanding is a much more specific term, keeping focus on the problems at stake rather than the business side of the issues. MIDA helps keeping the focus of the project and it was advantageous in all the projects; (ii) The need for a careful requirement analysis is also clear in MIDA. This is particularly useful when dealing with real industrial customers; (iii) The need to check regulatory and legal issues that may affect the project is also clear. This was specially important in the projects by Cruz et al. [40] and Silva et al. [42], because of the specificities of the data being used; (iv) Planning and data collection are also clear steps; and (v) The need to consider maintenance of the system and data governance is also clear in the project, when the models are deployed. In this particular case, the projects by Cruz et al. [40] and Costa et al. [41] included the deployment phase and therefore particular attention needed to be devoted to that phase.
MIDA is therefore an important contribution to the state of the art, as the updates and additions, when compared to CRISP-DM and similar methodologies, facilitate project management and help the on-boarding of young researchers, who benefit from a more up-to-date and straightforward roadmap for data analysis projects.

5.3. Limitations

During the research, no clear limitations of the methodology itself were identified, except perhaps that it has two more steps to consider, when compared to the original CRISP-DM. The main limitation, however, is the validation of the methodology itself. Comparison of the results of MIDA to other methodologies in statistically valid experiments with control groups are out of the scope of the current project and are, therefore, left as future work.

6. Conclusions

CRISP-DM remains one of the most widely adopted approaches for data-driven projects. However, its high-level structure, lack of explicit planning and data collection stages, and absence of considerations for modern regulatory constraints reduce its effectiveness in modern real-world industrial environments.
To overcome those shortcomings, we proposed MIDA—Method for Industrial Data Analysis—an eight-step methodology designed and refined through practical applications in engineering and industrial projects. MIDA builds upon the foundations of CRISP-DM but incorporates explicit stages for planning, data collection, legal analysis, and a more detailed and domain-oriented understanding of the problem. Its structure aims to provide a clearer roadmap for projects where multidisciplinary knowledge, regulatory compliance, and data heterogeneity are the norm.
MIDA has already been successfully applied in different real-world projects, such as the analysis of operation states in injection-molding equipment, estimation of forest biomass using deep learning approaches and encoding of EHR using ICD-10. These applications demonstrate the methodology’s flexibility and its suitability for different scenarios.
The advantages of MIDA include its explicit emphasis on planning, legal requirements, and structured data collection, as well as its balance between technical rigor and practical applicability. Nevertheless, MIDA may still need adaptation when used in domains radically different from those for which it was originated.
Future work includes validating the methodology across additional industrial sectors and refining specific tasks for particular domains if needed. Ongoing projects developed by the authors and collaborators will provide further feedback to improve MIDA and reinforce its role as a practical, domain-oriented alternative to the classical CRISP-DM framework.

Author Contributions

Conceptualization: M.M.; Methodology: M.M. and T.F.; Formal analysis and investigation: M.M. and T.F.; Writing—original draft preparation: M.M.; Writing—review and editing: T.F. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article.

Conflicts of Interest

The authors declare that they have no known competing financial or personal relationships with other people or organizations that could inappropriately influence the work reported in the paper. There is no professional or other personal interest of any nature or kind in any product or company that could be influencing the positions presented in, or the review of, the manuscript.

References

  1. Chapman, P.; Clinton, J.; Kerber, R.; Khabaza, T.; Reinartz, T.; Shearer, C.; Wirth, R. CRISP-DM 1.0; IBM: Armonk, NY, USA, 1999; Available online: https://public.dhe.ibm.com/software/analytics/spss/documentation/modeler/14.2/es/CRISP-DM.pdf (accessed on 31 January 2026).
  2. Piatetsky-Shapiro, G. Knowledge discovery in real databases: A report on the IJCAI-89 workshop. AI Mag. 1990, 11, 68–70. [Google Scholar]
  3. Fayyad, U.; Piatetsky-Shapiro, G.; Smyth, P. From data mining to knowledge discovery in databases. AI Mag. 1996, 17, 37–54. [Google Scholar]
  4. SAS Institute. Data Mining and the Case for Sampling; Technical Report; SAS Institute: Cary, NC, USA, 2003; Available online: https://www.scribd.com/document/339713338/SAS-SEMMA (accessed on 31 January 2026).
  5. Wirth, R.; Hipp, J. CRISP-DM: Towards a standard process model for data mining. In Proceedings of the 4th International Conference on the Practical Applications of Knowledge Discovery and Data Mining, Manchester, UK, 11–13 April 2000; Volume 1, pp. 29–39. Available online: http://www.cs.unibo.it/~danilo.montesi/CBD/Beatriz/10.1.1.198.5133.pdf (accessed on 31 January 2026).
  6. Azevedo, A.; Santos, M. KDD, SEMMA and CRISP-DM: A parallel overview. In IADIS European Conf. Data Mining; 2008; Volume 8, pp. 182–185. Available online: https://www.academia.edu/download/43395778/KDD-SEMMA-CRISP-DM_DM_2008.pdf (accessed on 31 January 2026).
  7. Shameti, K.; Cico, B. Comparison of methodological approaches: CRISP-DM vs. OSEMN methodology using linear regression and statistical analysis. In Balkan Conference in Informatics; Springer: Cham, Switzerland, 2024; pp. 47–60. Available online: https://link.springer.com/chapter/10.1007/978-3-031-84093-7_4 (accessed on 31 January 2026).
  8. Dineva, K.; Atanasova, T. OSEMN process for working over data acquired by iot devices mounted in beehives. Curr. Trends Nat. Sci. 2018, 7, 47–53. [Google Scholar]
  9. Kumari, K.; Bhardwaj, M.; Sharma, S. OSEMN approach for real time data analysis. Int. J. Eng. Manag. Res. 2020, 10, 107–110. [Google Scholar] [CrossRef] [Scilit]
  10. Martínez-Plumed, F.; Contreras-Ochando, L.; Ferri, C.; Hernández-Orallo, J.; Kull, M.; Lachiche, N.; Ramírez-Quintana, M.J.; Flach, P. CRISP-DM twenty years later: From data mining processes to data science trajectories. IEEE Trans. Knowl. Data Eng. 2021, 33, 3048–3061. [Google Scholar] [CrossRef] [Scilit]
  11. Schröer, C.; Kruse, F.; Gómez, J.M. A systematic literature review on applying CRISP-DM process model. Procedia Comput. Sci. 2021, 181, 526–534. [Google Scholar] [CrossRef] [Scilit]
  12. Shimaoka, A.M.; Ferreira, R.C.; Goldman, A. The evolution of CRISP-DM for data science: Methods, processes and frameworks. SBC Comput. Rev. 2024, 4, 28–43. [Google Scholar] [CrossRef] [Scilit]
  13. Zavaleta-Sánchez, E.; Domínguez-Sánchez, G.; Irene Loeza-Mejía, C.; Sánchez-DelaCruz, E. Comparative study of KDD and CRISP-DM methodologies for phishing identification. In International Congress on Information and Communication Technology; Springer: Singapore, 2024; pp. 317–330. Available online: https://link.springer.com/chapter/10.1007/978-981-97-3559-4_25 (accessed on 31 January 2026).
  14. Gómez Palacios, H.J.; Jiménez Toledo, R.A.; Pantoja, G.A.H.; Navarro, Á.A.M. A comparative between CRISP-DM and SEMMA through the construction of a MODIS repository for studies of land use and cover change. Adv. Sci. Technol. Eng. Syst. J. 2017, 2, 598–604. [Google Scholar] [CrossRef] [Scilit]
  15. Rosander, A. Evaluating Frameworks for Implementing Machine Learning in Signal Processing. Master’s Thesis, Kth Skolan för Electroteknik Och Datavetenskap, Stockholm, Sweden, 2018. Available online: https://www.diva-portal.org/smash/get/diva2:1250897/FULLTEXT01.pdf (accessed on 31 January 2026).
  16. Doede, N.; Merkel, P.; Kriwall, M.; Stonis, M.; Behrens, B. Implementation of an intelligent process monitoring system for screw presses using the CRISP-DM standard. Prod. Eng. 2025, 19, 77–88. [Google Scholar] [CrossRef] [Scilit]
  17. Jaggia, S.; Kelly, A.; Lertwachara, K.; Chen, L. Applying the CRISP-DM framework for teaching business analytics. Decis. Sci. J. Innov. Educ. 2020, 18, 612–634. [Google Scholar] [CrossRef] [Scilit]
  18. Solano, J.A.; Cuesta, D.J.L.; Ibáñez, S.F.U.; Coronado-Hernández, J.R. Predictive models assessment based on CRISP-DM methodology for students performance in colombia—Saber 11 test. Procedia Comput. Sci. 2022, 198, 512–517. [Google Scholar] [CrossRef] [Scilit]
  19. Purbasari, A.; Rinawan, F.R.; Zulianto, A.; Susanti, A.I.; Komara, H. CRISP-DM for data quality improvement to support machine learning of stunting prediction in infants and toddlers. In Proceedings of the 2021 8th International Conference on Advanced Informatics: Concepts, Theory and Applications (ICAICTA), Bandung, Indonesia, 29–30 September 2021; pp. 1–6. Available online: https://www.researchgate.net/profile/Ayi-Purbasari/publication/357107554_CRISP-DM_for_Data_Quality_Improvement_to_Support_Machine_Learning_of_Stunting_Prediction_in_Infants_and_Toddlers/links/64df7dac1351f5785b7305fc/CRISP-DM-for-Data-Quality-Improvement-to-Support-Machine-Learning-of-Stunting-Prediction-in-Infants-and-Toddlers.pdf (accessed on 31 January 2026).
  20. Rahmadi, L.; Hadiyanto; Sanjaya, R.; Prambayun, A. Crop prediction using machine learning with CRISP-DM approach. In Proceedings of Data Analytics and Management; Swaroop, A., Polkowski, Z., Correia, S.D., Virdee, B., Eds.; Springer Nature: Singapore, 2023; pp. 399–421. Available online: https://link.springer.com/chapter/10.1007/978-981-99-6550-2_31 (accessed on 31 January 2026).
  21. Bosnjak, Z.; Grljevic, O.; Bosnjak, S. CRISP-DM as a framework for discovering knowledge in small and medium sized enterprises’ data. In Proceedings of the 2009 5th International Symposium on Applied Computational Intelligence and Informatics; IEEE: New York, NY, USA, 2009; pp. 509–514. Available online: https://ieeexplore.ieee.org/abstract/document/5136302 (accessed on 31 January 2026).
  22. Kannengiesser, U.; Gero, J.S. Modelling the design of models: An example using CRISP-DM. Proc. Des. Soc. 2023, 3, 2705–2714. [Google Scholar] [CrossRef] [Scilit]
  23. Saltz, J.S. CRISP-DM for data science: Strengths, weaknesses and potential next steps. In Proceedings of the 2021 IEEE International Conference on Big Data (Big Data); IEEE: New York, NY, USA, 2021; pp. 2337–2344. Available online: https://ieeexplore.ieee.org/abstract/document/9671634 (accessed on 31 January 2026).
  24. Plotnikova, V.; Dumas, M.; Milani, F.P. Applying the CRISP-DM data mining process in the financial services industry: Elicitation of adaptation requirements. Data Knowl. Eng. 2022, 139, 102013. [Google Scholar] [CrossRef] [Scilit]
  25. Huber, S.; Wiemer, H.; Schneider, D.; Ihlenfeldt, S. DMME: Data mining methodology for engineering applications—A holistic extension to the CRISP-DM model. Procedia Cirp 2019, 79, 403–408. [Google Scholar] [CrossRef] [Scilit]
  26. Cazacu, M.; Titan, E. Adapting CRISP-DM for social sciences. BRAIN Broad Res. Artif. Intell. Neurosci. 2021, 11, 99–106. [Google Scholar] [CrossRef] [Scilit]
  27. Venter, J.; de Waal, A.; Willers, C. Specializing CRISP-DM for evidence mining. In IFIP International Conference on Digital Forensics; Springer: New York, NY, USA, 2007; pp. 303–315. Available online: https://link.springer.com/chapter/10.1007/978-0-387-73742-3_21 (accessed on 31 January 2026).
  28. Nagashima, H.; Kato, Y. APREP-DM: A framework for automating the pre-processing of a sensor data analysis based on CRISP-DM. In Proceedings of the 2019 IEEE International Conference on Pervasive Computing and Communications Workshops (PerCom Workshops); IEEE: New York, NY, USA, 2019; pp. 555–560. Available online: https://ieeexplore.ieee.org/abstract/document/8730785 (accessed on 31 January 2026).
  29. Ayele, W.Y. Adapting CRISP-DM for idea mining: A data mining process for generating ideas using a textual dataset. Int. J. Adv. Comput. Sci. Appl. 2020, 11, 20–32. [Google Scholar]
  30. Bokrantz, J.; Subramaniyan, M.; Skoogh, A. Realising the promises of artificial intelligence in manufacturing by enhancing CRISP-DM. Prod. Plan. Control 2024, 35, 2234–2254. [Google Scholar] [CrossRef] [Scilit]
  31. Kristoffersen, E.; Aremu, O.O.; Blomsma, F.; Mikalef, P.; Li, J. Exploring the relationship between data science and circular economy: An enhanced CRISP-DM process model. In Digital Transformation for a Sustainable Society in the 21st Century; Pappas, I.O., Mikalef, P., Dwivedi, Y.K., Jaccheri, L., Krogstie, J., Mymki, M., Eds.; Springer International Publishing: Cham, Switzerland, 2019; pp. 177–189. Available online: https://link.springer.com/chapter/10.1007/978-3-030-29374-1_15 (accessed on 31 January 2026).
  32. Acuña-Cid, H.A.; Ahumada-Tello, E.; Ovalle-Osuna, Ó.O.; Evans, R.; Hernández-Ríos, J.E.; Zambrano-Soto, M.A. CRISP-NET: Integration of the CRISP-DM model with network analysis. Mach. Learn. Knowl. Extr. 2025, 7, 101. [Google Scholar] [CrossRef] [Scilit]
  33. Catley, C.; Smith, K.; McGregor, C.; Tracy, M. Extending CRISP-DM to incorporate temporal data mining of multidimensional medical data streams: A neonatal intensive care unit case study. In Proceedings of the 2009 22nd IEEE International Symposium on Computer-Based Medical Systems; IEEE: New York, NY, USA, 2009; pp. 1–5. Available online: https://ieeexplore.ieee.org/abstract/document/5255394 (accessed on 31 January 2026).
  34. Asamoah, D.; Sharda, R. Adapting CRISP-DM Process for Social Network Analytics: Application to Healthcare. 2015. Available online: https://aisel.aisnet.org/amcis2015/BizAnalytics/GeneralPresentations/33/ (accessed on 31 January 2026).
  35. Ramos, J.L.C.; Rodrigues, R.L.; Silva, J.S.; de Oliveira, P.L.S. CRISP-EDM: Uma proposta de adaptação do modelo CRISP-DM para mineração de dados educacionais. In Simpósio Brasileiro de Informática na Educação (SBIE); SBC, 2020; pp. 1092–1101. Available online: https://sol.sbc.org.br/index.php/sbie/article/view/12865/12719 (accessed on 31 January 2026).
  36. Niaksu, O. CRISP data mining methodology extension for medical domain. Balt. J. Mod. Comput. 2015, 3, 92. [Google Scholar]
  37. Schäfer, F.; Zeiselmair, C.; Becker, J.; Otten, H. Synthesizing CRISP-DM and quality management: A data mining approach for production processes. In Proceedings of the 2018 IEEE International Conference on Technology Management, Operations and Decisions (ICTMOD); IEEE: New York, NY, USA, 2018; pp. 190–195. Available online: https://ieeexplore.ieee.org/abstract/document/8691266 (accessed on 31 January 2026).
  38. Jajal, P.; Jiang, W.; Tewari, A.; Kocinare, E.; Woo, J.; Sarraf, A.; Lu, Y.; Thiruvathukal, G.K.; Davis, J.C. Analysis of failures and risks in deep learning model converters: A case study in the onnx ecosystem. arXiv 2023, arXiv:2303.17708. [Google Scholar]
  39. Rodrigues, J.; Costa, I.; Farinha, J.T.; Mendes, M.; Margalho, L. Predicting motor oil condition using artificial neural networks and principal component analysis. Eksploat. Niezawodn. 2020, 22, 440–448. [Google Scholar] [CrossRef] [Scilit]
  40. Cruz, S.; Paulino, A.; Duraes, J.; Mendes, M. Real-time quality control of heat sealed bottles using thermal images and artificial neural network. J. Imaging 2021, 7, 24. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  41. Costa, J.; Silva, R.; Martins, G.; Barreiros, J.; Mendes, M. Analysis of the state and fault detection of a plastic injection machine—A machine learning-based approach. Algorithms 2025, 18, 521. [Google Scholar] [CrossRef] [Scilit]
  42. Silva, H.; Duque, V.; Macedo, M.; Mendes, M. Aiding ICD-10 encoding of clinical health records using improved text cosine similarity and PLM-ICD. Algorithms 2024, 17, 144. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.