1. Introduction
Rolling bearings are critical machine elements in industrial drive trains, manufacturing equipment, and energy systems, where tribological behavior directly affects efficiency [
1], reliability [
2,
3], and operating costs [
4]. In many applications, small changes in design [
5], operating conditions [
6], and tolerances [
7] can propagate into changes in system-level performance. Consequently, industry has a strong incentive to predict and optimize rolling bearing performance early in the development process [
8], when design decisions have the largest leverage on cost and functionality.
Physics-based simulations are a key enabler for such predictions in early phases, ranging from numerical contact models to elastohydrodynamic lubrication solvers, finite element analyses, and multi-body dynamics. However, classical simulation workflows face persistent challenges when used for design and optimization: simulations are often computationally expensive and capturing realistic operating envelopes typically requires extensive parameter sweeps and reasoned design of experiments (DoE) campaigns [
9]. As a result, engineers trade model fidelity for tractability or restrict analyses to a limited set of operating conditions, which can mask critical interactions and reduce confidence in early-stage decisions [
10].
From a product development perspective, this limitation contradicts the goal of frontloading engineering [
11]—i.e., bringing forward knowledge about design variations and their impact on key performance metrics to the earliest phases, when changes are comparatively inexpensive. The “rule of ten” based on Boehm [
12] and Clark and Fujimoto [
13] illustrates that the cost of design changes increases by the factor of ten with each development phase, making late discovery of performance issues costly [
14,
15]. In order to gain meaningful insights, engineers should ideally analyze many designs and operating conditions [
16], including tolerances [
7,
17] and uncertainties [
18], which requires a significant number of simulations. However, in practice, the required computational and modeling effort can be prohibitively high. Machine learning (ML) offers a pragmatic way to overcome these limitations by complementing physics-based simulations with data-driven models. According to Wang [
19], ML is defined as a branch of artificial intelligence that studies and builds models that can learn from data or the environment to make predictions about unknown data or make decisions based on environmental information.
In this review, the roles of ML in simulation-based design and optimization of rolling bearings are categorized as “replace” and “enhance” to distinguish between different levels of workflows.
In a
replacement role, ML is used to build a surrogate model that substitutes the physics-based simulation model for individual predictions. In this case, the surrogate primarily serves as a computationally efficient predictor, while the underlying physics-based simulations remain necessary for generating training data and for validation purposes [
20,
21,
22,
23,
24]. In an
enhancing role, ML is used to improve or extend the simulation workflow while the physics-based model remains an integral part of the overall methodology. This may include reducing the number of required simulation runs, increasing computational efficiency, or enabling downstream tasks such as optimization or design exploration based on simulation data [
25,
26,
27,
28].
ML surrogates can approximate computationally intensive simulations at a fraction of the cost, enabling rapid investigation of a large set of input parameters. In a broader context, ML can support sensitivity analyses, uncertainty quantification, and optimization by reducing the number of simulations required while maintaining predictive accuracy within a defined domain. The integration of simulation-driven data mining with ML enables the development of scalable workflows that provide early quantitative insights into the impact of input parameters (known as features) on bearing performance. As illustrated in
Figure 1, this approach significantly enhances the upfront incorporation of engineering and technical knowledge compared to traditional methods based solely on simulation.
ML encompasses a broad spectrum of methods structured around paradigms of supervised, unsupervised, and reinforcement learning. Across these paradigms, ML techniques address a range of task types, including classification and regression for predictive modeling, clustering to discover latent structures in unlabeled data, ranking to order alternatives by relevance or utility, and dimensionality reduction for compact representations of high-dimensional inputs. In addition, ML methods can uncover correlations and dependency patterns in data and support decision-making by matching observations with recommended actions or design choices within an optimization or control context [
19].
The application of those paradigms in the field of tribology and rolling bearings is investigated in various studies and review articles. Rosenkranz et al. [
29] characterized the state of the art in tribology as strongly dominated by artificial neural networks (ANNs), with particular maturity in condition monitoring applications such as automated particle analysis and fault diagnosis, while also highlighting growing use of ANN surrogate models to reduce the computational effort of thermo-hydrodynamic and thermo-elastohydrodynamic contacts. The authors emphasized persistent gaps—in particular, the uncertainty of experimental data, limited transferability between tribosystems, reduced physical transparency in ANN models, and the lack of a common data infrastructure—as well as the observation that deep learning models are not universally superior to simpler approaches for certain tasks.
Building on this broader tribology perspective, Marian and Tremmel [
30] noted that approximately 75% of the tribological publications they examined use ANNs for ML. Furthermore, vibration-based condition monitoring and fault diagnosis continue to be the predominant use cases in bearing applications. They also note a growing trend toward metamodeling computationally intensive contact problems and an emerging interest in physics-informed ML, identifying data availability and comparability as the main problems, exacerbated by limited transparency due to unavailable datasets and models, as well as the insufficient use of alternative methods such as support vector machines (SVMs) or random forests.
A systematic review by Singh et al. [
31] categorized the leading ML methods for extracting features from vibration signals in the field of rolling bearing fault prognosis and diagnosis. They identified a significant shift from traditional techniques such as ANNs and SVMs to deep learning (DL) for rolling bearing diagnosis, while introducing a systematic benchmarking framework and quality function deployment chart to address the urgent need for quantifiable algorithm selection in industrial AI applications. They emphasized that the literature remains heavily focused on diagnosis rather than prognosis and is often validated using benchmark datasets, with key gaps including limited validation under variable real operating conditions, persistent data-quality issues, insufficient multisensor fusion, and ongoing interpretability barriers that hinder industrial trust.
More recently, Marian and Tremmel [
32] addressed the rise of physics-informed ML as a central development trend, with physics-informed neural networks (PINNs) as a leading approach for meshless solutions of the Reynolds equation and for predicting lubrication film thickness and pressure distributions. They note method-specific limitations such as the discrepancy between low residual physics losses and actual prediction errors, and the need to extend models beyond standard formulations of the Reynolds equation to capture non-stationary flow behavior, the compressibility of lubricants, or shear thinning fluids, thus ensuring applicability in various technical contexts.
Complementing these observations, Paturi et al. [
33] confirmed the predominance of ANN-based modeling in tribology, particularly multilayer perceptrons trained with backpropagation, while also pointing to the rapid spread of convolutional neural networks (CNNs), recurrent neural networks (RNNs) and long short-term memory (LSTM) networks for bearing-related tasks such as fault diagnosis and remaining useful life (RUL) prediction. Combining these algorithms with genetic algorithms (GA) further reduces the prediction error.
Both Marian and Tremmel [
30] and Paturi et al. [
33] advocated for the open dissemination of datasets and models via repositories, while emphasizing that automating the currently manual data acquisition processes is essential to fully leverage the potential of ML in tribology.
From the perspective of computational tribology [
34] and materials research, Sose et al. [
35] highlighted the widespread use of ANNs for predicting friction and wear parameters, as well as the increasing coupling of ML as the objective function of evolutionary and swarm-based optimizers for the development of novel materials. The authors point to an important imbalance: despite the enormous amounts of data available from molecular simulations and related computational approaches, ML is still underutilized in simulation-oriented tribology, partly due to the lack of standardized analysis tools that can translate simulation data into learnable representations, as well as continuing challenges regarding interpretability and robustness in the case of noisy or incomplete data within experiments.
In their comprehensive review article, Ayman et al. [
36] outlined four priorities for strengthening the practical application of bearing prognostics in industry: (i) Moving beyond vibration measurements through multimodal learning that integrates additional sensing modalities such as temperature measurements and thermal imaging to better detect early-stage degradation; (ii) Addressing temporal and computational complexity as data-driven models become more sophisticated, including the development of compressed models suitable for real-time monitoring under limited computational resources; (iii) Transitioning from component-level prediction to system-level prediction that can account from multiple interacting components, which requires data from real-world manufacturing systems given the scarcity of benchmark datasets for complex systems; and (iv) supplementing point estimates of RUL with explicit uncertainty quantification—for example, via Bayesian models—to support more reliable decision-making in practice.
In summary, there are several review articles in the broader field of tribology. As Rosenkranz et al. [
29] noted, highly domain-specific expert knowledge is crucial for solving tribological issues, and this requirement becomes particularly apparent when data-driven methods are applied to different tribosystems and operating conditions. In contrast, review articles that explicitly target the design of rolling bearings are still relatively rare. As Marian et al. [
30] synthesized in their review, most published work on rolling bearings in the context of ML tends to refer to vibration theory rather than tribology, which is still the case today. This is a well-researched area, especially in the field of rolling bearing fault diagnosis. Simulations are often combined with ML methods to generate large and diverse training datasets across different operating conditions and bridge the domain gap between simulated and measured signals. This improves the robustness of fault diagnosis under noise and variable speeds while reducing dependence on costly experimental vibration data [
37,
38,
39,
40,
41,
42,
43,
44,
45,
46,
47]. However, the potential of ML-supported simulations to frontload engineering knowledge in rolling bearing design has not yet been systematically examined; this review addresses this gap. Studies that apply ML exclusively for condition monitoring or fault diagnosis are excluded. The review considers peer-reviewed journal and conference publications in English and systematically synthesizes methods, applications, and emerging research trends in this field.
In order to structure the review and ensure a comprehensible synthesis, three research questions (RQ) guide the analysis. First, in
RQ1, from a methodological perspective, the question is which ML methods have been applied to replace or enhance simulations in the design and optimization of rolling bearings, and how these approaches differ in terms of modeling capabilities, data requirements, and computational performance. Second, from an application and domain perspective,
RQ2 examines for which tribological phenomena and simulation domains—such as lubrication, contact mechanics, and dynamics—ML has been used in rolling bearing simulations, and which simulation outputs are typically predicted. Third, to capture the state of the art and its evolution,
RQ3 investigates current trends, limitations, and research gaps in ML-based bearing simulations and identifies the methodological and technological developments needed to advance knowledge further. The remainder of this review article is structured as follows:
Section 2 outlines the methodology,
Section 3 synthesizes results by application domain and ML workflow,
Section 4 discusses gaps and implications, and
Section 5 concludes with recommendations and an outlook.
2. Methods
After introducing the research problem in
Section 1—in particular, the identification of trends and applications of ML algorithms that enable early stage product development of rolling bearings in the context of simulations—a systematic literature review was conducted to answer research questions RQ1–RQ3. The main objective of this methodology is to ensure a consistent, transparent, and reproducible search, selection, and synthesis process and to clearly distinguish this review from existing tribology-oriented reviews by adhering to the PRISMA 2020 guideline [
48]. The review procedure was specified in a review protocol and is registered on the Open Science Framework [
49] to document the research strategy, eligibility criteria, and screening and extraction workflow. The resulting evidence base is reported in
Section 3 and interpreted in
Section 4, followed by a consolidated conclusion and an outlook on future research directions and practical implications for ML-supported simulation workflows in rolling bearing technology.
2.1. Literature Search Strategy and Study Selection
The literature search was conducted using the Scopus, Web of Science, and IEEE Xplore databases. In Scopus, the search query was applied to the “title-abstract-keywords” field, while in Web of Science and IEEE Xplore, the search was performed across all fields. Only English-language, peer-reviewed journal and conference publications published between 2005 and 2025 were considered, as the topic is relatively new and earlier publications would not contribute meaningfully to answering the research questions. The search strategy was divided into four comprehensive concept blocks: (i) the product domain, (ii) the physical domain, (iii) ML methods, and (iv) modeling and simulation. Accordingly, the product domain focused on rolling bearings, while the physical domain covered the relevance to the subject area using terms related to tribology, friction, lubrication, and dynamics. The methods were formulated in general terms to maximize the hit rate and included alternative spellings, acronyms, and full expressions, such as ML, Artificial Intelligence (AI), surrogate models, metamodels, data-driven, PINN, feature learning, GA, decision trees, ANN, SVM, DL, Bayesian and Kriging methods and k-Nearest Neighbors (kNN). The “modeling and simulation” block covered terms related to modeling, numerical analysis, design, computation, optimization, and established simulation approaches and tools, including elastohydrodynamic lubrication (EHL), the finite element method (FEM), and computational fluid dynamics (CFD), as well as analytical and numerical investigations.
To maintain focus on early design and optimization using simulations, keywords associated with prognostics and health management were excluded, including terms related to RUL, fault detection, and condition monitoring. These topics are primarily relevant in later lifecycle phases and were therefore outside the scope of this review. The complete search string, reflecting the four inclusive blocks and the exclusions, is provided in
Figure 2. The initial search was executed on 15 December 2025, followed by an update check on 14 January 2026 to capture newly published material.
2.2. Literature Scanning Process
After completing the searches and exporting all retrieved records, retracted records (
n = 3) and duplicates from Scopus, Web of Science, and IEEE Xplore (
n = 45) were removed, and the subsequent screening was carried out in Zotero. Studies were selected based on the predefined inclusion and exclusion criteria summarized in
Table 1. The selection process followed a two-stage screening procedure. In the first stage, the titles and abstracts of all records found were reviewed based on the criteria. A total of 157 records were screened, resulting in 20 reports sought for retrieval after title and abstract screening. Four reports could not be retrieved, and the full texts of the remaining 16 reports were assessed to confirm eligibility. Five reports were excluded because no ML was included. The remaining 11 reports were included for methodological data extraction. To further reduce the likelihood of missing relevant work, a backward reference search was performed by screening the reference lists of the included full-text studies; no additional studies were identified.
This systematic review adheres to the core reporting principles of PRISMA 2020 and incorporates key procedural elements of the framework by Kitchenham and Charters [
50], including a predefined screening workflow, explicit eligibility criteria, and a structured data extraction scheme aligned with research questions RQ1–RQ3. A formal tool for assessing potential bias and an overall quality assessment were not applied because the included studies are methodologically heterogeneous, which would make a uniform assessment difficult to interpret. Instead, methodological rigor was captured through consistently extracted indicators, including dataset size, ML approach, the reported inputs and outputs, prediction metrics, bearing type, simulation method, main findings, study overview, and bibliometric information. Due to the substantial variability in study designs and reported performance metrics, a quantitative meta-analysis was not feasible; therefore, findings were integrated through a narrative synthesis across method classes and tribological simulation domains. The overall scope was intentionally defined to capture ML approaches that replace or enhance simulations in the design and optimization of rolling bearings, while studies focused exclusively on condition monitoring or diagnostics were excluded in order to maintain the focus on simulation-centered product development. Search strings and eligibility criteria are reported explicitly, and all screening decisions were documented to ensure traceability of study selection. Finally, the PRISMA flow diagram is provided in
Figure 3.
In addition, dataset size categories were defined for this review based on the rule of thumb commonly used in surrogate modeling, according to which the recommended sample size is approximately 10–100 times [
51,
52,
53] the number of input features
[
54]. Based on recent studies, Alwosheel et al. [
54] suggested a factor
for the rule of thumb, indicating that
may be too optimistic. Since there are no universal numerical thresholds in the literature, we introduced five practically motivated intervals applicable to all kinds of rolling bearing simulations, reflecting typical feature dimensionalities (
), to enable consistent comparison across studies, see
Table 2.
3. Results
Despite the systematic search, some relevant studies may not have been identified due to limitations in database indexing and variability in keywords and terminology. Across all 11 studies included, the evidence base remains limited and sparse over time, with only a few publications per year and a noticeable concentration in recent years, as shown in
Figure 4. Therefore, a deeper quantitative trend analysis was not pursued. Instead,
Section 3.1 synthesizes the applications of ML in rolling bearing simulations to address the research questions RQ1–RQ2, and
Section 3.2 summarizes trends and emerging directions in ML related to rolling bearing simulations with regard to research question RQ3. Finally,
Section 3.3 consolidates methodological patterns into a set of practical recommendations that can be adapted to different bearing types, domains, and simulation tools.
3.1. Applications of ML in Rolling Bearing Simulations
The rolling bearing types covered in the included studies are distributed heterogeneously across the most common bearing types. A total of five bearing types are included, representing both point and line contacts. As summarized in
Table 3, the studies use design parameters and operating conditions as input features. In most cases, both feature groups are combined in the design of experiments to capture interaction effects and thus assess how design variations affect performance indicators under changing operating conditions such as speed or lubrication.
Building on the distribution of bearing types and the associated input feature groups, we next examine how these inputs are translated into simulation objectives across tribological domains.
Table 4 therefore assigns the studies included to the underlying simulation domain and the output quantities, highlighting which performance metrics are most frequently focused on in ML-supported rolling bearing simulations.
There is a clear correlation between the simulation area and the domain-specific output quantities, which is primarily determined by the capabilities and limitations of the underlying simulation tools. Tribology oriented simulations predominantly target churning power losses, film thickness, stress-related measures, and friction parameters, while investigations into bearing dynamics primarily predict dynamic response indicators. From a tool perspective, tribological and dynamic simulations are therefore largely implemented as separate workflows. One notable exception is the use of CABA3D multi-body simulation tool [
63] in [
16], which builds on the scientific work of Vesselinov [
64]. This enables closer coupling between dynamic and tribological performance quantities.
3.1.1. Contact Mechanics
The tribology-oriented subfield encompasses contact mechanics, EHL modeling, and CFD, with ML predominantly used to reduce the marginal costs of investigating design and operating variations, while physics-based simulation is retained as a reference.
Within the field of contact mechanics, Lostado et al. [
61,
62] used FEM simulations to train supervised learning models that predict contact pressure [
61] and contact stress ratio [
62] in tapered roller bearings. The models are learned from operating and geometric descriptors (preload, load, CoF, angular and longitudinal position [
61]; preload, radial and axial load, and torque [
62]) and subsequently used as computationally efficient surrogates. The common goal of both studies is to represent highly nonlinear mechanical behavior, with particular emphasis on contact stress and contact-pressure distributions, while avoiding the computational burden of conventional FEM workflows. To this end, the authors followed a standard data mining pipeline consisting of data generation, model training, and testing. First, a set of representative operating cases was simulated using FEM to build a training database. Several regression techniques were then compared, including linear regression, regression trees, SVMs, ANNs, and GPR. Among the models evaluated, M5P achieves an RMSE < 1.71% [
61], while an SVM with an RBF kernel achieves an RMSE < 4.94% [
62]. A characteristic feature of the latter publication is the explicit integration of GA [
62]. To determine an optimal operating load configuration for TRBs, the study employs GA, a metaheuristic optimization paradigm inspired by natural selection and biological evolution. GAs are well suited for global search in complex, high dimensional design spaces by iteratively evolving a population of candidate solutions through a structured cycle of coding/decoding, initialization, evaluation, selection, crossover, and mutation [
65]. In this case, the GA is used to maximize radial load capacity while enforcing key performance constraints—specifically, maintaining contact stress ratios of approximately 25% in both rows of rollers. To keep the optimization computationally tractable, the GA is coupled with surrogate models rather than FE (finite element) simulations. In combination, the compact dataset and multiple surrogate candidates illustrate an efficient optimization workflow within this class of contact-mechanics problems.
Zhao et al. [
60] applied ML in a reliability-oriented bearing simulation setting of main shaft bearings in wind turbines, whereby the greatest challenge lies not in a single high-precision run, but in the large number of evaluations required for sensitivity and probabilistic assessment. The study addresses the subsurface stress response of a cylindrical roller bearing and explicitly considers variability arising from machining and assembly. In this workflow, FEM models are used as the primary simulator and Kriging surrogate models (DS2) are trained to approximate the stress quantity over a space defined by the structure thicknesses, the friction conditions, and the interference conditions as inputs. The target output is the subsurface stress reliability, enabling Monte Carlo propagation of uncertainty through the surrogate rather than repeated FEM evaluations. The specified runtime underscores the motivation: full probabilistic evaluation required 13.113 s, and the surrogate-based approach is approximately 780 times faster than FEM with a relative error of 3.8 × 10
−5.
Wang et al. [
57] adopted a similar approach, combining Kriging surrogate modeling with a Latin hypercube design (LHD) and GAs to optimize the sealing performance of a DGBB. The study addresses a fundamental trade-off in the underlying optimization problem: If the contact pressure is too low, the film thickness between the contact surfaces increases, which can promote grease leakage; conversely, if the contact pressure is too high, the film thickness decreases and wear between the contact partners is aggravated, thereby reducing the service life of the oil seal. The input features comprise both design parameters—in particular, the angle of inclination of the sealing lip and radial interference—and operating conditions, including the coefficient of friction, lubricant temperature, and rotational speed. The outputs are defined as maximum von Mises stress and maximum contact pressure, which serve as quantitative indicators of seal performance. Different operating conditions were verified through a three-level orthogonal experiment using the three operating condition features, and the resulting high accuracy was reported as evidence of the correctness of both the thermally stress-coupled FEM model and the Kriging surrogate [
57]. Overall, the contact mechanics related studies exemplify how ML can replace and enhance FE simulation workflows in tasks such as prediction, optimization, and uncertainty quantification.
3.1.2. Lubrication
In the field of tribology, with a focus on lubrication regimes, EHL and CFD simulations are used to predict film thickness, stress and friction-related parameters, and churning power losses, respectively.
Wirsching et al. [
25] extended surrogate modeling to the tribological domain and developed a Kriging metamodel (DS3) to support design optimization of an EHL roller face/rib contact in TRBs in order to improve energy efficiency. The authors justify their approach with the high computational costs incurred by repeatedly solving lubrication problems when investigating design alternatives. They use an LHD for sampling EHL simulation data points to generate training data and a metamodel of optimal prognosis (MOP) that maps key design features—roller face radius, rib radius, and eccentricity—to multiple performance targets, namely pressure, minimum lubricant gap, and coefficient of friction. MOP automatically selects the most accurate surrogate model based on a coefficient of prognosis (COP). Following this approximation, a multi-objective GA is applied to the metamodels to find global optima, followed by verification through recalculations in EHL simulations. Predictive performance is indicated by the coefficient of prognosis (COP), with COP over 90%, and an overall surrogate error below 2%. Limitations include the focus on a single TRB type under purely axial loading and the use of simplified, isothermal, fully flooded assumptions that neglect surface roughness and thermal effects, which may limit transferability to radial or transient operating conditions. Focusing on similar outputs, such as film thickness and friction, Walker et al. [
59] conducted EHL simulations with the same goal—a data-driven alternative to repeated numerical evaluations of film formation and friction. Using a large dataset (DS4) generated from an LHD-sampled EHL model, the authors trained an ANN to map features—such as load, entrainment speed, reduced radius, Young’s modulus, pressure viscosity coefficient, reference viscosity, Poisson’s ratio, density, and contact length—over a broad design space to central film thickness viscous friction, and boundary friction. To ensure generalization and prevent overfitting, the authors employed early stopping and regularization techniques. The specified runtimes contrast with the numerical and surrogate evaluation: the EHL simulation requires about 5 s per case, while the trained surrogate provides predictions in about 3 ms. The analytical solution is only negligibly faster than the ANN, but significantly less accurate.
Chen et al. [
58] combined Kriging and LHD to create a surrogate that can be used by GA to optimize targets, similar to Wang et al. [
57], with a focus on reducing the turnaround time of CFD evaluations. The core of their study was to investigate and optimize the air-oil two-phase flow characteristics in planetary gear bearings of helicopters using under-race lubrication while minimizing energy dissipation. The study integrated experimental testing with advanced numerical modeling to analyze how various design and operating conditions (oil feed hole parameter, number of holes, inlet velocity, bearing revolution speed) influence internal lubrication performance, such as oil volume fraction and churning power loss. One limitation is that parameters such as load and heat, which affect lubrication properties, were not considered.
3.1.3. Dynamics
Whereas tribology research emphasizes contact mechanics and lubrication, studies focusing on dynamics utilize ML to accelerate the evaluation of rotor bearing responses, the modeling of cage motion, and the prediction of vibrations in environments characterized by repetitive dynamic simulations.
The research by Han et al. [
55] focused on the accurate identification of dynamic bearing parameters and unbalance information in rotor systems. Their Kriging surrogate model (DS2) provided an explicit mapping between the input parameters—including stiffness and damping coefficients, rotational speed—and the resulting unbalance vibration responses as targets. To locate the globally optimal solution within this surrogate space, different evolutionary optimization algorithms were used in the study. The underlying simulation model is based on Timoshenko beam finite elements to describe the motion of the elastic shaft, rigid disks, and bearings and is later verified experimentally. Despite the model’s efficiency, the authors identified limitations regarding noise sensitivity, particularly for damping coefficients, which exhibited higher identification errors compared to stiffness when 10% Gaussian noise was introduced. In comparison of the combined method Kriging Surrogate Model and Evolutionary Algorithm (KSMEA), the KSMEA is 11 to 12 times faster than EA without surrogates (DEA, PSO, and GA) while maintaining the smallest modeling error.
In a similar vein, Srinivas et al. [
56] focused on predicting vibration responses in rotor systems equipped with external dampers, such as squeeze film dampers and elastic rings. The methodology employed an ANN (DS3) architecture with four hidden layers, each with 256 nodes. The simulation data was generated using a 35-element Timoshenko beam FE model that incorporates material damping and structural flexibility of squirrel cages and ring dampers. The model utilizes the following input features: clearance, depth, unbalance amplitude and phase, oil viscosity, and rotational speed. The target was the bearing vibration response. Of 99 test data points, 91 had an error below 10%, indicating a high predictive accuracy of the surrogate model after hyperparameter tuning. A significant limitation acknowledged by the authors is the high variance in vibration magnitudes (ranging from 1 to 191 µm), which necessitated filtering the training data to a specific range (50–150 µm) to achieve convergence. Furthermore, the study currently relies on numerical data and lacks validation using data from experimental setups, which represents an important area for future research.
Schwarz et al. [
16,
20] examined the dynamics of inner bearings and provide an exemplary contribution to current research on cage dynamics in ACBB. They proposed a new classification method to identify operating regimes associated with cage instability. Their work combined multibody simulation in CABA3D with supervised ML methods to predict cage motion regimes and key performance indicators under a wide range of operating conditions and design variants. The cage was represented as an elastic FE structure reduced with the Craig and Bampton [
66] method, and a systematic sampling of parameters such as rotational speed, loads, torque, cage/rolling element geometries, and pocket and guiding clearances were used to generate a comprehensive database (DS4) of simulated cage responses. Based on this, classification models were trained to distinguish between stable, unstable, and circling cage motion [
20], while regression models predicted quantities such as friction torque, contact forces, cage acceleration, and components of a cage dynamics indicator (CDI) [
16]. The study employed quadratic discriminant analysis (QDA) for labeling simulation outcomes and an AdaBoostM1 ensemble classifier using decision trees as weak learners to classify CDI [
20]. In the regression study, ensemble learners such as Random Forest and XGBoost, as well as an ANN, were additionally evaluated within a hyperparameter tuning workflow using GA [
16]. The prediction models reached high quality, with the classifier achieving an accuracy of about ninety percent for the motion regime classification [
20] and regression models obtaining coefficients of determination between 0.75 and 0.94 for most targets [
16]. Despite achieving high prediction accuracy and a predictions that were 36,000 times faster than MBS, the authors acknowledge limitations, specifically regarding the reliance on nominal geometry models that neglect real-world manufacturing deviations, which could be the reason for small differences in experimental verification [
20]. Furthermore, the initial data generation remains computationally intensive, requiring approximately 10 h per simulation [
16], resulting in approximately 40,000 h for this dataset, which could be extended to other bearing types and input features.
A detailed overview of the eleven articles, sorted by year of publication, is provided in
Table 5.
Section 3.1 shows that the studies included are clustered according to simulation domain and tools, and that each domain is mainly associated with a set of its own target results.
Section 3.2 therefore shifts from applications to synthesis, summarizing the common methodological trends, recurring limitations, and open research gaps across all studies.
3.2. Gaps and Remaining Challenges in Machine Learning Assisted Bearing Simulations
It is evident that ML in rolling bearing simulations enables frontloading of engineering knowledge by quantifying the effects of variations in design parameters and operating conditions at early stages of development. Nevertheless, it should be noted that several studies were identified during the screening process in which rolling bearing design optimization was conducted without ML by applying evolutionary algorithms (EAs). EAs are population-based, stochastic search methods that generate iteratively improved candidate solutions through selection and variation operators (primarily crossover and mutation). The most important subparadigms include evolutionary strategies, evolutionary programming, and GA, which differ mainly in their representation and weighting of recombination versus mutation [
67]. These contributions typically optimize static or dynamic load ratings under analytical formulations used as constraints, often without executing numerical simulations [
68,
69,
70,
71,
72,
73,
74].
However, other approaches are required for industrial implementation of frontloading, as direct EA-based optimization cannot easily control simulation tools in a practical and efficient manner. In simulation-driven contexts, optimization is typically performed on surrogate models [
25,
55,
57,
58,
62], since directly embedding an optimizer in simulations is generally inefficient. This is particularly relevant because established rolling bearing manufacturers—who have the greatest influence on product optimization—rely on commercial or internal simulation programs whose predictive accuracy in relation to subsequent field tests is substantially higher than that of purely analytical equations. Examples of established bearing simulation tools that form the backbone of classical bearing design include BEARINX (Schaeffler), SimPro (SKF), Syber (Timken), and the Bearing Calculation Tool (NSK). In addition, smaller bearing manufacturers may employ commercial software packages such as MESYS, KISSsoft, Romax Spin, and FVA-Workbench. More detailed dynamic simulations are possible with tools such as BEAST [
75] (SKF), LaMBDA [
76] (FVA), and CABA3D [
16], although these simulations are computationally more expensive.
Against this background, even ML-supported workflows encounter practical limitations when coupled with simulation environments. Therefore, the following points address the remaining challenges and future potential with regard to RQ3:
Most of the studies included do not report on new experiments, but this is not necessarily a shortcoming, if the underlying simulation model has been independently validated and the ML model has been verified using the simulation results. In such cases, statistical analyses can still provide actionable trends that help make subsequent field trials more efficient and cost-effective.
From a ML perspective, several studies remain inadequately documented. Future work should report the surrogate modeling workflow in a reproducible manner, including data preprocessing, train/validation/test split, hyperparameter optimization strategy and final simulation settings, as well as measures to reduce overfitting. Moreover, the lack of consistent reporting on surrogate quality metrics limits comparability between studies.
Providing the extracted datasets—even a simple table with input features and targets (e.g., as a CSV)—along with a clear description of the simulation model would enable reuse across the community. Such availability would facilitate the extension of studies to additional bearing types and feature sets and enable systematic improvements of ML models among comparable simulation models.
Some studies already cover a wide range of operating conditions (e.g., load collectives, viscosity and friction parameters, and speed bands). Nevertheless, expanding the coverage—particularly toward boundary regimes—would improve the robustness of the conclusions and reduce the risk of errors when applying them in boundary regions.
Although surrogates reduce marginal evaluation costs, most workflows require a space-filling DoE to generate the training database in advance, so initial computational investments scale with both the number of samples and the runtime of the simulation. Efficiency could be improved through adaptive sampling and multi-fidelity modeling, whereby adaptive sampling preferably adds new simulation points in informative regions (e.g., high error or high sensitivity), and multi-fidelity models combine low-fidelity simulations with fewer high-fidelity runs to achieve target accuracy at lower overall cost.
While dimensional design parameters are commonly varied, manufacturing imperfections such as surface roughness and form deviations are rarely taken into account. This presents a clear opportunity to improve realism and strengthen the link between simulations and practical performance.
Nonlinear mappings are generally well captured, but interpretability remains an open question: In the case of strong nonlinearities, the same design parameter can have qualitatively different effects in different areas of the design space, complicating engineering interpretation and decision-making in the field of rolling bearings.
Beyond improving computational efficiency, future simulation and ML workflows should prioritize closer coupling between domains—particularly tribological and dynamic simulations—to enable integrated evaluation of efficiency, dynamic stability, and durability within a single simulation framework.
Addressing these aspects can further advance simulation-driven rolling bearing design by improving product efficiency while reducing development time. However, these benefits will only translate into industry if the underlying methods are communicated in a form that practitioners can reliably apply. Therefore, the following section presents a set of methodological building blocks designed to support the practical introduction of ML-supported work steps for rolling bearing simulations.
3.3. Methodological Approaches for Future Rolling Bearing Simulations
This section is intended for both academic researchers and industrial practitioners. It consolidates recurring methodological patterns into a set of practical recommendations that can be adapted to different bearing types, domains, and simulation tools.
Accordingly, the section describes (i) how simulation-driven ML use cases can be formulated in bearing design and optimization, (ii) how data mining strategies can be set up, (iii) how ML models can be selected and trained in a way that remains transparent and reproducible, (iv) how models can be validated using the underlying simulator and, if available, experimental evidence, and (v) how surrogates can be integrated into technical processes such as optimization, sensitivity analysis, uncertainty propagation, and design space exploration. Based on the included studies, these steps are particularly relevant when ML models are trained on outputs from rolling bearing simulation environments, where simulation runtimes are high and parameter interactions are often strongly nonlinear. The goal is to provide a compact toolbox that supports consistent implementation and helps avoid common pitfalls; see
Figure 5.
- (i)
The use case formulation is the step in which the technical problem is translated into a precise task, as these early decisions determine the possible modeling approach and the amount and type of data to be generated. The process starts with defining the scope of the simulation: Which domain shall be covered (tribology, bearing dynamics, rating life/load rating, system analysis) and should individual or coupled simulation tools be used? For rolling bearings, this choice is particularly important because the relevant target variables differ substantially between domains. Depending on the application, the ML model may be intended to predict friction-related quantities, film thickness, pressure distributions, dynamic responses, or rating life. Likewise, the dominant input variables may include design variations and operating conditions. Next, the intended objective of ML is formulated in measurable terms—predicting specific simulation results, enabling exploration of the design space, supporting uncertainty/sensitivity analyses, or performing optimizations—which directly implies whether the task is a regression or classification task and which performance metrics are meaningful. The input space is then defined by listing the features for the DoE, their bounds, constraints, and dimensionality, as these factors determine both the complexity of the model and sampling requirements. In rolling bearing simulations, this step should explicitly account for physical admissibility, since not all parameter combinations are meaningful or numerically stable; for example, combinations of load, speed, viscosity, or contact geometry may lead to unrealistic operating states or non-convergent simulations. Based on this dimensionality and the computational costs per simulation run, as well as all available experimental data, an expected sampling budget is determined using rules of thumb. Another important step is to define the tools for the ML process (e.g., Python™ or MATLAB®). Finally, the required credibility is explicitly defined by specifying whether the model must be verified using the simulator or additionally validated using experiments, including acceptable error limits for the intended use. At the same time, a streamlined documentation scheme is established to keep the workflow reproducible and extendable when variables and variants are retrained or added.
- (ii)
Space-filling DoE is the next step in creating an initial dataset. The goal is to cover the permissible input space as evenly as possible with a limited number of simulation runs. Typical options include LHDs and low discrepancy sequences such as Sobol or Halton, which provide good coverage in moderate dimensions. This initial sampling is typically performed before model training, as it creates a robust base dataset for the initial surrogate model and for identifying nonlinearities, interactions, and constraints. Adaptive sampling is not part of this step. It only becomes relevant after the initial surrogate model has been evaluated, when new samples are targeted at areas with high errors, high uncertainty, or high relevance to the engineering objective. For rolling bearings, this procedure is especially relevant because the response surfaces are often shaped by threshold-like transitions, such as regime changes in lubrication or contact-state changes under varying loads. Consequently, uniform sampling alone may overlook critical technical regions unless the parameter ranges are selected based on domain knowledge. For full automation, the ML workflow should be linked to the simulation from this step onwards.
- (iii)
The development of a traceable and reproducible ML model begins with data curation and harmonization; i.e., consolidation of all simulation runs into a single versioned dataset with consistent variable names, units, sign conventions, and data types, accompanied by a data dictionary. Non-converged or physically invalid simulation cases should not be removed. Instead, they should be flagged and documented, including the exact filtering rules and the number of samples affected, to maintain transparency and enable later sensitivity checks. The curated dataset is then processed through a deterministic feature engineering pipeline. Feature scaling should be explicitly defined and applied consistently (e.g., standardization or min–max scaling), with the adjusted scaling parameters stored to ensure that future training or inference operations use identical transformations. If the input space is high-dimensional or highly correlated, dimensionality reduction can be applied to improve model stability and sample efficiency; for example, through feature selection or principal component analysis. Finally, the setup of supervised learning must reflect whether it is a regression or classification task. Kriging, tree-based ensembles such as random forest or gradient boosting, and neural networks are suitable for regression. Typical options for classification are SVM, tree-based classifiers, and discriminant analysis, or neural classifiers. Hyperparameter optimization is performed using reproducible methods such as grid search, random search, Bayesian optimization, early stopping, or L1/L2 regularization with fixed random seeds and k-fold cross-validation. The search space, optimization algorithm, number of trials, and evaluation metric must be documented. The best configuration, including all adjusted parameters, is stored together with the model for reconstruction. For both types of tasks, the complete set of modeling decisions, including selected feature sets, preprocessing parameters, candidate model families, and all training configurations, should be documented with fixed random seeds so that the surrogate can be reconstructed and extended without ambiguities.
- (iv)
Validation must distinguish between verification using the simulator and validation using experimental evidence, where available. Verification checks whether the surrogate reproduces the simulator results across the defined design and operating range. To ensure unbiased evaluation, all studies should maintain a strict separation between training and evaluation by reserving a hold-out test set that is not used during model selection or hyperparameter optimization. The performance of the surrogate is then quantified using appropriate metrics—such as RMSE, MAPE, R2, and residual analysis for regression, or accuracy and confusion matrix for classification—supplemented by diagnostic checks that reveal systematic errors. If the errors are too high, an adaptive sampling process can optionally be applied to create new data in uncertain areas. If experimental data is available, validation extends the assessment to physical evidence by comparing simulator and surrogate predictions with measurements under matched conditions. This step requires transparent handling of measurement uncertainty, noise, and variability. Rather than aiming for a perfect match, experimental validation should provide information on whether deviations are within acceptable tolerance for the intended technical decision. If experiments are only possible to a limited extent, targeted sampling at representative operating points can still convey credibility, provided that the selection rationale and uncertainty bounds are reported. In rolling bearing research, this is particularly relevant because experimental validation is often expensive and limited to selected operating points, while the simulation models themselves may already contain simplifying assumptions.
- (v)
Integration of surrogates into design processes by using the surrogate as a fast, reusable predictor. In design space exploration, the surrogate can be sampled densely to map sensitivities and trade-offs between design parameters and operating conditions, thereby revealing feasible regions and dominant influencing factors at an early stage of development. For optimization GA, PSO or gradient-based methods allow for efficient solution of multi-objective problems directly on the surrogate. For sensitivity analysis, global methods such as Sobol indices are practical, as many evaluations are required. The surrogate provides these at negligible cost, allowing engineers to prioritize which loads, geometric tolerances, lubrication parameters, or speed ranges are most important and simplify the design problem accordingly. For uncertainty propagation, the surrogate supports Monte Carlo simulation or reliability analyses under variable inputs without unreasonable runtimes.
To transfer these capabilities to bearing design, the ML model should be integrated into existing industrial simulation tools as a decision-making aid: (1) Development of a validated surrogate model for the performance variables that are relevant to design decisions; (2) Use of screening to eliminate regions and reduce the dimensionality; (3) Performing optimization and robustness analyses to identify optimized designs under relevant operating conditions; and (4) Validating only the final shortlist with simulations. Once this workflow is established, a new paradigm called online learning [
77,
78,
79] could be introduced to maintain and extend the validity of the surrogate model over time by incrementally updating the models; for example, when additional operating conditions, new lubricants, or new bearing variants are evaluated. This workflow reduces turnaround time, increases robustness, and enables early, quantitative frontloading of performance knowledge for rolling bearings without the need for continuous, expensive simulation loops. Overall, this framework provides a practical entry point for integrating surrogates into rolling bearing design workflows, while remaining readily extensible to more sophisticated, problem-solving implementations.
4. Discussion
This systematic review aims to clarify how ML is currently used to support early-stage development of rolling bearings, when simulations are the primary source of data. In the 11 studies included, the evidence base remains small and heterogeneous in terms of domains, simulators, and reported evaluation practices, but it consistently shows that ML can frontload quantitative knowledge by making the investigation of design and operating variations computationally manageable.
RQ1 (Methods). The included works are dominated by regression surrogates trained on simulation data, reflecting the central objective of replacing repeated, costly simulation calls with fast predictors. Frequently used surrogate families include Kriging models, neural networks, and tree-based methods, with some studies benchmarking multiple candidates before selecting the best model for a given target quantity. A smaller subset addresses classification, notably for regime identification in bearing dynamics. The reporting on the methods suggests that model selection is largely determined by (i) the nonlinearity of the underlying input–output relationship, (ii) the size of the dataset and sampling strategy, and (iii) the intended downstream use; e.g., design space exploration or optimization loops. However, comparability between studies is limited by inconsistent reporting of preprocessing, data partitioning strategies, and hyperparameter optimization procedures, as well as non-harmonized performance metrics.
RQ2 (Domains and Results). The literature can be divided into three areas of application—contact mechanics, lubrication, and dynamics—each closely linked to the capabilities and limitations of the underlying simulation tools. In contact mechanics, FE simulation datasets are used to predict stress and pressure outputs, enabling efficient optimization and sensitivity workflows without repeated FE evaluations. Studies in the field of lubrication primarily address film thickness, pressure, friction, and churning power losses using EHL or CFD solvers. Dynamics studies apply ML to predict rotor bearing behavior and evaluate cage motion, with repeated dynamic simulations motivating surrogates for rapid conclusions in studies with many parameters. Overall, the mapping from domain to target output is not just a research preference, but a practical consequence of the architecture of simulation tools: tribological and dynamic workflows are typically implemented as separate pipelines, and tighter coupling remains the exception rather than the rule.
RQ3 (Trends, Gaps, and Needs). Several recurring challenges emerge as obstacles to broader application and greater scientific consolidation. First, many studies rely on simulation verification and provide only limited experimental grounding. While this may be acceptable if the simulator is independently validated, the chain of credibility should be explicitly stated. Second, documentation and reproducibility remain inconsistent: insufficient reporting of training/validation/testing protocols, hyperparameter tuning, and preprocessing decisions limits reusability and makes it difficult to interpret performance claims across domains. Third, limited data availability—often restricted to “upon request” or not provided at all—prevents benchmarking and can slow cumulative progress. Fourth, generalization beyond the operating conditions studied is rarely established. Extending coverage to boundary regimes is particularly important, as bearing performance often exhibits regime shifts that amplify extrapolation risk. Therefore, PINNs [
80] could improve the robustness of surrogate models, provided that the underlying physics is well understood and can be expressed by a suitable loss function. Fifth, the initial costs of dataset generation remain a central barrier: although surrogates reduce marginal evaluation costs, most workflows still require a space-filling DoE with potentially expensive simulation runs before benefits are realized. In this context, adaptive sampling and multi-fidelity strategies appear to be practical methodological developments for reducing overall costs while maintaining accuracy, especially with expensive simulators. Finally, realism gaps persist, as manufacturing deviations and non-ideal geometries are rarely represented, and interpretability under strongly nonlinear mappings remains an open technical question when parameter effects differ qualitatively across the design space.
Despite the limited number of studies, a coherent engineering message emerges: ML is most beneficial when it is treated as support for decisions around established simulation tools rather than a replacement for physics-based reasoning. In this role, surrogates enable dense evaluation for design space exploration, optimization, sensitivity analysis, and uncertainty propagation, while simulations remain essential for (i) generating training data, (ii) confirming designs, and (iii) maintaining confidence through regular verification and, where possible, validation. The methodological building blocks outlined in
Section 3.3 provide an essential foundation for credible and reproducible industrial investigations. However, detailed elaboration of specific application scenarios is still required.
Considering the input groups covered in the included studies, combining design and operating conditions provides the most flexible basis for simulations, as it reflects the variability found in practice. At the same time, a reduced input space under fixed operating conditions in defined industrial environments, such as continuous duty applications may be sufficient. Looking ahead, industrial-scale online learning represents a promising way to put these concepts into practice. In such a configuration, the server infrastructure could continuously generate informative simulation cases—expanding the feature space and extending coverage to additional bearing types—and incrementally update a shared surrogate that aggregates knowledge across domains and products. Over time, this evolving model could serve as a comprehensive, company-internal knowledge base for rolling bearing behavior, enabling rapid design-space exploration.
However, the technical capacity for early detection only translates into true frontloading if it is embedded in a functional organizational framework. As Fricke et al. [
15] emphasize, recognizing the need for change is insufficient without the corresponding “decision discipline”. Consequently, the potential of simulation-based insights is only realized if identified changes are translated into binding decisions and strictly implemented within the development process.
Limitations
In this systematic review, only peer-reviewed journal and conference publications indexed in Scopus, Web of Science, and IEEE Xplore were considered. Gray literature was not searched, so the review may not have captured all relevant evidence. In addition, the search strategy was operationalized using defined keywords and database-specific field settings, which may result in studies with atypical terminology being overlooked. Therefore, we performed a backward search, which did not identify any additional articles. The relatively small number of published studies reflects the early stage of research in this specific field. While this underlines the relevance and potential for future research, it also limits the ability to derive broadly generalizable conclusions. The findings should therefore be interpreted as indicative of emerging trends and methodological directions rather than as statistically robust or universally applicable results. Since the included studies are methodologically heterogeneous in terms of simulation accuracy, learning objectives, and reported metrics, a quantitative meta-analysis and a uniform assessment of bias scoring were not feasible; instead, the results were summarized narratively, which may increase the reliability of the qualitative interpretation. Data extraction was performed according to a predefined framework, but despite careful review, unintended extraction or classification errors cannot be completely ruled out. Any inaccuracies identified after publication will be transparently corrected in an updated supplementary file or, if necessary, in an erratum.
5. Conclusions
This systematic review addresses a largely unexplored area by focusing on the early integration of ML into rolling bearing simulation workflows. In contrast to existing studies, which predominantly focus on condition monitoring and tribological applications, this work provides a structured analysis of simulation-driven ML approaches in early design stages. The main contributions of this paper are threefold: (i) the identification and structuring of existing approaches across lubrication, contact mechanics, and dynamics; (ii) the derivation of a methodological framework for simulation-driven rolling bearing ML workflows; and (iii) the systematic identification of key research gaps and challenges that currently limit broader scientific consolidation and industrial application.
The analysis of the eleven included studies shows that ML is primarily used to approximate simulation outputs based on generated datasets, enabling efficient exploration of complex design spaces. The reviewed studies can be grouped into three main application domains—lubrication, contact mechanics, and dynamics—in which ML approaches are used to replace or enhance simulation processes, typically based on datasets generated through DoE. Across all domains, surrogate-based optimization emerges as the dominant use case. At the same time, the reviewed studies differ substantially in terms of simulation models, input parameter spaces, data generation strategies, and evaluation metrics. This heterogeneity currently limits comparability and prevents the derivation of generalizable guidelines for selecting appropriate ML approaches under specific conditions.
To address these challenges, this work proposes a structured methodological framework that supports the consistent implementation of ML workflows regarding rolling bearing simulations. The framework integrates key elements such as use case formulation for ML in rolling bearing simulations, the setup of data mining and sampling strategies, transparent and reproducible model development, systematic verification and validation using the underlying simulator and, where available, experimental evidence, as well as the integration of surrogate models into technical processes such as optimization, sensitivity analysis, uncertainty propagation, and design space exploration. In this way, it provides practical guidance for both academic research and industrial application, with the aim of improving transparency, reproducibility, and robustness in future studies.
The review further reveals several critical research gaps. These include insufficient methodological transparency in many studies, limited documentation of data generation and ML workflows, and a lack of standardized benchmarks for comparing different approaches. In addition, the influence of manufacturing deviations and real-world variability is rarely considered, although these factors can significantly affect the transferability of ML models to real applications. Furthermore, aspects such as boundary operating regimes and the coupling between tribological and dynamic simulation domains remain insufficiently addressed in the current simulation workflows. Addressing these issues is essential for advancing the reliability and applicability of ML in rolling bearing simulations.
From an industrial perspective, the integration of surrogate models into simulation workflows enables earlier and more extensive design space exploration, supports sensitivity and robustness analyses, and reduces computational effort in iterative development processes. This allows for a more efficient frontloading of performance knowledge in rolling bearing design. In this context, approaches such as online learning offer additional potential to continuously extend model capabilities by incorporating newly generated simulation data over time. In such a setting, simulation environments could iteratively generate additional, informative data points for newly introduced features or bearing types and integrate these into an evolving ML model. At the same time, it must be noted that ML models are inherently limited by the fidelity of the underlying simulation models on which they are trained. As simulation methods are further developed and refined, previously trained surrogate models may become outdated and require retraining to reflect the improved physical modeling. Online learning can provide a structured mechanism to continuously update and align surrogate models with evolving simulation capabilities.
Over time, the use of online learning could accumulate extensive knowledge about the behavior of rolling bearings and support its integration into company-internal AI systems. Engineers could then interact with such systems via intuitive interfaces to explore how design parameters and operating conditions influence the performance of previously trained rolling bearing configurations. Overall, this work provides a structured foundation for the systematic use of ML in rolling bearing simulations and highlights key directions for future research to enable more robust design and optimization workflows.