Rather than predicting force fluctuations alone, the proposed model is effectively learning the energetic signature of experimentally measured wear transitions. The monotonic increase in cutting power with VB, particularly after the mid-life stage, reflects: (1) increased flank contact length; (2) enhanced ploughing contribution; (3) reduced effective rake angle; and (4) greater frictional dissipation.
Therefore, Pc_Mean is not merely a statistical descriptor, but a physically grounded proxy of wear progression.
3.2.1. Feature Relevance and Physical Interpretation
Figure 16 and
Figure 17 present the global feature importance ranking and SHAP-based interpretability analysis, respectively. The average cutting power (
) emerged as the dominant predictor, reflecting the direct coupling between mechanical energy consumption and tool wear progression. As wear advances, the frictional contact area increases, the effective rake angle degrades, and ploughing contributions intensify, leading to a systematic rise in cutting power. This relationship has been consistently reported in power-based tool monitoring studies and is strongly correlated with flank wear and edge rounding dynamics [
57,
58].
Skewness (
) ranked among the most influential descriptors, indicating that asymmetry in the power distribution captures non-stationary behavior associated with intermittent adhesion, chip segmentation instability, and micro-chipping events. Higher-order statistical moments have been widely recognized as sensitive indicators of nonlinear dynamics and incipient instabilities in machining processes, including chatter onset and tool degradation [
59,
60].
Amplitude-related metrics such as the maximum-to-minimum ratio (
) and harmonic mean (
) exhibited significant relevance. These features quantify abrupt excursions and low-energy intermittency in the signal, which may arise from transient friction events, localized thermal fluctuations, or periodic chip adhesion-detachment cycles. Similar observations have been reported in studies correlating envelope statistics and peak ratios with tool damage progression and cutting instability [
61,
62].
Trend-based descriptors, including the root mean square error of the fitted trend (
), capture the regularity of power evolution over time. Elevated values indicate nonlinear or unstable behavior, typically emerging during accelerated wear stages when diffusion, oxidation, and crater growth destabilize the tribological interface [
5,
32]. Lower-ranked features, such as trend intercept and cumulative absolute differences, appear to provide redundant information relative to more expressive global and distributional descriptors.
The convergence between XGBoost feature ranking and SHAP explanations enhances model transparency and supports physically interpretable decision-making. Explainable artificial intelligence (XAI) is increasingly regarded as a prerequisite for industrial certification, operator trust, and regulatory compliance in safety-critical manufacturing environments [
21,
22].
3.2.3. Implications for Intelligent Machining and Sustainability
The combination of high predictive performance, strong interpretability, and minimal sensor dependency positions the proposed framework as a viable candidate for edge-level implementation in intelligent machining systems. Ensemble models such as XGBoost are computationally efficient, memory-light, and well-suited for embedded inference, enabling real-time monitoring without reliance on cloud infrastructure or high-bandwidth data transfer [
19,
20].
From a sustainability perspective, compatibility with advanced cooling strategies such as MQL and ICT supports the transition toward low-fluid and closed-loop thermal management without compromising monitoring reliability [
48,
49]. Reliable early detection of tool degradation enables optimized tool utilization, reduced scrap generation, lower rework rates, and improved energy efficiency, aligning with Industry 4.0 and sustainable manufacturing paradigms [
53,
54].
To evaluate the performance of the model, several metrics were used, including accuracy, precision, recall, F1 Score, AUC-ROC (Area under the Receiver Operator Characteristic curve) and KS (Kolmogorov–Smirnov).
Table 7 shows the results found from the application of XGBoost followed by an RFECV.
| | Accuracy | Precision | Recall | F1 | AUC-ROC | KS |
| Training | 0.9647 | 0.92913 | 0.83098 | 0.8772 | 0.9877 | 0.8878 |
| Test | 0.95896 | 0.91667 | 0.804878 | 0.8571 | 0.9925 | 0.9403 |
| Validation | 0.93283 | 0.761905 | 0.80000 | 0.7804 | 0.9592 | 0.84738 |
Although accuracy is not the most appropriate metric in isolation to evaluate the performance of models in unbalanced problems, the results obtained were quite high in all the sets evaluated: 96.47% in training, 95.90% in the test and 93.28% in validation. These numbers indicate that the model was successful in correctly classifying most examples, even when exposed to data that was not used during training. The slight performance reduction in the test and validation sets is expected and shows that the model is not overly adjusted to the training set.
Precision, which measures the proportion of true positives among positive predictions, was 92.91 percent in training, 91.67 percent in the test, and 76.19 percent in validation. The steeper drop in validation may indicate that the model had greater difficulty maintaining the quality of positive predictions in completely new data, which may be a result of greater complexity or variability in these data.
Revocation, which represents the ability to identify all positive cases, remained more stable among the groups: 83.10% in training, 80.49% in the test and 80.00% in validation. These results suggest that the model maintains good sensitivity even outside the training set, managing to capture a consistent proportion of true positives.
The F1 score, a metric that represents the balance between accuracy and recall, was 87.73% in training, 85.71% in the test and 78.05% in validation. These values reflect solid performance in the training and test sets, with a more significant reduction in validation, consistent with the observed behavior in accuracy.
The AUC-ROC metric, which evaluates the model's ability to distinguish between positive and negative classes regardless of the decision threshold, showed high values: 98.78% in training, 99.26% in the test, and 95.92% in validation. Even with a slight drop in validation, the value remains above 95.0%, which demonstrates an excellent capacity for discrimination in all sets.
Finally, the KS (Kolmogorov–Smirnov), which measures the separation between the distributions of the positive and negative classes, also had significant performances: 0.8879 in training, 94.04% in the test and 84.74% in validation. This indicates that the model continues to be able to distinguish well between classes, even in more challenging validation environments.
The graph in
Figure 16 shows the importance of the model variables in order of importance, and the graph in
Figure 17 shows the Shapley Additive Explanations
(SHAP) diagram of the variables.
The analysis of the graphs reveals that the Pc_Mean variable, i.e., the average of the values captured during the machining process, was consistently the most relevant variable both in terms of global importance (XGBoost) and individual impact (SHAP values). SHAP values are a tool based on game theory that allows you to interpret machine learning models by attributing to each variable a local contribution to the prediction. This approach shows not only how much each variable influences the model’s output (magnitude), but also in which direction it impacts the prediction. In the case analyzed, as an example, high values of Pc_Mean and Pc_Skewness (features importances of XGBoost) are associated with increases in prediction (represented in red in the SHAP graphs), while low values (in blue) tend to reduce it.
The graph in
Figure 16 reveals that the Pc_Mean variable, i.e., the average of the values captured during the machining process, was consistently the most relevant variable both in terms of global importance (XGBoost) and individual impact (SHAP values).
Asymmetry (Pc_Skewness) also appears with high relevance. Asymmetry measures the symmetry of the distribution of time series values and is associated with anomalous vibration patterns or non-uniform variations in the shear load. High asymmetry values indicate distortions in process stability, which can be predictive of tool defects or degradation [
59].
The variables Pc_MaxMin_Ratio and Pc_Harmonic_Mean appear with significant contributions. The first quantifies the relative amplitude of the signal, and the second smooths the impact of extreme values, favoring the stability of time series with less variability. These parameters indicate rapid or abrupt variations in the process, which can often be associated with sudden changes in the machining of the material or failures in coolant.
The Pc_Trend_RMSE metric, which evaluates the mean square error of the trend, represents the regularity of the evolution of the process over time. High values indicate unpredictable, nonlinear fluctuations. Variables such as Pc_Trend_Intercept and Pc_AbsDiff_Sum showed a lower relative contribution. The first represents the starting point of the linear trend of the process, and its low importance suggests that the initial value of the time series does not contribute so much to the prediction of failures or deviations in the process. The second reflects the sum of the absolute differences between the consecutive points and may indicate local variability, but apparently, even though this variable was important, this metric had less influence on the model, perhaps because it is redundant compared to other, more global metrics such as skewness and RMSE.
As a conclusion, it was observed that the model is very sensitive to cutting forces, and the increase in the average (Pc_Mean) is strongly associated with a possible end of tool life, which makes this variable a critical indicator for continuous monitoring of the process. In addition, characteristics such as asymmetry (Pc_Skewness) and the ratio between maxima and minima (Pc_MaxMin_Ratio) also demonstrate high relevance, indicating that distortions in the distribution of data and abrupt variations in signal amplitude are potential signs of wear or instability. Variables such as Pc_Trend_RMSE and Pc_Harmonic_Mean reinforce this interpretation by pointing out that unpredictable fluctuations and smoothed extreme oscillations are equally relevant to predict anomalies. On the other hand, metrics such as Pc_Trend_Intercept and Pc_AbsDiff_Sum had a slighter impact, suggesting that the initial value of the trend and the local variation between consecutive points are not determinants in isolation, and are probably covered by broader variables. Thus, the model indicates that the stability and strength of the force signal are the main parameters to accurately estimate the behavior of the tool over time.
3.2.4. Industrial Implementation Considerations
Cutting power is attractive for industrial tool-condition monitoring because it can often be obtained directly from machine-tool controllers as spindle power or drive current, eliminating the need for dynamometers or additional multi-sensor hardware. In a practical deployment, power is sampled continuously and processed using a sliding window (e.g., 1 s windows updated at a chosen step size), from which the same statistical and trend features used in this study are computed and passed to the trained classifier for real-time inference. Tool-change decisions can be based either on the predicted class label or on the predicted probability of the worn class, enabling conservative early-warning thresholds that reduce the risk of missed detections. The present model is formulated as a binary end-of-life classifier using the adopted wear criterion (VB ≥ 0.6 mm). Extension to intermediate thresholds (e.g., VB = 0.3 mm) can be achieved by relabeling the dataset into multiple wear bands or reformulating the model as a regression on VB, while retaining the same feature pipeline. Changes in cutting parameters (Vc, f, ap) are expected to shift absolute power levels; therefore, broader multi-parameter validation (e.g., DOE-based expansion) and/or normalization/adaptation strategies are identified as necessary steps for future industrial-scale generalization.
From an industrial implementation perspective, the reliance on cutting power eliminates the need for intrusive and costly sensors, such as piezoelectric dynamometers, which are often impractical for shop-floor production. The cutting power signal can be acquired directly via the machine tool’s PLC or non-invasive current sensors. This compatibility, combined with the low computational overhead of the XGBoost model, facilitates the deployment of ‘Edge AI’ solutions directly on the machine controller. Consequently, the system supports real-time monitoring, ensuring that tool condition is continuously assessed with minimal latency, allowing for immediate reaction to safeguard the workpiece and machine integrity.