Power consumption was calculated as power consumed during milling process, from which was subtracted power consumption while the CNC machine was in idle mode. Cutting forces were more straightforward as the measuring tool has its own software that showed the entire experimental process with detailed graphs of any components of cutting forces. To work with data, random noise and volatility were reduced. This was done to reveal a clearer, more coherent signal for data analysis. All calculations were made by software, baring the exception of making it an absolute value, as depending on which side forces were applied one was negative. A negative value only meant that forces were applied in the other direction.
3.1. Statistical Evaluation
The statistical significance of machining parameters was evaluated using analysis of variance (ANOVA), a method originally developed by Fisher [
28] and widely applied in engineering experimental designs like Montgomery [
29] to determine the relative contribution of cutting parameters to response variables.
For purpose of statistical evaluation, we used the program STATISTICA version 14.0.0.15 (TIBCO Software Inc., Palo Alto, CA, USA). This program was chosen based on availability, technical understanding of software, utilities and personal preference.
Table 3,
Table 4,
Table 5,
Table 6 and
Table 7 summarise descriptive statistics of measured variables. The main tool chosen for analysis was ANOVA (analysis of variance), in order to determine statistically significant differences between the means of independent groups for forces
Fx,
Fy,
Fz, overall force
Fc, and cutting power
P. It is essential for comparing multiple groups simultaneously. For analysis purposes, we took all force components as absolute values, as during the experiment different combinations of feed speed (
f), rotation of milling tool (
n) and different types of milling (conventional, climb) were used.
Analysis of variance was conducted to examine the groups of data and if there are significant corelations. The Kolmogorov–Smirnov test of normality revealed that measured data are normally distributed and suitable for ANOVA.
Table 8 shows comparisons between the data acquired during the experiment. The highest influence of technological parameters was on the measured force
Fx.Figure 5 shows how the change of force
Fx is influenced by milling parameters. Climb milling had a component of this force greater than conventional. According to Duncan’s test, the influence of individual independent observed factors on this force component is statistically significant (probability of similarity
p < 0.05), except for the values at the feed speed of 6 and 10 m/min
−1 Figure 6 shows that with increase of feed speed, measured force
Fy increases; also, the force decreases with increased revolutions of the cutting tool. Here, on the contrary, climb milling had a smaller component of this force than conventional milling. Duncan’s test showed a statistically significant effect (
p < 0.05) of independent technological parameters on this force component.
Figure 7 shows how cutting parameters influence cutting force in direction
Fz; the graph shows all cutting parameters are in range of each other with little change less than 2 N. According to Duncan’s test, the influence of individual independent observed factors on this force component is statistically significant (
p < 0.05), except for the values at the feed speed of 6 and 8 m/min. All presented dependencies are related to the theory of wood composites milling, with individual components oriented according to the coordinate system of the force sensor. Their size is also influenced by the shape of the cutting edge of the tool and the change in the cross-section of the removed layer of material according to the type of milling. In terms of cutting energy specifically, the results can also be influenced by the material moisture and temperature. These were constant (8% and 22 °C) and standard for the selected machining conditions. The influence of vibrations was minimized by applying a low filter to the measured signal.
Figure 8 shows the dependence of computed force
Fc on technological parameters. Computed force increases with increase of feed speed and decreases with increase of rotation of tool. This fact is also true for theory of cutting, when cutting force decreases with increase in cutting speed due to decreasing thickness of removed material. This dependency shows that force does not change much with change of milling type, although Duncan’s test also showed a statistically significant effect of individual factors on this force.
Figure 9 shows the dependence of the consumed power required for material removal by the technological parameters of milling. With increasing speed of rotation of the tool, the consumed electrical power increases mainly due to higher currents flowing through the electrical winding of the electric spindle.
F-test determines whether differences between group means are statistically significant by comparing variability between groups to variability within groups. In ANOVA, SS (Sum of Squares) measures total variation, MS (Mean Square) is SS divided by its degrees of freedom, F is the ratio of between-group variance to within-group variance (MS between/MS within), and p is the probability of observing such an F value if there is actually no real difference between groups.
From the data in
Table 9, all examined factors have a statistically significant effect on the force
Fx. Spindle speed
n (rpm) demonstrates a strong influence (F = 284.98,
p < 0.001), while feed rate
f (m/min) also shows statistical significance (F = 6.00,
p = 0.003). Milling type exhibits the highest F value (F = 24,853.98,
p < 0.001), indicating that it is the dominant factor affecting
Fx. Based on the magnitude of the F statistics, milling type contributes most substantially to variations in Fy, followed by spindle speed, whereas feed rate has a comparatively smaller effect.
All investigated factors are highly statistically significant (
p < 0.001) for force
Fy according to results in
Table 10. Milling type yields the largest F value (F = 23,848.25), indicating the strongest influence on
Fy. Feed rate (F = 1145.58) and spindle speed (F = 883.48) also demonstrate pronounced and statistically significant effects. These results confirm that Fz is strongly dependent on both cutting conditions and the selected milling type.
The F-test results in
Table 11 reveal statistically significant effects of all three factors (
p ≤ 0.001) for parameters
Fz. Spindle speed shows the greatest impact (F = 360.90), followed by feed rate (F = 45.66). Although milling type is statistically significant (F = 27.38,
p = 0.000001), its effect is comparatively smaller. Therefore, Fz is predominantly influenced by spindle speed, with secondary contributions from feed rate and milling type.
In
Table 12, results indicate that spindle speed (F = 614.56,
p < 0.001), feed rate (F = 311.32,
p < 0.001), and milling type (F = 17.31,
p < 0.001) all significantly influence parameter Fc. Among these factors, spindle speed exhibits the strongest effect, followed by feed rate, while milling type contributes to a lesser extent. Therefore, Fc is predominantly governed by cutting conditions.
The F-test for power consumption in
Table 13 indicates that spindle speed (F = 210.13,
p < 0.001) and feed rate (F = 104.58,
p < 0.001) significantly influence the response. In contrast, milling type does not exhibit a statistically significant effect (F = 0.01,
p = 0.931). This suggests that parameter P is determined mainly by cutting conditions rather than by the type of milling process.
3.2. Predictive Modelling
Matlab (R2024b Update 5) Regression Learner was used as a tool for creation of predictive models. The main goal was to predict total force Fc. This program allows a wide selection of different model types, their training, and testing, and is then followed by comparison. Main parameters used for comparison are Root Mean Square Error (RMSE) and coefficient of determination R2. The parameters of RMSE of created models should be as close as possible to a value of 0. It is also important to compare RMSE values for the training and testing datasets. If trained and tested data have large differences between each other, it may indicate that the model has memorised the training data and is therefore unsuited for new data.
Measured data were randomly separated into two groups: trained (80%) and tested (20%) data. Data used as predictors for model training were spindle speed (n), feed speed (f) and milling type. The dependent variable was total force created by the milling tool during the process of material removal from workpieces. To choose the right model, initial training and testing were carried out using all available models in the Matlab environment, after which models were subsequently compared. The testing data originated from measured datasets obtained during experiments; this data is about 20% from total measured data. Data were randomly selected from each combination of machining parameters for testing data. Bayesian optimization as optimizer and expected improvement per second as acquisition function for 30 iterations were used for model optimisation.
After filtering clearly unsuitable models (RMSE ≥ 0.7 and R
2 ≤ 0.8), the remaining models were further compared with respect to their size in bytes. The size of the resulting models for Boosted Trees (Model No. 2.16) and Bagged Trees (Model No. 2.17) exceeded 18 × 10
4 bytes, while the other models were below 1.5 × 10
4 bytes, as shown in
Figure 10. With a larger dataset, this imbalance would be even more apparent, and for this reason models with this high requirement were removed from further analysis.
The filtered models (RMSE ≤ 0.7 and R
2 ≥ 0.8) are shown in
Figure 11, with their RMSE and R
2 being very similar.
The values of RMSE (Root Mean Square Error), MSE (Mean Square Error), R
2 (Coefficient of determination) and MAE (Mean Absolute Error) for both validation and test are in
Table 14. A 5-fold cross-validation method was used as a resampling technique for the model’s performance evaluation. The table also contains training time and model size. From their comparison, the Fine Tree model appears to be optimal, which has a low RMSE, short training time (3.88 s) and reasonable model size (8919 bytes).
Based on the comparison in
Table 14, the Fine Tree model (Model No. 2.5) was further analysed.
Figure 12 shows the predicted response compared to the true response. These have an approximate linear relationship, where their distance from the black line represents the prediction model error. Fine Trees are variants of decision tree models, which belong to the class of non-linear models used in supervised learning.
Figure 13,
Figure 14 and
Figure 15 show the comparison between measured and predicted values for individual predictors in box graph format. From a first look at the figures, the learned model predicts a value that is within the values of experimental measurements. The blue circle in the
Figure 13 is extreme value. There is a slight shift that may be a consequence of the amount of training data. By further expanding and learning, the model can decrease in inaccuracy. It should be emphasized that not all predictors and their change appeared to be statistically significant during ANOVA analysis.
3.3. Comparison of Residual Values for Validated and Tested Data
The residuals values should be randomly distributed around zero, without any visible pattern (such as a linear trend). If a pattern appears, it may indicate that the model has memorized the training data rather than learning general relationships, which limits its ability to perform well on new unseen data. The residual comparisons for the Fine Tree model (Model No. 2.5), analysed by individual predictors for the validated and tested values, are presented in Figures. It can be said that model validation was performed by comparing errors (residuals) for training and testing data. Red circles represent extreme values.
Figure 16 shows a comparison of validation residuals (a) and test residuals (b) for the RPM predictor. The distribution for validation residuals is around the zero value in the interval ±0.5 N; some extreme values are +1 and −1.5 N, respectively. The distribution for test residuals has more positive values, but also lies around a value of 0.5, i.e., corresponds to the validated residuals. This means that the model can be considered as suitably trained.
Figure 17 shows a comparison of validation residuals (a) and test residuals (b) for the feed speed predictor. The distribution for validation residuals is around the zero value in the interval ±0.5 N; some extreme values are +1.5 and −1.5 N, respectively. The test residuals are slightly shifted toward positive values; however, most of them also fall within ±0.5 N. Overall, the distribution of test residuals corresponds well to that of the validation residuals. This means that the model can be considered as suitably trained.
Figure 18 shows a comparison of validation residuals (a) and test residuals (b) for the direction of milling predictor. The distribution for validation residuals is around the zero value in the interval ±0.5 N; some extreme values are +1.5 and −1.5 N, respectively. The distribution for test residuals has more positive values, especially for climb milling. Nevertheless, most values remain within ±0.5 N, indicating that the distribution of test residuals is similar to that of the validation residuals. This means that the model can be considered as suitably trained.