Author Contributions
Conceptualization, A.R.S.; Methodology, A.R.S.; Software, P.L.; Validation, P.L., A.M.L. and A.R.S.; Formal analysis, P.L. and J.M.; Investigation, P.L., J.M., F.P. and A.R.S.; Data curation, P.L. and A.R.S.; Writing—original draft, A.R.S.; Writing—review and editing, J.M., F.P., A.M.L. and A.R.S.; Visualization, P.L., F.P., A.M.L. and A.R.S.; Supervision, J.M., F.P., A.M.L. and A.R.S.; Project administration, A.R.S.; Funding acquisition, A.R.S. All authors have read and agreed to the published version of the manuscript.
Figure 1.
Workflow of the proposed six-step booming-noise detection methodology.
Figure 1.
Workflow of the proposed six-step booming-noise detection methodology.
Figure 2.
Tachometer captured RPM vs. time signal for vehicle 2.
Figure 2.
Tachometer captured RPM vs. time signal for vehicle 2.
Figure 3.
Microphone captured signal for vehicle 2 (left side).
Figure 3.
Microphone captured signal for vehicle 2 (left side).
Figure 4.
Microphone captured signal for vehicle 2 (right side).
Figure 4.
Microphone captured signal for vehicle 2 (right side).
Figure 5.
Time–frequency representation captured at the driver’s left ear in vehicle 1, using a binaural microphone setup.
Figure 5.
Time–frequency representation captured at the driver’s left ear in vehicle 1, using a binaural microphone setup.
Figure 6.
Extracted engine orders from vehicle 1 in-cabin recordings.
Figure 6.
Extracted engine orders from vehicle 1 in-cabin recordings.
Figure 7.
Order spectrum derived from vehicle 3 measurements.
Figure 7.
Order spectrum derived from vehicle 3 measurements.
Figure 8.
Broadband noise component in vehicle 1 post decomposition.
Figure 8.
Broadband noise component in vehicle 1 post decomposition.
Figure 9.
Broadband noise component in vehicle 3 post decomposition.
Figure 9.
Broadband noise component in vehicle 3 post decomposition.
Figure 10.
Example mission profiles generated from time–RPM anchor points defined in
Table 3.
Figure 10.
Example mission profiles generated from time–RPM anchor points defined in
Table 3.
Figure 11.
Comparison of original and modified second-order curves with injected low-frequency modulations. (a) vehicle 1; (b) vehicle 2; (c) vehicle 3.
Figure 11.
Comparison of original and modified second-order curves with injected low-frequency modulations. (a) vehicle 1; (b) vehicle 2; (c) vehicle 3.
Figure 12.
Modified second-order amplitude curves after applying the Hann window for booming simulation. (a) vehicle 1; (b) vehicle 2; (c) vehicle 3.
Figure 12.
Modified second-order amplitude curves after applying the Hann window for booming simulation. (a) vehicle 1; (b) vehicle 2; (c) vehicle 3.
Figure 13.
Smoothed order 2 profiles without added low-frequency randomness. (a) vehicle 1; (b) vehicle 2; (c) vehicle 3.
Figure 13.
Smoothed order 2 profiles without added low-frequency randomness. (a) vehicle 1; (b) vehicle 2; (c) vehicle 3.
Figure 14.
Smoothed second-order curves incorporating random low-frequency modulations. (a) vehicle 1; (b) vehicle 2; (c) vehicle 3.
Figure 14.
Smoothed second-order curves incorporating random low-frequency modulations. (a) vehicle 1; (b) vehicle 2; (c) vehicle 3.
Figure 15.
Smoothed second-order curves with low-frequency modulations, following Hann window-based amplitude shaping. (a) vehicle 1; (b) vehicle 2; (c) vehicle 3.
Figure 15.
Smoothed second-order curves with low-frequency modulations, following Hann window-based amplitude shaping. (a) vehicle 1; (b) vehicle 2; (c) vehicle 3.
Figure 16.
Decomposed contributions of synthesized engine orders and broadband noise to the total sound pressure.
Figure 16.
Decomposed contributions of synthesized engine orders and broadband noise to the total sound pressure.
Figure 17.
Reconstructed pressure amplitude profile of the full vehicle sound after synthesis.
Figure 17.
Reconstructed pressure amplitude profile of the full vehicle sound after synthesis.
Figure 18.
Spectrogram representations of synthesized full vehicle sound profiles.
Figure 18.
Spectrogram representations of synthesized full vehicle sound profiles.
Figure 19.
CNN architecture showing feature extraction and classification stages.
Figure 19.
CNN architecture showing feature extraction and classification stages.
Figure 20.
Accuracy and loss trends for training with 1250 data samples.
Figure 20.
Accuracy and loss trends for training with 1250 data samples.
Figure 21.
Evolution of training and validation accuracy and loss using an expanded dataset of 12,500 examples.
Figure 21.
Evolution of training and validation accuracy and loss using an expanded dataset of 12,500 examples.
Figure 22.
Minimum, maximum and average test accuracy values for different numbers of convolutional layers.
Figure 22.
Minimum, maximum and average test accuracy values for different numbers of convolutional layers.
Figure 23.
Training and validation accuracy and loss curves for the binary (Yes/No) classification task over 30 epochs.
Figure 23.
Training and validation accuracy and loss curves for the binary (Yes/No) classification task over 30 epochs.
Figure 24.
Test accuracy statistics (minimum, maximum, mean) across different convolutional depths for the YesSevere/YesMild/No classification.
Figure 24.
Test accuracy statistics (minimum, maximum, mean) across different convolutional depths for the YesSevere/YesMild/No classification.
Figure 25.
Learning curves showing training and validation accuracy and loss for the three-class classification task using 30 epochs.
Figure 25.
Learning curves showing training and validation accuracy and loss for the three-class classification task using 30 epochs.
Figure 26.
Scatter plot of 30 misclassified test cases for Network 4 (Yes/No classification, smoothed data).
Figure 26.
Scatter plot of 30 misclassified test cases for Network 4 (Yes/No classification, smoothed data).
Figure 27.
Distribution of misclassified Network 4 samples by RPM range.
Figure 27.
Distribution of misclassified Network 4 samples by RPM range.
Figure 28.
Scatter plot of 30 misclassified test cases for Network 7 (YesSevere/YesMild/No classification, non-smoothed data).
Figure 28.
Scatter plot of 30 misclassified test cases for Network 7 (YesSevere/YesMild/No classification, non-smoothed data).
Figure 29.
Distribution of misclassified Network 7 samples by RPM range.
Figure 29.
Distribution of misclassified Network 7 samples by RPM range.
Table 1.
Comparison of the most relevant related works.
Table 1.
Comparison of the most relevant related works.
| Article | Main Contribution | Advantages Over This Work | Limitations Compared to This Work |
|---|
| Kim et al. (2023) [14] | FE and modal analysis of tailgate-induced booming in EVs. | Provides a physically validated approach for structural improvements. | No AI; limited to structural methods and EV context. |
| Altinsoy (2022) [15] | Psychoacoustic study of booming perception in ICE/HEVs based on listening tests. | Strong perceptual grounding; real-world relevance. | Requires human evaluation; no automation or classification. |
| Song et al. (2025) [16] | Big data + TPA/OTPA framework for NVH optimization in vehicle production. | Production-scale applicability; sensor fusion. | Still dependent on large sensor setups; no classification approach. |
| Chu et al. (2023) [17] | CNN with MFCC and augmentation for environmental sound classification. | Lightweight architecture; high accuracy with little data. | Uses generic datasets; not tailored for vehicle interior sounds |
| Tsalera et al. (2021) [18] | Transfer learning with pre-trained CNNs for audio classification (GoogLeNet, YAMNet). | Efficient training; flexible use of vision models. | Not domain-specific; tested mostly on urban and domestic sounds. |
| Souza et al. (2024) [19] | MBD + ANN to predict tonal sound prominence in EV powertrains. | Bypasses need for real measurements; ideal for early design. | Focus on tonal sounds in EVs; no classification of booming phenomena. |
Table 2.
Key parameters of the test track layout used during data acquisition.
Table 2.
Key parameters of the test track layout used during data acquisition.
| Parameter | Value | Unit | Notes |
|---|
| Surface type | Asphalt | — | Uniform across the full circuit |
| Straight section length (each) | 400 × 2 | m | Two identical straights |
| Lane length—Lane 1/2/3 | 2074/2097/2120 | m | Measured along central line |
| Lane width—Lane 1/2/3 | 3.75/3.75/4.00 | m | Lane 3 slightly wider for high-speed tests |
| Curve radius—North/South | 186.5/113.5 | m | North is wider than south |
| Max cornering speed—Lane 3 (north/south curve) | 117/96 | km/h | Under free cornering conditions |
Table 3.
Input anchor points for mission profile generation.
Table 3.
Input anchor points for mission profile generation.
| Phase | Time [s] | Engine Speed [RPM] |
|---|
| Start | 0.0 | 1.05 |
| Finish | 6.5 | 4.57 |
Table 4.
Limits of parameters G and ΔRPM used in second-order amplitude modulation with a Hann window.
Table 4.
Limits of parameters G and ΔRPM used in second-order amplitude modulation with a Hann window.
| Parameter | Range [min, max] |
|---|
| G | [0.0, 1.5] |
| ΔRPM | [0, 560] RPM |
Table 5.
CNN architecture and training parameters.
Table 5.
CNN architecture and training parameters.
| Component | Configuration |
|---|
| Input | Grayscale spectrogram images |
| Convolutional layers | 6 convolutional layers |
| Filters | Increasing depth across layers following a doubling scheme |
| Kernel size | 3 × 3 |
| Activation function | ReLU |
| Padding | Same |
| Pooling | Max pooling (stride = 2) |
| Regularization | Dropout (0.5) and L2 regularization (λ = 0.0001) |
| Fully connected layer | 1 dense layer |
| Output layer | Softmax (3-class)/Sigmoid (binary) |
| Loss function | Categorical cross-entropy/Binary cross-entropy |
| Optimizer | Stochastic Gradient Descent with momentum |
| Learning rate | Initial value with decay applied during training |
| Batch size | 32 |
| Number of epochs | 50 |
| Validation strategy | Hold-out validation set (separate from training data) |
| Framework | MATLAB 2019 |
Table 6.
Selected hyperparameters and optimization settings for CNN training.
Table 6.
Selected hyperparameters and optimization settings for CNN training.
| Category | Parameter | Value |
|---|
| Optimization | Optimizer | SGD with momentum |
| Momentum | 0.9 |
| Initial learning rate | 0.001 |
| Decay schedule | ×0.1 every 10 epochs |
| Mini-batch size | 32 |
| Architecture | Convolution kernel | 3 × 3 |
| Filters per layer nn | 8 × 2 (n − 1) |
| Padding | Same |
| Convolution stride | 1 |
| Max-pooling (kernel, stride) | 2 × 2, stride 2 |
| Regularization | Dropout | 50% |
| L2 coefficient (λ) | 0.0001 |
Table 7.
Accuracy results for smoothed and non-smoothed test cases across CNN architectures.
Table 7.
Accuracy results for smoothed and non-smoothed test cases across CNN architectures.
| Network ID | Convolutional Layers | Epochs | Accuracy (Smoothed Test Set) | Accuracy (Non-Smoothed Test Set) |
|---|
| 1 | 3 | 10 | 94.20% | 84.51% |
| 2 | 4 | 10 | 94.76% | 84.78% |
| 3 | 5 | 10 | 95.72% | 86.11% |
| 4 | 6 | 10 | 96.20% | 85.25% |
Table 8.
Test accuracy of Network 3 on non-smoothed cases, broken down by vehicle.
Table 8.
Test accuracy of Network 3 on non-smoothed cases, broken down by vehicle.
| Network ID | Overall Accuracy (Non-Smoothed) | Vehicle 1 | Vehicle 2 | Vehicle 3 |
|---|
| 3 | 86.11% | 81.33% | 93.37% | 83.73% |
Table 9.
Accuracy results for ternary classification on smoothed vs. non-smoothed cases.
Table 9.
Accuracy results for ternary classification on smoothed vs. non-smoothed cases.
| Network ID | Convolutional Layers | Epochs | Accuracy (Smoothed Test Set) | Accuracy (Non-Smoothed Test Set) |
|---|
| 6 | 6 | 10 | 93.44% | 70.34% |
| 7 | 7 | 10 | 92.92% | 78.16% |
Table 10.
Classification accuracy (%) of Network 7 on raw (unsmoothed) datasets.
Table 10.
Classification accuracy (%) of Network 7 on raw (unsmoothed) datasets.
| Dataset | Vehicle 3 | Vehicle 2 | Vehicle 1 | Overall |
|---|
| Accuracy | 75.10 | 90.55 | 69.02 | 78.20 |
Table 11.
Regression-based performance indicators for the CNN model.
Table 11.
Regression-based performance indicators for the CNN model.
| Metric | Description |
|---|
| Overall accuracy | Percentage of samples whose predicted booming intensity falls within the correct perceptual range after thresholding |
| Training loss | Mean regression loss during training iterations |
| Validation loss | Mean regression loss on validation data |
| Convergence stability | Consistency of accuracy and loss across epochs |