Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (12)

Search Parameters:
Keywords = Mel time–frequency spectrum

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
11 pages, 1725 KB  
Article
Tool Wear Detection in Milling Using Convolutional Neural Networks and Audible Sound Signals
by Halil Ibrahim Turan and Ali Mamedov
Machines 2026, 14(1), 59; https://doi.org/10.3390/machines14010059 - 2 Jan 2026
Cited by 4 | Viewed by 1385
Abstract
Timely tool wear detection has been an important target for the metal cutting industry for decades because of its significance for part quality and production cost control. With the shift toward intelligent and sustainable manufacturing, reliable tool-condition monitoring has become even more critical. [...] Read more.
Timely tool wear detection has been an important target for the metal cutting industry for decades because of its significance for part quality and production cost control. With the shift toward intelligent and sustainable manufacturing, reliable tool-condition monitoring has become even more critical. One of the main challenges in sound-based tool wear monitoring is the presence of noise interference, instability and the highly volatile nature of machining acoustics, which complicates the extraction of meaningful features. In this study, a Convolutional Neural Network (CNN) model is proposed to classify tool wear conditions in milling operations using acoustic signals. Sound recordings were collected from tools at different wear stages under two cutting speeds, and Mel-Frequency Cepstral Coefficients (MFCCs) were extracted to obtain a compact representation of the short-term power spectrum. These MFCC matrices enabled the CNN to learn discriminative spectral patterns associated with wear. To evaluate model stability and reduce the effects of algorithmic randomness, training was repeated three times for each cutting speed. For the 520 rpm dataset, the model achieved an average validation accuracy of 96.85 ± 2.07%, while for the 635 rpm dataset it achieved 93.69 ± 2.07%. The results demonstrate the feasibility of using acoustic signals, despite inherent noise challenges, as a complementary approach for identifying suitable tool replacement intervals in milling. Full article
(This article belongs to the Special Issue Intelligent Tool Wear Monitoring)
Show Figures

Figure 1

21 pages, 7017 KB  
Article
Multi-Scale Frequency-Adaptive-Network-Based Underwater Target Recognition
by Lixu Zhuang, Afeng Yang, Yanxin Ma and David Day-Uei Li
J. Mar. Sci. Eng. 2024, 12(10), 1766; https://doi.org/10.3390/jmse12101766 - 5 Oct 2024
Cited by 4 | Viewed by 1655
Abstract
Due to the complexity of underwater environments, underwater target recognition based on radiated noise has always been challenging. This paper proposes a multi-scale frequency-adaptive network for underwater target recognition. Based on the different distribution densities of Mel filters in the low-frequency band, a [...] Read more.
Due to the complexity of underwater environments, underwater target recognition based on radiated noise has always been challenging. This paper proposes a multi-scale frequency-adaptive network for underwater target recognition. Based on the different distribution densities of Mel filters in the low-frequency band, a three-channel improved Mel energy spectrum feature is designed first. Second, by combining a frequency-adaptive module, an attention mechanism, and a multi-scale fusion module, a multi-scale frequency-adaptive network is proposed to enhance the model’s learning ability. Then, the model training is optimized by introducing a time–frequency mask, a data augmentation strategy involving data confounding, and a focal loss function. Finally, systematic experiments were conducted based on the ShipsEar dataset. The results showed that the recognition accuracy for five categories reached 98.4%, and the accuracy for nine categories in fine-grained recognition was 88.6%. Compared with existing methods, the proposed multi-scale frequency-adaptive network for underwater target recognition has achieved significant performance improvement. Full article
(This article belongs to the Section Ocean Engineering)
Show Figures

Figure 1

18 pages, 3164 KB  
Article
Cough Detection Using Acceleration Signals and Deep Learning Techniques
by Daniel Sanchez-Morillo, Diego Sales-Lerida, Blanca Priego-Torres and Antonio León-Jiménez
Electronics 2024, 13(12), 2410; https://doi.org/10.3390/electronics13122410 - 20 Jun 2024
Cited by 8 | Viewed by 5174
Abstract
Cough is a frequent symptom in many common respiratory diseases and is considered a predictor of early exacerbation or even disease progression. Continuous cough monitoring offers valuable insights into treatment effectiveness, aiding healthcare providers in timely intervention to prevent exacerbations and hospitalizations. Objective [...] Read more.
Cough is a frequent symptom in many common respiratory diseases and is considered a predictor of early exacerbation or even disease progression. Continuous cough monitoring offers valuable insights into treatment effectiveness, aiding healthcare providers in timely intervention to prevent exacerbations and hospitalizations. Objective cough monitoring methods have emerged as superior alternatives to subjective methods like questionnaires. In recent years, cough has been monitored using wearable devices equipped with microphones. However, the discrimination of cough sounds from background noise has been shown a particular challenge. This study aimed to demonstrate the effectiveness of single-axis acceleration signals combined with state-of-the-art deep learning (DL) algorithms to distinguish intentional coughing from sounds like speech, laugh, or throat noises. Various DL methods (recurrent, convolutional, and deep convolutional neural networks) combined with one- and two-dimensional time and time–frequency representations, such as the signal envelope, kurtogram, wavelet scalogram, mel, Bark, and the equivalent rectangular bandwidth spectrum (ERB) spectrograms, were employed to identify the most effective approach. The optimal strategy, which involved the SqueezeNet model in conjunction with wavelet scalograms, yielded an accuracy and precision of 92.21% and 95.59%, respectively. The proposed method demonstrated its potential for cough monitoring. Future research will focus on validating the system in spontaneous coughing of subjects with respiratory diseases under natural ambulatory conditions. Full article
Show Figures

Figure 1

17 pages, 5633 KB  
Article
Audio General Recognition of Partial Discharge and Mechanical Defects in Switchgear Using a Smartphone
by Dongyun Dai, Quanchang Liao, Zhongqing Sang, Yimin You, Rui Qiao and Huisheng Yuan
Appl. Sci. 2023, 13(18), 10153; https://doi.org/10.3390/app131810153 - 9 Sep 2023
Cited by 4 | Viewed by 2159
Abstract
Mechanical defects and partial discharge (PD) defects can appear in the indoor switchgear of substations or distribution stations, making the switchgear a safety hazard. However, traditional acoustic methods detect and identify these two types of defects separately, ignoring the general recognition of audio [...] Read more.
Mechanical defects and partial discharge (PD) defects can appear in the indoor switchgear of substations or distribution stations, making the switchgear a safety hazard. However, traditional acoustic methods detect and identify these two types of defects separately, ignoring the general recognition of audio signals. In addition, the process of using testing equipment is complex and costly, which is not conducive to timely testing and widespread application. To assist technicians in making a quick preliminary diagnosis of defect types for switchgear, improve the efficiency of the subsequent overhaul, and reduce the cost of detection, this paper proposes a general audio recognition method for identifying defects in switchgear using a smartphone. Using this method, we can analyze and identify audio and video files recorded with smartphones and synchronously distinguish background noise, mechanical vibration, and PD audio signals, which have good applicability within a certain range. When testing the feasibility of using smartphones to identify three types of audio signal, through characterizing 12 sets of live audio and video files provided by technicians, it was found that there were similarities and differences in these characteristics, such as the autocorrelation, density, and steepness of the waveforms in the time domain, and the band energy and harmonic components of the frequency spectrum, and new combinations of features were proposed as applicable. To compare the recognition performance for features in the time domain, frequency band energy, Mel-frequency cepstral coefficient (MFCC), and this method, feature vectors were input into a support vector machine (SVM) for a recognition test, and the recognition results showed that the the present method had the highest recognition accuracy. Finally, a set of mechanical defects and PD defects were set up for a switchgear, for practical verification, which proved that this method was general and effective. Full article
Show Figures

Figure 1

16 pages, 4388 KB  
Article
Automatic COVID-19 Detection from Cough Sounds Using Multi-Headed Convolutional Neural Networks
by Wei Wang, Qijie Shang and Haoyuan Lu
Appl. Sci. 2023, 13(12), 6976; https://doi.org/10.3390/app13126976 - 9 Jun 2023
Cited by 4 | Viewed by 2655
Abstract
Novel coronavirus disease 2019 (Corona Virus Disease 2019, COVID-19) is rampant all over the world, threatening human life and health. Currently, the detection of the presence of nucleic acid from SARS-CoV-2 is mainly based on the nucleic acid test as the standard. However, [...] Read more.
Novel coronavirus disease 2019 (Corona Virus Disease 2019, COVID-19) is rampant all over the world, threatening human life and health. Currently, the detection of the presence of nucleic acid from SARS-CoV-2 is mainly based on the nucleic acid test as the standard. However, this method not only takes up a lot of medical resources but also takes a long time to achieve detection results. According to medical analysis, the surface protein of the novel coronavirus can invade the respiratory epithelial cells of patients and cause severe inflammation of the respiratory system, making the cough of COVID-19 patients different from that of healthy people. In this study, the cough sound is used as a large-scale pre-screening method before the nucleic acid test. Firstly, the Mel spectrum features, Mel Frequency Cepstral Coefficients, and VGG embeddings features of cough sound are extracted and oversampling technology is used to balance the dataset for classes with a small number of samples. In terms of the model, we designed multi-headed convolutional neural networks to predict audio samples, and adopted an early stop method to avoid the over-fitting problem of the model. The performance of the model is measured by the binary cross-entropy loss function. Our model performs well on the dataset of the AICovidVN 115M challenge that its accuracy rate is 98.1%, and on the dataset of the University of Cambridge that its accuracy rate is 91.36%. Full article
Show Figures

Figure 1

14 pages, 5743 KB  
Article
Classifying Heart-Sound Signals Based on CNN Trained on MelSpectrum and Log-MelSpectrum Features
by Wei Chen, Zixuan Zhou, Junze Bao, Chengniu Wang, Hanqing Chen, Chen Xu, Gangcai Xie, Hongmin Shen and Huiqun Wu
Bioengineering 2023, 10(6), 645; https://doi.org/10.3390/bioengineering10060645 - 25 May 2023
Cited by 29 | Viewed by 3712
Abstract
The intelligent classification of heart-sound signals can assist clinicians in the rapid diagnosis of cardiovascular diseases. Mel-frequency cepstral coefficients (MelSpectrums) and log Mel-frequency cepstral coefficients (Log-MelSpectrums) based on a short-time Fourier transform (STFT) can represent the temporal and spectral structures of original heart-sound [...] Read more.
The intelligent classification of heart-sound signals can assist clinicians in the rapid diagnosis of cardiovascular diseases. Mel-frequency cepstral coefficients (MelSpectrums) and log Mel-frequency cepstral coefficients (Log-MelSpectrums) based on a short-time Fourier transform (STFT) can represent the temporal and spectral structures of original heart-sound signals. Recently, various systems based on convolutional neural networks (CNNs) trained on the MelSpectrum and Log-MelSpectrum of segmental heart-sound frames that outperform systems using handcrafted features have been presented and classified heart-sound signals accurately. However, there is no a priori evidence of the best input representation for classifying heart sounds when using CNN models. Therefore, in this study, the MelSpectrum and Log-MelSpectrum features of heart-sound signals combined with a mathematical model of cardiac-sound acquisition were analysed theoretically. Both the experimental results and theoretical analysis demonstrated that the Log-MelSpectrum features can reduce the classification difference between domains and improve the performance of CNNs for heart-sound classification. Full article
Show Figures

Figure 1

17 pages, 8594 KB  
Article
Array-Based Underwater Acoustic Target Classification with Spectrum Reconstruction Based on Joint Sparsity and Frequency Shift Invariant Feature
by Chenxiang Lu, Xiangyang Zeng, Qiang Wang, Lu Wang and Anqi Jin
J. Mar. Sci. Eng. 2023, 11(6), 1101; https://doi.org/10.3390/jmse11061101 - 23 May 2023
Cited by 1 | Viewed by 2457
Abstract
The target spectrum, which is commonly used in feature extraction for underwater acoustic target classification, can be improperly recovered via conventional beamformer (CBF) owing to its frequency-variant spatial response and lead to degraded classification performance. In this paper, we propose a target spectrum [...] Read more.
The target spectrum, which is commonly used in feature extraction for underwater acoustic target classification, can be improperly recovered via conventional beamformer (CBF) owing to its frequency-variant spatial response and lead to degraded classification performance. In this paper, we propose a target spectrum reconstruction method under a sparse Bayesian learning framework with joint sparsity priors that can not only achieve high-resolution target separation in the angular domain but also attain beamwidth constancy over a frequency range at no cost of reducing angular resolution. Experiments on real measured array data show the recovered spectrum via our proposed method can effectively suppress interference and preserve more detailed spectral structures than CBF. This indicates our method is more suitable for target classification because it has the capability of retaining more representative and discriminative characteristics. Moreover, due to target motion and the underwater channel effect, the frequency of prominent spectral line components can be shifted over time, which is harmful to classification performance. To overcome this problem, we proposed a frequency shift-invariant feature extraction method with the help of elaborately designed frequency shift-invariant filter banks. The classification experiments demonstrate that our proposed methods outperform traditional CBF and Mel-frequency features and can help improve underwater recognition performance. Full article
(This article belongs to the Section Ocean Engineering)
Show Figures

Figure 1

10 pages, 4169 KB  
Article
Rotating Machinery State Recognition Based on Mel-Spectrum and Transfer Learning
by Fan Li, Zixiao Lu, Junyue Tang, Weiwei Zhang, Yahui Tian, Zhongyu Cui, Fei Jiang, Honglang Li and Shengyuan Jiang
Aerospace 2023, 10(5), 480; https://doi.org/10.3390/aerospace10050480 - 18 May 2023
Cited by 3 | Viewed by 2342
Abstract
During drilling into the soil, the rotating mechanical structure will be affected by soil particles and external disturbances, affecting the health of the rotating mechanical structure. Therefore, real-time monitoring of the operational status of rotating mechanical structures is of great significance. This paper [...] Read more.
During drilling into the soil, the rotating mechanical structure will be affected by soil particles and external disturbances, affecting the health of the rotating mechanical structure. Therefore, real-time monitoring of the operational status of rotating mechanical structures is of great significance. This paper proposes a working state recognition method based on Mel-spectrum and transfer learning, which uses the mechanical vibration signal’s time domain and frequency domain information to identify the mechanical structure’s working state. Firstly, we cut the signal at window length, and then the Mel-spectrum of the truncated signal is obtained through the Fourier transform and Mel-scale filter bank. Finally, we adopted the method of transfer learning. The pre-trained model VGG16 is adjusted to extract and classify the features of the Mel-spectrum. Experimental results show that the framework maintains an accuracy of more than 90% for vibration signals under minor window conditions, which verifies the real-time reliability of the method. Full article
(This article belongs to the Special Issue Space Sampling and Exploration Robotics)
Show Figures

Figure 1

17 pages, 8438 KB  
Article
A Semi-Supervised Speech Deception Detection Algorithm Combining Acoustic Statistical Features and Time-Frequency Two-Dimensional Features
by Hongliang Fu, Hang Yu, Xuemei Wang, Xiangying Lu and Chunhua Zhu
Brain Sci. 2023, 13(5), 725; https://doi.org/10.3390/brainsci13050725 - 26 Apr 2023
Cited by 6 | Viewed by 3014
Abstract
Human lying is influenced by cognitive neural mechanisms in the brain, and conducting research on lie detection in speech can help to reveal the cognitive mechanisms of the human brain. Inappropriate deception detection features can easily lead to dimension disaster and make the [...] Read more.
Human lying is influenced by cognitive neural mechanisms in the brain, and conducting research on lie detection in speech can help to reveal the cognitive mechanisms of the human brain. Inappropriate deception detection features can easily lead to dimension disaster and make the generalization ability of the widely used semi-supervised speech deception detection model worse. Because of this, this paper proposes a semi-supervised speech deception detection algorithm combining acoustic statistical features and time-frequency two-dimensional features. Firstly, a hybrid semi-supervised neural network based on a semi-supervised autoencoder network (AE) and a mean-teacher network is established. Secondly, the static artificial statistical features are input into the semi-supervised AE to extract more robust advanced features, and the three-dimensional (3D) mel-spectrum features are input into the mean-teacher network to obtain features rich in time-frequency two-dimensional information. Finally, a consistency regularization method is introduced after feature fusion, effectively reducing the occurrence of over-fitting and improving the generalization ability of the model. This paper carries out experiments on the self-built corpus for deception detection. The experimental results show that the highest recognition accuracy of the algorithm proposed in this paper is 68.62% which is 1.2% higher than the baseline system and effectively improves the detection accuracy. Full article
(This article belongs to the Special Issue Neural Network for Speech and Gesture Semantics)
Show Figures

Figure 1

16 pages, 6773 KB  
Article
Defect Identification Method for Transformer End Pad Falling Based on Acoustic Stability Feature Analysis
by Shuai Han, Bowen Wang, Sizhuo Liao, Fei Gao and Mo Chen
Sensors 2023, 23(6), 3258; https://doi.org/10.3390/s23063258 - 20 Mar 2023
Cited by 9 | Viewed by 2408
Abstract
A transformer’s acoustic signal contains rich information. The acoustic signal can be divided into a transient acoustic signal and a steady-state acoustic signal under different operating conditions. In this paper, the vibration mechanism is analyzed, and the acoustic feature is mined based on [...] Read more.
A transformer’s acoustic signal contains rich information. The acoustic signal can be divided into a transient acoustic signal and a steady-state acoustic signal under different operating conditions. In this paper, the vibration mechanism is analyzed, and the acoustic feature is mined based on the transformer end pad falling defect to realize defect identification. Firstly, a quality–spring–damping model is established to analyze the vibration modes and development patterns of the defect. Secondly, short-time Fourier transform is applied to the voiceprint signals, and the time–frequency spectrum is compressed and perceived using Mel filter banks. Thirdly, the time-series spectrum entropy feature extraction algorithm is introduced into the stability calculation, and the algorithm is verified by comparing it with simulated experimental samples. Finally, stability calculations are performed on the voiceprint signal data collected from 162 transformers operating in the field, and the stability distribution is statistically analyzed. The time-series spectrum entropy stability warning threshold is given, and the application value of the threshold is demonstrated by comparing it with actual fault cases. Full article
Show Figures

Figure 1

12 pages, 10333 KB  
Article
A Method of Speech Coding for Speech Recognition Using a Convolutional Neural Network
by Mariusz Kubanek, Janusz Bobulski and Joanna Kulawik
Symmetry 2019, 11(9), 1185; https://doi.org/10.3390/sym11091185 - 19 Sep 2019
Cited by 37 | Viewed by 6159
Abstract
This work presents a new approach to speech recognition, based on the specific coding of time and frequency characteristics of speech. The research proposed the use of convolutional neural networks because, as we know, they show high resistance to cross-spectral distortions and differences [...] Read more.
This work presents a new approach to speech recognition, based on the specific coding of time and frequency characteristics of speech. The research proposed the use of convolutional neural networks because, as we know, they show high resistance to cross-spectral distortions and differences in the length of the vocal tract. Until now, two layers of time convolution and frequency convolution were used. A novel idea is to weave three separate convolution layers: traditional time convolution and the introduction of two different frequency convolutions (mel-frequency cepstral coefficients (MFCC) convolution and spectrum convolution). This application takes into account more details contained in the tested signal. Our idea assumes creating patterns for sounds in the form of RGB (Red, Green, Blue) images. The work carried out research for isolated words and continuous speech, for neural network structure. A method for dividing continuous speech into syllables has been proposed. This method can be used for symmetrical stereo sound. Full article
Show Figures

Graphical abstract

17 pages, 3284 KB  
Article
Resource-Efficient Pet Dog Sound Events Classification Using LSTM-FCN Based on Time-Series Data
by Yunbin Kim, Jaewon Sa, Yongwha Chung, Daihee Park and Sungju Lee
Sensors 2018, 18(11), 4019; https://doi.org/10.3390/s18114019 - 18 Nov 2018
Cited by 24 | Viewed by 9198
Abstract
The use of IoT (Internet of Things) technology for the management of pet dogs left alone at home is increasing. This includes tasks such as automatic feeding, operation of play equipment, and location detection. Classification of the vocalizations of pet dogs using information [...] Read more.
The use of IoT (Internet of Things) technology for the management of pet dogs left alone at home is increasing. This includes tasks such as automatic feeding, operation of play equipment, and location detection. Classification of the vocalizations of pet dogs using information from a sound sensor is an important method to analyze the behavior or emotions of dogs that are left alone. These sounds should be acquired by attaching the IoT sound sensor to the dog, and then classifying the sound events (e.g., barking, growling, howling, and whining). However, sound sensors tend to transmit large amounts of data and consume considerable amounts of power, which presents issues in the case of resource-constrained IoT sensor devices. In this paper, we propose a way to classify pet dog sound events and improve resource efficiency without significant degradation of accuracy. To achieve this, we only acquire the intensity data of sounds by using a relatively resource-efficient noise sensor. This presents issues as well, since it is difficult to achieve sufficient classification accuracy using only intensity data due to the loss of information from the sound events. To address this problem and avoid significant degradation of classification accuracy, we apply long short-term memory-fully convolutional network (LSTM-FCN), which is a deep learning method, to analyze time-series data, and exploit bicubic interpolation. Based on experimental results, the proposed method based on noise sensors (i.e., Shapelet and LSTM-FCN for time-series) was found to improve energy efficiency by 10 times without significant degradation of accuracy compared to typical methods based on sound sensors (i.e., mel-frequency cepstrum coefficient (MFCC), spectrogram, and mel-spectrum for feature extraction, and support vector machine (SVM) and k-nearest neighbor (K-NN) for classification). Full article
(This article belongs to the Section Internet of Things)
Show Figures

Figure 1

Back to TopTop