Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (4)

Search Parameters:
Keywords = distant speech processing

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
19 pages, 9045 KB  
Article
Audio Pre-Processing and Beamforming Implementation on Embedded Systems
by Jian-Hong Wang, Phuong Thi Le, Shih-Jung Kuo, Tzu-Chiang Tai, Kuo-Chen Li, Shih-Lun Chen, Ze-Yu Wang, Tuan Pham, Yung-Hui Li and Jia-Ching Wang
Electronics 2024, 13(14), 2784; https://doi.org/10.3390/electronics13142784 - 15 Jul 2024
Cited by 5 | Viewed by 3983
Abstract
Since the invention of the microphone by Barina in 1876, there have been numerous applications of audio processing, such as phonographs, broadcasting stations, and public address systems, which merely capture and amplify sound and play it back. Nowadays, audio processing involves analysis and [...] Read more.
Since the invention of the microphone by Barina in 1876, there have been numerous applications of audio processing, such as phonographs, broadcasting stations, and public address systems, which merely capture and amplify sound and play it back. Nowadays, audio processing involves analysis and noise-filtering techniques. There are various methods for noise filtering, each employing unique algorithms, but they all require two or more microphones for signal processing and analysis. For instance, on mobile phones, two microphones located in different positions are utilized for active noise cancellation (one for primary audio capture and the other for capturing ambient noise). However, a drawback is that when the sound source is distant, it may lead to poor audio capture. To capture sound from distant sources, alternative methods, like blind signal separation and beamforming, are necessary. This paper proposes employing a beamforming algorithm with two microphones to enhance speech and implementing this algorithm on an embedded system. However, prior to beamforming, it is imperative to accurately detect the direction of the sound source to process and analyze the audio from that direction. Full article
(This article belongs to the Special Issue Recent Advances in Audio, Speech and Music Processing and Analysis)
Show Figures

Figure 1

15 pages, 11962 KB  
Article
Speech Enhancement Based on Two-Stage Processing with Deep Neural Network for Laser Doppler Vibrometer
by Chengkai Cai, Kenta Iwai and Takanobu Nishiura
Appl. Sci. 2023, 13(3), 1958; https://doi.org/10.3390/app13031958 - 2 Feb 2023
Cited by 5 | Viewed by 3430
Abstract
The development of distant-talk measurement systems has been attracting attention since they can be applied to many situations such as security and disaster relief. One such system that uses a device called a laser Doppler vibrometer (LDV) to acquire sound by measuring an [...] Read more.
The development of distant-talk measurement systems has been attracting attention since they can be applied to many situations such as security and disaster relief. One such system that uses a device called a laser Doppler vibrometer (LDV) to acquire sound by measuring an object’s vibration caused by the sound source has been proposed. Different from traditional microphones, an LDV can pick up the target sound from a distance even in a noisy environment. However, the acquired sounds are greatly distorted due to the object’s shape and frequency response. Due to the particularity of the degradation of observed speech, conventional methods cannot be effectively applied to LDVs. We propose two speech enhancement methods that are based on two-stage processing with deep neural networks for LDVs. With the first proposed method, the amplitude spectrum of the observed speech is first restored. The phase difference between the observed and clean speech is then estimated using the restored amplitude spectrum. With the other proposed method, the low-frequency components of the observed speech are first restored. The high-frequency components are then estimated by the restored low-frequency components. The evaluation results indicate that they improved the observed speech in sound quality, deterioration degree, and intelligibility. Full article
(This article belongs to the Special Issue Audio and Acoustic Signal Processing)
Show Figures

Figure 1

21 pages, 860 KB  
Article
Application of Fusion of Various Spontaneous Speech Analytics Methods for Improving Far-Field Neural-Based Diarization
by Sergei Astapov, Aleksei Gusev, Marina Volkova, Aleksei Logunov, Valeriia Zaluskaia, Vlada Kapranova, Elena Timofeeva, Elena Evseeva, Vladimir Kabarov and Yuri Matveev
Mathematics 2021, 9(23), 2998; https://doi.org/10.3390/math9232998 - 23 Nov 2021
Cited by 1 | Viewed by 3173
Abstract
Recently developed methods in spontaneous speech analytics require the use of speaker separation based on audio data, referred to as diarization. It is applied to widespread use cases, such as meeting transcription based on recordings from distant microphones and the extraction of the [...] Read more.
Recently developed methods in spontaneous speech analytics require the use of speaker separation based on audio data, referred to as diarization. It is applied to widespread use cases, such as meeting transcription based on recordings from distant microphones and the extraction of the target speaker’s voice profiles from noisy audio. However, speech recognition and analysis can be hindered by background and point-source noise, overlapping speech, and reverberation, which all affect diarization quality in conjunction with each other. To compensate for the impact of these factors, there are a variety of supportive speech analytics methods, such as quality assessments in terms of SNR and RT60 reverberation time metrics, overlapping speech detection, instant speaker number estimation, etc. The improvements in speaker verification methods have benefits in the area of speaker separation as well. This paper introduces several approaches aimed towards improving diarization system quality. The presented experimental results demonstrate the possibility of refining initial speaker labels from neural-based VAD data by means of fusion with labels from quality estimation models, overlapping speech detectors, and speaker number estimation models, which contain CNN and LSTM modules. Such fusing approaches allow us to significantly decrease DER values compared to standalone VAD methods. Cases of ideal VAD labeling are utilized to show the positive impact of ResNet-101 neural networks on diarization quality in comparison with basic x-vectors and ECAPA-TDNN architectures trained on 8 kHz data. Moreover, this paper highlights the advantage of spectral clustering over other clustering methods applied to diarization. The overall quality of diarization is improved at all stages of the pipeline, and the combination of various speech analytics methods makes a significant contribution to the improvement of diarization quality. Full article
Show Figures

Figure 1

21 pages, 3799 KB  
Article
Electrical Brain Responses Reveal Sequential Constraints on Planning during Music Performance
by Brian Mathias, William J. Gehring and Caroline Palmer
Brain Sci. 2019, 9(2), 25; https://doi.org/10.3390/brainsci9020025 - 28 Jan 2019
Cited by 12 | Viewed by 5664
Abstract
Elements in speech and music unfold sequentially over time. To produce sentences and melodies quickly and accurately, individuals must plan upcoming sequence events, as well as monitor outcomes via auditory feedback. We investigated the neural correlates of sequential planning and monitoring processes by [...] Read more.
Elements in speech and music unfold sequentially over time. To produce sentences and melodies quickly and accurately, individuals must plan upcoming sequence events, as well as monitor outcomes via auditory feedback. We investigated the neural correlates of sequential planning and monitoring processes by manipulating auditory feedback during music performance. Pianists performed isochronous melodies from memory at an initially cued rate while their electroencephalogram was recorded. Pitch feedback was occasionally altered to match either an immediately upcoming Near-Future pitch (next sequence event) or a more distant Far-Future pitch (two events ahead of the current event). Near-Future, but not Far-Future altered feedback perturbed the timing of pianists’ performances, suggesting greater interference of Near-Future sequential events with current planning processes. Near-Future feedback triggered a greater reduction in auditory sensory suppression (enhanced response) than Far-Future feedback, reflected in the P2 component elicited by the pitch event following the unexpected pitch change. Greater timing perturbations were associated with enhanced cortical sensory processing of the pitch event following the Near-Future altered feedback. Both types of feedback alterations elicited feedback-related negativity (FRN) and P3a potentials and amplified spectral power in the theta frequency range. These findings suggest similar constraints on producers’ sequential planning to those reported in speech production. Full article
(This article belongs to the Special Issue Advances in the Neurocognition of Music and Language)
Show Figures

Figure 1

Back to TopTop