Next Article in Journal
Dispersion Turning Attenuation Microfiber for Flowrate Sensing
Next Article in Special Issue
AuCFSR: Authentication and Color Face Self-Recovery Using Novel 2D Hyperchaotic System and Deep Learning Models
Previous Article in Journal
Wireless Sensor Network-Based Rockfall and Landslide Monitoring Systems: A Review
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Communication

A Pre-Training Framework Based on Multi-Order Acoustic Simulation for Replay Voice Spoofing Detection

1
Department of Computer Engineering, Chosun University, Gwangju 61452, Republic of Korea
2
Digital Analysis Division, National Forensic Service, Wonju 26460, Republic of Korea
*
Author to whom correspondence should be addressed.
Sensors 2023, 23(16), 7280; https://doi.org/10.3390/s23167280
Submission received: 23 July 2023 / Revised: 10 August 2023 / Accepted: 15 August 2023 / Published: 20 August 2023
(This article belongs to the Special Issue Sensors in Multimedia Forensics)

Abstract

Voice spoofing attempts to break into a specific automatic speaker verification (ASV) system by forging the user’s voice and can be used through methods such as text-to-speech (TTS), voice conversion (VC), and replay attacks. Recently, deep learning-based voice spoofing countermeasures have been developed. However, the problem with replay is that it is difficult to construct a large number of datasets because it requires a physical recording process. To overcome these problems, this study proposes a pre-training framework based on multi-order acoustic simulation for replay voice spoofing detection. Multi-order acoustic simulation utilizes existing clean signal and room impulse response (RIR) datasets to generate audios, which simulate the various acoustic configurations of the original and replayed audios. The acoustic configuration refers to factors such as the microphone type, reverberation, time delay, and noise that may occur between a speaker and microphone during the recording process. We assume that a deep learning model trained on an audio that simulates the various acoustic configurations of the original and replayed audios can classify the acoustic configurations of the original and replay audios well. To validate this, we performed pre-training to classify the audio generated by the multi-order acoustic simulation into three classes: clean signal, audio simulating the acoustic configuration of the original audio, and audio simulating the acoustic configuration of the replay audio. We also set the weights of the pre-training model to the initial weights of the replay voice spoofing detection model using the existing replay voice spoofing dataset and then performed fine-tuning. To validate the effectiveness of the proposed method, we evaluated the performance of the conventional method without pre-training and proposed method using an objective metric, i.e., the accuracy and F1-score. As a result, the conventional method achieved an accuracy of 92.94%, F1-score of 86.92% and the proposed method achieved an accuracy of 98.16%, F1-score of 95.08%.
Keywords: voice spoofing; acoustic configuration; deep learning voice spoofing; acoustic configuration; deep learning

Share and Cite

MDPI and ACS Style

Go, C.; Park, N.I.; Jeon, O.-Y.; Chun, C. A Pre-Training Framework Based on Multi-Order Acoustic Simulation for Replay Voice Spoofing Detection. Sensors 2023, 23, 7280. https://doi.org/10.3390/s23167280

AMA Style

Go C, Park NI, Jeon O-Y, Chun C. A Pre-Training Framework Based on Multi-Order Acoustic Simulation for Replay Voice Spoofing Detection. Sensors. 2023; 23(16):7280. https://doi.org/10.3390/s23167280

Chicago/Turabian Style

Go, Changhwan, Nam In Park, Oc-Yeub Jeon, and Chanjun Chun. 2023. "A Pre-Training Framework Based on Multi-Order Acoustic Simulation for Replay Voice Spoofing Detection" Sensors 23, no. 16: 7280. https://doi.org/10.3390/s23167280

APA Style

Go, C., Park, N. I., Jeon, O.-Y., & Chun, C. (2023). A Pre-Training Framework Based on Multi-Order Acoustic Simulation for Replay Voice Spoofing Detection. Sensors, 23(16), 7280. https://doi.org/10.3390/s23167280

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop