Next Article in Journal
Study on Multi-Crack Damage Evolution and Fatigue Life of Corroded Steel Wires Inside In-Service Bridge Suspenders
Previous Article in Journal
A Novel Robust Hybrid Control Strategy for a Quadrotor Trajectory Tracking Aided with Bioinspired Neural Dynamics
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

CycleDiffusion: Voice Conversion Using Cycle-Consistent Diffusion Models

Artificial Intelligence Laboratory, Department of Computer Science and Engineering, Korea University, Seoul 02841, Republic of Korea
*
Author to whom correspondence should be addressed.
Appl. Sci. 2024, 14(20), 9595; https://doi.org/10.3390/app14209595
Submission received: 3 October 2024 / Revised: 15 October 2024 / Accepted: 18 October 2024 / Published: 21 October 2024
(This article belongs to the Section Computing and Artificial Intelligence)

Abstract

Voice conversion (VC) refers to the technique of modifying one speaker’s voice to mimic another’s while retaining the original linguistic content. This technology finds its applications in fields such as speech synthesis, accent modification, medicine, security, privacy, and entertainment. Among the various deep generative models used for voice conversion, including variational autoencoders (VAEs) and generative adversarial networks (GANs), diffusion models (DMs) have recently gained attention as promising methods due to their training stability and strong performance in data generation. Nevertheless, traditional DMs focus mainly on learning reconstruction paths like VAEs, rather than conversion paths as GANs do, thereby restricting the quality of the converted speech. To overcome this limitation and enhance voice conversion performance, we propose a cycle-consistent diffusion (CycleDiffusion) model, which comprises two DMs: one for converting the source speaker’s voice to the target speaker’s voice and the other for converting it back to the source speaker’s voice. By employing two DMs and enforcing a cycle consistency loss, the CycleDiffusion model effectively learns both reconstruction and conversion paths, producing high-quality converted speech. The effectiveness of the proposed model in voice conversion is validated through experiments using the VCTK (Voice Cloning Toolkit) dataset.
Keywords: cycle consistency; diffusion model; voice conversion cycle consistency; diffusion model; voice conversion

Share and Cite

MDPI and ACS Style

Yook, D.; Han, G.; Chang, H.-P.; Yoo, I.-C. CycleDiffusion: Voice Conversion Using Cycle-Consistent Diffusion Models. Appl. Sci. 2024, 14, 9595. https://doi.org/10.3390/app14209595

AMA Style

Yook D, Han G, Chang H-P, Yoo I-C. CycleDiffusion: Voice Conversion Using Cycle-Consistent Diffusion Models. Applied Sciences. 2024; 14(20):9595. https://doi.org/10.3390/app14209595

Chicago/Turabian Style

Yook, Dongsuk, Geonhee Han, Hyung-Pil Chang, and In-Chul Yoo. 2024. "CycleDiffusion: Voice Conversion Using Cycle-Consistent Diffusion Models" Applied Sciences 14, no. 20: 9595. https://doi.org/10.3390/app14209595

APA Style

Yook, D., Han, G., Chang, H.-P., & Yoo, I.-C. (2024). CycleDiffusion: Voice Conversion Using Cycle-Consistent Diffusion Models. Applied Sciences, 14(20), 9595. https://doi.org/10.3390/app14209595

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop