This is an early access version, the complete PDF, HTML, and XML versions will be available soon.
Open AccessArticle
VQ-CycleDiffusion: Vector Quantized Cycle-Consistent Diffusion Models for Voice Conversion
by
Dongsuk Yook
Dongsuk Yook 1,*,
Semin Kim
Semin Kim 2 and
Hyung-Pil Chang
Hyung-Pil Chang 1
1
Artificial Intelligence Laboratory, Department of Computer Science and Engineering, Korea University, Seoul 02841, Republic of Korea
2
Artificial Intelligence Laboratory, Department of Data Science, Korea University, Seoul 02841, Republic of Korea
*
Author to whom correspondence should be addressed.
Appl. Sci. 2026, 16(18), 9314; https://doi.org/10.3390/app16189314 (registering DOI)
Submission received: 5 August 2026
/
Revised: 13 September 2026
/
Accepted: 14 September 2026
/
Published: 19 September 2026
Abstract
This paper proposes a novel hybrid voice conversion framework, termed VQ-CycleDiffusion, which integrates vector quantization (VQ) with CycleDiffusion to address the inherent limitations of continuous latent representations in diffusion-based models. While conventional diffusion-based voice conversion (VC) models provide superior speech quality, they often suffer from speaker information leakage due to their reliance on continuous latent spaces, which fail to completely separate linguistic content from source speaker characteristics. To overcome this issue, we introduce a discrete bottleneck via VQ to extract multi-speaker linguistic representations, which are then used as conditional inputs for CycleDiffusion after a carefully designed codeword transformation between the source and target speakers. Experimental results on the VCTK corpus demonstrate that the proposed model significantly outperforms the conventional method in terms of speaker similarity, as confirmed by both i-vector and x-vector cosine similarity metrics, while maintaining spectral reconstruction performance comparable to that of the baseline model, as evidenced by stable Mel-cepstral distance (MCD) values. Furthermore, the quality of the converted speech was not degraded, as demonstrated by the predicted mean opinion score (MOS) evaluation. Ablation studies reveal that initializing the codebook using k-means clustering and focusing on diffusion model fine-tuning play key roles in maximizing performance. These results validate that the proposed method effectively disentangles linguistic content from speaker characteristics while preserving the high-fidelity generation capability of diffusion models.
Share and Cite
MDPI and ACS Style
Yook, D.; Kim, S.; Chang, H.-P.
VQ-CycleDiffusion: Vector Quantized Cycle-Consistent Diffusion Models for Voice Conversion. Appl. Sci. 2026, 16, 9314.
https://doi.org/10.3390/app16189314
AMA Style
Yook D, Kim S, Chang H-P.
VQ-CycleDiffusion: Vector Quantized Cycle-Consistent Diffusion Models for Voice Conversion. Applied Sciences. 2026; 16(18):9314.
https://doi.org/10.3390/app16189314
Chicago/Turabian Style
Yook, Dongsuk, Semin Kim, and Hyung-Pil Chang.
2026. "VQ-CycleDiffusion: Vector Quantized Cycle-Consistent Diffusion Models for Voice Conversion" Applied Sciences 16, no. 18: 9314.
https://doi.org/10.3390/app16189314
APA Style
Yook, D., Kim, S., & Chang, H.-P.
(2026). VQ-CycleDiffusion: Vector Quantized Cycle-Consistent Diffusion Models for Voice Conversion. Applied Sciences, 16(18), 9314.
https://doi.org/10.3390/app16189314
Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details
here.
Article Metrics
Article metric data becomes available approximately 24 hours after publication online.