Previous Article in Journal
Advances in Nonlinear Dynamics and Vibration Control of Clearance-Containing Flexible Oscillating Mechanisms
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
This is an early access version, the complete PDF, HTML, and XML versions will be available soon.
Article

VQ-CycleDiffusion: Vector Quantized Cycle-Consistent Diffusion Models for Voice Conversion

1
Artificial Intelligence Laboratory, Department of Computer Science and Engineering, Korea University, Seoul 02841, Republic of Korea
2
Artificial Intelligence Laboratory, Department of Data Science, Korea University, Seoul 02841, Republic of Korea
*
Author to whom correspondence should be addressed.
Appl. Sci. 2026, 16(18), 9314; https://doi.org/10.3390/app16189314 (registering DOI)
Submission received: 5 August 2026 / Revised: 13 September 2026 / Accepted: 14 September 2026 / Published: 19 September 2026
(This article belongs to the Section Computing and Artificial Intelligence)

Abstract

This paper proposes a novel hybrid voice conversion framework, termed VQ-CycleDiffusion, which integrates vector quantization (VQ) with CycleDiffusion to address the inherent limitations of continuous latent representations in diffusion-based models. While conventional diffusion-based voice conversion (VC) models provide superior speech quality, they often suffer from speaker information leakage due to their reliance on continuous latent spaces, which fail to completely separate linguistic content from source speaker characteristics. To overcome this issue, we introduce a discrete bottleneck via VQ to extract multi-speaker linguistic representations, which are then used as conditional inputs for CycleDiffusion after a carefully designed codeword transformation between the source and target speakers. Experimental results on the VCTK corpus demonstrate that the proposed model significantly outperforms the conventional method in terms of speaker similarity, as confirmed by both i-vector and x-vector cosine similarity metrics, while maintaining spectral reconstruction performance comparable to that of the baseline model, as evidenced by stable Mel-cepstral distance (MCD) values. Furthermore, the quality of the converted speech was not degraded, as demonstrated by the predicted mean opinion score (MOS) evaluation. Ablation studies reveal that initializing the codebook using k-means clustering and focusing on diffusion model fine-tuning play key roles in maximizing performance. These results validate that the proposed method effectively disentangles linguistic content from speaker characteristics while preserving the high-fidelity generation capability of diffusion models.
Keywords: vector quantization; cycle consistency; diffusion model; voice conversion vector quantization; cycle consistency; diffusion model; voice conversion

Share and Cite

MDPI and ACS Style

Yook, D.; Kim, S.; Chang, H.-P. VQ-CycleDiffusion: Vector Quantized Cycle-Consistent Diffusion Models for Voice Conversion. Appl. Sci. 2026, 16, 9314. https://doi.org/10.3390/app16189314

AMA Style

Yook D, Kim S, Chang H-P. VQ-CycleDiffusion: Vector Quantized Cycle-Consistent Diffusion Models for Voice Conversion. Applied Sciences. 2026; 16(18):9314. https://doi.org/10.3390/app16189314

Chicago/Turabian Style

Yook, Dongsuk, Semin Kim, and Hyung-Pil Chang. 2026. "VQ-CycleDiffusion: Vector Quantized Cycle-Consistent Diffusion Models for Voice Conversion" Applied Sciences 16, no. 18: 9314. https://doi.org/10.3390/app16189314

APA Style

Yook, D., Kim, S., & Chang, H.-P. (2026). VQ-CycleDiffusion: Vector Quantized Cycle-Consistent Diffusion Models for Voice Conversion. Applied Sciences, 16(18), 9314. https://doi.org/10.3390/app16189314

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Article metric data becomes available approximately 24 hours after publication online.
Back to TopTop