1. Introduction
Recent neural text-to-speech (TTS) systems have achieved substantial improvements in naturalness, intelligibility, and speaker fidelity. Reliable emotional rendering nevertheless remains difficult for applications such as virtual avatars, cinematic dubbing, and human–computer interaction. Emotional speech is conveyed through coupled variations in pitch, energy, duration, rhythm, pause structure, and speaking rate, rather than through spectral coloration alone [
1,
2,
3]. A practical emotional adaptation method must therefore increase affective expressiveness while preserving speaker timbre, pronunciation stability, alignment robustness, and the acoustic priors learned from large-scale neutral speech.
In flow-matching TTS systems with monotonic text-to-mel alignment, emotional variation affects both the vector field that generates mel-spectrogram trajectories and the phoneme-level duration pattern that shapes the rhythm, speaking rate, and pause structure. This coupling motivates a dual-track formulation in which emotional adaptation is introduced through two residual paths: one for acoustic vector field prediction and one for duration prediction. Representing emotion relative to neutral speech further allows the control signal to describe a shift from a neutral anchor toward the target emotional state, rather than an absolute style vector.
We propose the dual-track residual framework (DTRF) for residual strength-controlled emotional speech synthesis with a frozen flow-matching TTS backbone. The DTRF is built on a neutral-adapted, MAS-aligned Matcha-TTS-style OT-CFM backbone and keeps the text encoder, duration predictor, speaker embedding, and flow-matching decoder fixed during emotional adaptation. An asymmetric Emo-ControlNet injects zero-initialized acoustic residuals, while an EDA predicts the log-duration residuals. During inference, a shared scalar factor
scales both residual paths, providing a practical interface for controllable emotion rendering. Flow-matching TTS, controllable emotional synthesis, prosody modeling, and frozen backbone adaptation are reviewed in
Section 2.
The main contributions are as follows:
We propose the DTRF, a frozen backbone residual adaptation framework that injects emotion into both acoustic vector field prediction and phoneme duration prediction for flow-matching TTS.
We introduce a neutral-anchor emotion residual representation that converts absolute emotion embeddings into relative deviations from neutral speech, formulating emotional control as a shift from neutral speech toward the target emotion.
We design a shared residual strength control mechanism in which a scalar factor jointly scales the acoustic residual and the duration residual during inference, enabling controllable emotion rendering without updating the acoustic backbone.
On the English subset of the Emotional Speech Dataset (ESD), the DTRF is evaluated against internal and external baselines using objective metrics, subjective listening tests, and representation-level analyses. The remainder of the paper is organized as follows:
Section 2 reviews the related work;
Section 3 presents the DTRF;
Section 4 reports the experimental results; and
Section 5 and
Section 6 discuss the findings and conclude the study.
2. Related Work
2.1. Flow-Matching TTS
Continuous flow matching (CFM) has emerged as an effective paradigm for generative modeling by learning continuous vector fields that transform simple prior distributions into data distributions [
4,
5]. In speech generation, flow-matching formulations have been shown to support high-quality and flexible synthesis. Voicebox demonstrates large-scale text-guided speech generation, while VoiceFlow applies rectified flow matching to efficient text-to-speech generation [
6,
7]. Matcha-TTS further adapts Optimal Transport Conditional Flow Matching (OT-CFM) to non-autoregressive TTS by using monotonic alignment search (MAS) to obtain text-to-mel alignment and phoneme-level duration supervision [
8]. E2 TTS and F5-TTS also demonstrate the scalability of flow matching-style speech generation under prompt-based settings [
9,
10]. These studies demonstrate that flow matching-based TTS can achieve strong acoustic quality and flexible generation; however, most are designed for general speech synthesis and do not address controllable emotional adaptation of a frozen acoustic backbone. EmoCtrl-TTS integrates continuous affective trajectories into flow matching-based TTS [
11], showing the potential of flow-matching formulations for expressive control beyond neutral speech generation; nevertheless, frozen backbone emotional adaptation within this paradigm remains less systematically studied.
2.2. Controllable Emotional TTS
Controllable emotional TTS aims to generate speech with specified affective characteristics while preserving naturalness, intelligibility, and speaker identity. Existing emotion control strategies generally involve a trade-off between emotional expressiveness and timbre preservation. Discrete label or condition-based emotional TTS methods provide a simple and intuitive conditioning interface, but their control granularity is often limited because emotion is represented as a small set of categorical labels [
12,
13,
14,
15,
16]. Continuous affect representations, such as arousal–valence or valence–arousal–dominance (VAD), provide a finer description of emotional states and have been used to analyze or control expressive speech trajectories [
1,
11,
17,
18]. These continuous representations are particularly useful for modeling gradual changes in emotion intensity, but maintaining speaker consistency under continuous affective manipulation remains challenging.
Reference-based and emotion transfer methods can capture rich expressive patterns from reference speech, but they often require high-quality reference signals at inference time and may suffer from identity–emotion entanglement. In such cases, emotional style transfer may unintentionally leak source speaker characteristics or degrade the target speaker identity [
19,
20,
21,
22,
23,
24]. To alleviate this issue, prior studies have explored cross-speaker transfer, representation disentanglement, joint acoustic–emotion modeling, and prompt-based expressive control to improve emotional expressiveness while reducing speaker leakage [
25,
26]. Recent explorations of self-supervised disentanglement and self-refinement further indicate that emotion–speaker disentanglement remains an active research direction in cross-speaker expressive TTS [
27].
Self-supervised speech representations provide another route to emotional TTS; emotion2vec learns general-purpose speech emotion embeddings from large-scale self-supervised pretraining [
28], while DiEmo-TTS distills self-supervised emotional representations into a TTS system for cross-speaker emotional transfer [
29]. These studies show that high-quality emotion representations can support expressive synthesis and improve cross-speaker emotional transfer. However, how to inject such representations into a frozen pretrained acoustic model without disrupting its speaker and linguistic priors remains underexplored.
Recent work has also moved beyond utterance-level emotion labels toward local, time-varying, and multi-resolution control. By representing affect at the word, phoneme, frame, or hierarchical levels, these methods can support local emphasis, gradual emotion changes, and temporally varying expression [
1,
11,
25]. ControlNet-style extensions to flow-matching TTS, including TTS-CtrlNet, further investigate time-varying emotion control while keeping the acoustic backbone frozen [
30]. Nevertheless, many existing approaches rely on different backbone update strategies or training settings, and they do not specifically address emotional adaptation under a frozen neutral TTS backbone with explicit duration-level residual control.
2.3. Prosody and Duration Modeling
Prosody modeling is central to expressive and emotional speech synthesis. In non-autoregressive TTS systems, duration predictors and length regulators determine phoneme-level timing and control the expansion from phoneme-level linguistic representations to frame-level acoustic representations [
31,
32]. For emotional speech, the phoneme duration, rhythm, tempo, pause structure, and speaking rate carry affective cues that spectral or timbre-related modification alone does not fully capture. Emotional TTS therefore requires temporal prosody modeling in addition to acoustic emotion modeling; injecting emotion only into the acoustic generation pathway may leave the neutral timing pattern of the base model largely unchanged.
When the acoustic backbone and its associated duration predictor remain frozen, this limitation becomes more pronounced because the base duration model may be biased toward neutral or source-domain speaking patterns. Explicit duration-level adaptation is therefore needed alongside acoustic residual control; yet, to our knowledge, no prior frozen backbone emotional TTS method has jointly addressed phoneme-level duration residuals under a neutral adapted flow-matching backbone.
2.4. Frozen Backbone Residual Control
Adaptation methods for pretrained generative models aim to add new capabilities while preserving previously learned knowledge. Weight-space approaches, such as emotion arithmetic and task arithmetic, support interpolation or compositional adaptation by manipulating model weights or task vectors [
33,
34,
35]. However, such approaches may require storing, merging, or managing multiple model variants, and their behavior under limited emotional adaptation data can be difficult to control. Direct full-parameter updating can adapt a model to a target emotional domain, but it may also reduce the generalization ability of the pretrained model or weaken previously learned acoustic, linguistic, and speaker-related priors [
36,
37]. Parameter-efficient methods reduce adaptation costs by updating only a small subset of parameters or lightweight modules, offering a practical way to adapt pretrained models while limiting overfitting under limited data [
37,
38,
39].
Another line of work introduces controllability through auxiliary residual branches. ControlNet uses zero-initialized branches to add conditional control to a frozen image generation backbone while preserving the original generative prior [
40]. This design principle is relevant to frozen backbone adaptation because residual branches can introduce new control signals without directly updating the main network. TTS-CtrlNet extends this idea to flow-matching TTS and demonstrates that a frozen acoustic backbone can be augmented with trainable ControlNet-style branches for time-varying emotion control [
30]. However, existing residual control and parameter-efficient adaptation strategies have not fully explored dual-path emotional adaptation that jointly models acoustic residuals and phoneme-level duration residuals within a neutral-adapted Matcha-TTS-style OT-CFM backbone.
The preceding studies address important aspects of controllable emotional speech synthesis, including flow matching-based speech generation, emotion representation learning, cross-speaker emotional transfer, prosody modeling, and frozen backbone adaptation. Nevertheless, a gap remains for settings in which (1) the acoustic backbone must remain frozen after neutral speaker adaptation, (2) emotion must be injected into both acoustic generation and phoneme-level timing, and (3) speaker identity and intelligibility should be preserved under limited emotional training data. The DTRF is developed for this setting through dual-track residual adaptation with neutral-anchor emotion residuals, an asymmetric Emo-ControlNet, an EDA, and a shared inference-time scale factor
. Quantitative comparisons with representative internal and external systems under a unified evaluation protocol are reported in
Section 4.
3. Materials and Methods
3.1. Framework
3.1.1. Overall Framework
We present the DTRF for emotional adaptation of a pretrained Matcha-TTS [
8] backbone. As illustrated in
Figure 1, the Matcha-Base backbone is kept frozen during emotion control training to preserve the acoustic, linguistic, and speaker-related priors learned during pretraining and neutral adaptation. Emotional expressiveness is introduced through trainable residual modules instead of full-parameter updating.
The DTRF consists of three components: a neutral-anchor emotion residual representation, an asymmetric Emo-ControlNet for acoustic vector field control, and an emotional duration adapter (EDA) for duration-level prosody control. The neutral anchor converts absolute emotion embeddings into relative deviations from neutral speech. The Emo-ControlNet injects zero-initialized acoustic residual controls into the frozen flow prediction network, while the EDA predicts emotion-dependent log-duration residuals. Together, these two paths enable control over both spectral realization and temporal prosody under a frozen backbone setting.
3.1.2. Model Training
Training is formulated as a supervised text-to-speech task under the OT-CFM framework [
4]. For each training sample, we denote the ground-truth waveform as
x, the text transcript as
y, the speaker identity as
s, and the relative emotion residual as
, giving the input tuple
.
The waveform x is converted into a target mel-spectrogram , where is the number of mel-frequency bins and N is the acoustic frame length. The transcript y is converted into a phoneme sequence and passed through the frozen text encoder , which outputs the phoneme-level representation and the base log-duration prediction , where M is the phoneme length.
Following Matcha-TTS, we use monotonic alignment search (MAS) [
8,
41,
42,
43] to obtain a hard monotonic phoneme-to-frame alignment between
and
. This alignment expands
into the frame-level prior condition
and provides the MAS-derived duration target
for duration supervision.
Figure 2 summarizes the training pipeline, including MAS alignment, OT-CFM vector field learning, the Emo-ControlNet acoustic residual branch, and the EDA duration branch.
Within OT-CFM, Gaussian noise
and a random timestep
are sampled. The intermediate state is constructed using the
-parameterized interpolation used in our implementation:
The corresponding target vector field is
The frozen flow prediction network receives
and predicts the base vector field
. In parallel, the Emo-ControlNet receives the channel-wise concatenation
, where
denotes concatenation along the feature or channel dimension. The 768-dimensional residual emotion feature
is projected into the timestep-embedding space by
. In our implementation,
is a two-layer multilayer perceptron (MLP) with SiLU activation: Linear
–SiLU–Linear
. The emotion projection is fused with the timestep embedding via element-wise addition:
The Emo-ControlNet branch produces multi-scale control features at the residual injection points:
where
denotes the set of injection layers. Each zero-initialized residual head
projects
to the hidden dimension of the corresponding frozen decoder layer. The resulting training-time vector field is
The total training objective is a direct sum of the flow-matching loss, the prior loss, and the duration loss:
The flow-matching loss is computed as the mean-squared error between the predicted and target vector fields:
Following Matcha-TTS, the prior loss regularizes the aligned linguistic prior
toward the target mel-spectrogram
under a unit-variance Gaussian assumption:
where
denotes the valid-frame mask. The duration loss is
where
is obtained from MAS,
denotes the corresponding log-duration target, and
is defined in Equation (
13). During this stage, gradients are routed only to the Emo-ControlNet and the EDA; all Matcha-Base backbone parameters remain frozen. No additional intensity scaling is sampled during training, which is equivalent to training the residual paths at the standard scale
.
3.1.3. Model Inference
During inference, the DTRF takes a target transcript
, a target speaker identity
s, a relative emotion residual
, and a residual strength scale
as inputs. The transcript specifies linguistic content, the speaker identity anchors the target timbre,
specifies the target affective direction, and
scales the magnitude of both residual paths.
Figure 3 illustrates the inference pipeline.
The input transcript
is first encoded into the phoneme-level representation
and the base log-duration prediction
. The EDA predicts the emotional log-duration residual
. The inference-time log-duration prediction is obtained by applying
in the log-duration domain:
The frame-level duration is then discretized as follows:
where
denotes the ceiling operation. The prior representation
is expanded according to
to obtain the aligned acoustic condition
, where
is the generated mel-spectrogram length.
The acoustic trajectory is initialized from Gaussian noise
. Given the ordinary differential equation (ODE).
we use the explicit Euler solver with the number of function evaluations (NFE) fixed to 50 uniform integration steps from
to
for all formal evaluations. At each step, the final vector field is computed as follows:
where
is computed using the same control branch as in training, with
defined in Equation (
3). The state is updated by
. After integration, the generated mel-spectrogram is converted into waveform audio using a pretrained high-fidelity generative adversarial network (HiFi-GAN) vocoder.
3.2. Emotion Injection
3.2.1. Injection Mechanism
The emotion injection strategy is designed to reduce speaker–emotion entanglement and to distribute affective control across both spectral and temporal dimensions. We first extract an utterance-level emotion representation
using emotion2vec [
28]. Instead of using the absolute representation directly, we compute a global neutral anchor by averaging neutral utterance embeddings from the emotion control training split across speakers:
For a target emotional utterance, the relative emotion residual is defined as follows:
This residualization makes the control signal describe a deviation from neutral speech toward the target emotion, rather than an absolute style vector that may contain speaker-related information.
The acoustic residual path is implemented by an asymmetric Emo-ControlNet [
30]. The branch is adapted to the one-dimensional flow-matching decoder of Matcha-TTS and produces residual control features for acoustic vector field prediction. To reduce adaptation costs, the bypass branch retains only the two downsampling blocks and the middle block of the baseline U-Net [
44]. Following the ControlNet initialization principle [
40], these bypass blocks are initialized by deep-copying the corresponding frozen U-Net weights, and all residual output heads are initialized to zero.
In our implementation, residual controls are injected at three locations, , corresponding to the two downsampling blocks and the middle block of the frozen U-Net. At each injection point , the control feature matches the temporal resolution and channel dimension of the corresponding frozen U-Net feature, with 256 channels in our configuration. Each residual head is implemented as a zero-initialized 1D convolution with a kernel size of one, stride of one, and no padding. The residual output is added element-wise to the corresponding hidden feature in the frozen decoder. The zero initialization ensures that the control branch does not perturb the frozen backbone at the beginning of training.
The duration residual path is implemented by the EDA. For duration control, the EDA receives the phoneme-level prior representation
and the emotion residual
, where
M is the phoneme sequence length. The emotion residual is repeated along the phoneme-time axis to form
and concatenated with
along the channel dimension. The EDA input therefore has
channels. In our implementation, the EDA consists of Conv1d
, ReLU activation, LayerNorm, Conv1d
, ReLU activation, and a zero-initialized Conv1d
output layer. The predicted log-duration residual is
During training, the log-duration prediction used for duration supervision is obtained via additive fusion:
In the linear duration domain, this corresponds to multiplicative duration modulation before the final discretization step used during inference. This formulation allows the duration path to model emotion-related changes in speaking rate, segmental lengthening, and rhythm while keeping the frozen base duration predictor unchanged.
3.2.2. Inference-Time Residual Scaling
To provide controllable emotion rendering during inference, the DTRF introduces a scalar factor
that scales both residual paths. This factor is not restricted to acoustic conditioning; it jointly scales the acoustic vector field residual in Equation (
9) and the duration residual in Equation (
7). The acoustic residual injection points are detailed in
Figure 4, while the shared scaling of the acoustic and duration residuals is shown in
Figure 3. The same control parameter therefore adjusts both spectral realization and temporal prosody.
When , both residual paths are disabled, and the model approximately recovers the behavior of the frozen Matcha-Base backbone under the tested configuration. When , the model uses the residual magnitude seen during training. Values larger than increase the contribution of the learned emotion residuals and may make the target affect more salient, but they can also move the generated speech away from the stable operating region of the frozen backbone, leading to trade-offs in naturalness, speaker similarity, and intelligibility. In this work, is treated as an inference-time residual-strength scaling parameter; it moves the output from a neutral-like operating point toward the target emotion by jointly scaling acoustic and duration residuals, rather than acting as a fully disentangled perceptual intensity axis.
4. Results
4.1. Experimental Set-Up
4.1.1. Datasets and Data Splits
Experiments were conducted primarily on the English subset of ESD [
45], which contains 10 native English speakers, including five male and five female speakers, and five emotion categories: neutral, happy, angry, sad, and surprise. For emotion control training and evaluation, we used the filtered ESD emotion control file lists with 13,300 training utterances, 200 validation utterances, and 500 test utterances. Neutral utterances were used for neutral backbone adaptation and neutral-anchor construction, while emotion control evaluation focused on the four non-neutral categories: angry, sad, happy, and surprise.
The evaluation followed a seen speaker but text-disjoint setting. All 10 ESD speakers appeared in the training, validation, and test partitions, but the evaluation utterances and transcripts were disjoint from the training data. Utterances were grouped by normalized transcript, and samples with the same normalized linguistic content were assigned to the same subset whenever applicable to reduce text leakage. Since the ESD speaker identities were introduced during neutral backbone adaptation, this setting should be interpreted as seen speaker emotional adaptation rather than speaker-disjoint or fully zero-shot speaker generation. The split statistics are summarized in
Table 1.
For neutral backbone adaptation, we additionally used the Voice Cloning Toolkit (VCTK) corpus [
46], together with neutral speech from the ESD speakers. We followed the standard Matcha-TTS preprocessing and VCTK file list configuration without additional filtering rules. All audio used for training and evaluation was resampled to 22.05 kHz.
4.1.2. Backbone Adaptation and Emotion Training
We started from a Matcha-TTS model pretrained on VCTK. Before emotion control training, we adapted the model to the ESD speaker space while preserving the acoustic prior learned from VCTK. The speaker lookup table was expanded from 109 VCTK speakers to 119 identities by adding the 10 ESD speakers. During this stage, the CFM decoder and speaker embedding table were optimized, while the text encoder and other non-decoder modules remained frozen. To preserve the original VCTK speaker space, the first 109 rows of the speaker embedding table were locked by gradient masking and value restoration so that only the newly added ESD speaker embeddings were effectively updated.
The neutral adaptation data consisted of neutral ESD utterances mixed with VCTK utterances. No emotional utterances were used in this stage. The resulting neutral-adapted model is denoted as Matcha-Base, and it was used as the frozen backbone for emotion control training. For neutral adaptation, we used AdamW with a learning rate of and a weight decay of 0.01. The model was trained for 20,000 steps, and the checkpoint with the lowest validation loss was selected as Matcha-Base.
During emotion control training, Matcha-Base was frozen. The text encoder, main CFM decoder, main U-Net, and speaker embedding module were not updated. Only the Emo-ControlNet and the Emotional Duration Adapter were trained so that emotional expressiveness was introduced through residual control paths while preserving the acoustic and speaker-related structure learned during pretraining and neutral adaptation.
Emotion representations were extracted using the pretrained emotion2vec model (v2.0.4) [
28]. For each audio file, we extracted one 768-dimensional utterance-level embedding without additional VAD-based endpoint trimming or silence removal. When required by the extraction backend, the waveform was resampled to 16 kHz before emotion2vec inference; otherwise, the default preprocessing of the emotion2vec interface was used. The global neutral anchor was computed by averaging neutral ESD utterance embeddings from the training data across speakers. If an extracted feature were returned as a temporal matrix, then it was first averaged over time to obtain a single utterance-level vector. The DTRF then used the residual between the target emotion embedding and the global neutral anchor as the emotion control input.
Under our configuration, the trainable emotion control modules contain approximately 27M parameters, corresponding to 78.78% of the baseline decoder and 29.61% of the full model parameters. This represents a 70.39% reduction in trainable parameters relative to full-parameter updating, rather than a parameter count comparable to low-rank or bottleneck-adapter designs. Emotion control training was performed for 50,000 steps on a single NVIDIA RTX 4090 GPU with a batch size of 32. We used the Adam optimizer with a fixed learning rate of , no learning rate schedule, no weight decay, and no gradient clipping. The checkpoint with the lowest validation loss was selected for final evaluation.
The training loss followed
Section 3.1.2 and was the direct sum of the duration, prior, and flow-matching losses with equal weights. We set
in the OT-CFM interpolation. Mel-spectrograms used 80 bins, a 1024-point FFT, a window size of 1024, and a hop size of 256. The text encoder channel size was 192, the speaker embedding dimension was 64, and the decoder dropout rate was 0.05. The EDA input had
channels (
Section 3.2), formed by concatenating the 80-dimensional phoneme-level prior
and the repeated 768-dimensional emotion residual. Here,
denotes the phoneme-level acoustic prior rather than the 192-dimensional text encoder output. During inference, we used the explicit Euler ODE solver with NFE fixed to 50 for all formal evaluations. All internal systems shared the same acoustic feature configuration, vocoder, and inference step setting, so system comparisons were conducted under a fixed sampling budget. For this adaptation-focused comparison, we report the number of trainable parameters together with this fixed NFE setting. To further characterize the actual inference cost,
Section 4.4 reports the mel-spectrogram generation time, vocoder time, end-to-end synthesis time, real-time factor, generated waveform duration, and peak allocated GPU memory during inference under the same hardware and evaluation protocol.
4.1.3. Objective Evaluation Metrics
We evaluated generated speech from four perspectives: speaker similarity, emotion recognizability, text-content preservation, and affective trajectory structure. Speaker similarity is measured by Speaker Encoder Cosine Similarity (SECS). Speaker embeddings were extracted using a pretrained WavLM-based speaker verification encoder [
47], and each generated utterance was compared with a neutral reference utterance from the same target speaker.
Emotion recognizability was evaluated using an independent speech emotion recognition (SER) evaluator initialized from a WavLM Base Plus checkpoint. The evaluator was trained as a four-class classifier over the non-neutral ESD categories, namely angry, sad, happy, and surprise; neutral utterances were excluded. It was trained only on real ESD waveforms from the SER training file lists and was not used for emotion conditioning, neutral-anchor construction, or optimization in the DTRF. The best checkpoint was selected according to the unweighted average recall (UAR) validation macro. On the held-out real ESD SER test split containing 1457 four-emotion utterances after filtering, the evaluator achieved 91.21% accuracy and 91.22% SER-UAR. This split validated the independent SER evaluator only and was separate from the 500-utterance synthesis test set used for system comparison. We used SER-UAR as the primary objective emotion classification metric for synthesized speech. SER-Conf. denotes the mean posterior probability assigned by the independent SER evaluator to the target emotion class, and it was used in the residual strength () analysis.
We also report the Emotion2vec Cosine Similarity (E2V-CS) as an auxiliary embedding space measure of emotion alignment. For each generated utterance, an emotion2vec embedding was extracted and compared with the corresponding target-emotion anchor via cosine similarity. The target-emotion anchors were computed as the mean emotion2vec embedding of real training set utterances for each non-neutral emotion category. E2V-CS was macro-averaged over the four non-neutral emotion categories, excluding neutral samples. Since emotion2vec was also used to construct the conditioning residuals, E2V-CS was treated as a diagnostic measure and interpreted together with SER-UAR and subjective emotion ratings.
Text content preservation was estimated using the Automatic Speech Recognition Word Error Rate (ASR-WER). Generated utterances were transcribed using Whisper under the English automatic speech recognition setting [
48], and the word error rate (WER) was computed against the reference transcript. Since emotional speaking styles can affect ASR robustness, the ASR-WER was used as an automatic intelligibility proxy rather than as a direct perceptual measure.
Finally, we conducted a post hoc valence–arousal–dominance (VAD) projection to visualize affective trajectory trends. The projection was derived from emotion label posterior scores produced by the ModelScope emotion2vec+ base emotion recognition model [
28]. These posterior scores were mapped to VAD coordinates by a probability-weighted average over a fixed emotion-to-VAD prototype table.
For the objective comparison, all automatic metrics are summarized as corpus-level aggregates over the 500-utterance synthesized test set. SECS and E2V-CS were averaged over utterances, whereas the SER-UAR and ASR-WER were computed from aggregate classification and recognition results. The point estimates reported in
Table 2 were computed on the full test set using these aggregation rules. To quantify uncertainty, we additionally report the 95% confidence intervals estimated via bootstrap resampling over test utterances with 2000 replicates. For each bootstrap sample, SECS and E2V-CS were recomputed using the same utterance-level averaging and emotion-level macro-averaging rules as above; the SER-UAR was recomputed as the macro-averaged recall over the four non-neutral emotion categories, and the ASR-WER was recomputed from the aggregated word-level errors and reference word counts. The reported intervals correspond to the 2.5th and 97.5th percentiles of the bootstrap distribution. Bootstrap resampling was used only to estimate confidence intervals and did not alter the point estimates.
4.1.4. Subjective Evaluation Protocol
We conducted mean opinion score (MOS) listening tests with 10 participants. Each participant evaluated 42 audio samples: 32 samples for the main comparison experiment, covering eight system conditions across four emotion categories, and 10 samples for the residual strength () analysis, covering five control settings across two emotion categories. The eight main comparison conditions were GT, Matcha-Base, GST-Based, Full-Update, F5-TTS, DiEmo-TTS, Emosphere++, and DTRF. Here, GT denotes ground-truth recordings included as a natural speech reference rather than a synthesizer baseline. For each sample, listeners provided three independent ratings: naturalness MOS (NMOS), speaker similarity MOS (SMOS), and emotion MOS (EMOS).
All ratings were collected on a 5-point integer scale, where 1 denotes poor quality and 5 denotes excellent quality. The listening interface presented an audio player and three separate rating fields, and each sample could be replayed before rating. All systems were evaluated using the same interface and rating instructions. Participants were instructed to complete the test in a quiet environment using headphones or earphones at a comfortable and consistent playback volume.
Before the formal test, the participants were given rating instructions for the three dimensions. The NMOS evaluates perceived naturalness and the artifact level, the EMOS evaluates how clearly the utterance conveys the intended target emotion, and the SMOS evaluates perceived similarity to the target speaker. For the SMOS, a neutral reference utterance from the corresponding target speaker was provided together with the evaluated sample. For each model–emotion–metric condition, the MOS was computed as the mean score over the 10 participants. For the overall NMOS and SMOS, listener-level means were first computed across emotion categories for each system and then averaged across the participants.
We report the mean with the 95% confidence interval, computed using the Student-t distribution with degrees of freedom, where denotes the number of participants. The main comparison uses one utterance per model–emotion condition under the fixed listening protocol described above. Because the listening test involved only 10 participants, used one utterance per model–emotion condition in the main comparison, and did not include pairwise significance testing, the reported MOS differences should be interpreted as descriptive point estimates rather than as evidence of statistically significant superiority among closely matched systems.
4.1.5. Baseline Systems
We compare the DTRF with three internal baselines (Matcha-Base, Full-Update, and GST-Based) and three external representative systems (F5-TTS, DiEmo-TTS, and Emosphere++); the objective and subjective results are summarized in
Table 2 and
Table 3. The internal baselines used the same data split, acoustic feature configuration, vocoder, and evaluation protocol as the DTRF. Matcha-Base denotes the neutral-adapted Matcha-TTS backbone without emotion control injection and characterizes the frozen backbone before residual emotional adaptation. Full-Update uses the same emotion control architecture as the DTRF but jointly updates the acoustic backbone and emotion control modules on the emotion control training set. Global Style Token (GST)-Based keeps Matcha-Base frozen and trains a global style token conditioning module, including the reference encoder, attention layer, style tokens, and linear projection layer. During inference, GST-Based does not use test set reference audio; instead, emotion-specific GST prototypes are computed by averaging training set GST embeddings for each emotion category.
For external comparison, we included three systems with publicly available implementations that covered complementary perspectives. F5-TTS was included as a recent flow-matching TTS system, and it provides a generative reference under a high-quality flow-based formulation. Although it is not specialized for discrete emotion control, it serves as the primary flow-matching TTS baseline for overall speech quality, intelligibility, and emotion recognizability under a prompt-based control interface. DiEmo-TTS was included because it integrates self-supervised emotion representations into emotional TTS, which is directly relevant to our use of emotion2vec-based neutral-anchor residuals. Emosphere++ was included as a representative controllable emotional TTS system with explicit emotion modeling and prosody-aware synthesis. Together, these systems enable comparison against a flow-matching TTS reference, a self-supervised emotion representation-based emotional TTS model, and a dedicated controllable emotional synthesis model. Because the external systems differ in training data, architecture, speaker-conditioning interface, and waveform decoder, the comparison is intended as a representative system-level reference rather than a strictly controlled architectural ablation on the frozen Matcha-Base backbone studied here. Ground-truth recordings were additionally included in subjective evaluation as a natural speech reference upper bound, but they were not treated as a synthesizer baseline. All external outputs were processed with the same downstream metric pipeline as the internal systems. For the DTRF and the internal baselines, waveform reconstruction was performed with the same pretrained HiFi-GAN vocoder. All generated waveforms were resampled to 22.05 kHz when needed and evaluated using the same SECS, SER-UAR, E2V-CS, ASR-WER, and MOS protocols.
4.2. Main Results
4.2.1. Comparison with Baselines
Table 2 and
Table 3 summarize the objective and subjective comparison with the representative systems, respectively. The comparison includes three internal baselines built around the same Matcha-Base backbone, namely Matcha-Base, Full-Update, and GST-Based, as well as external systems including Emosphere++ [
49], DiEmo-TTS [
29], and F5-TTS [
10]. Ground-truth recordings were included as the reference condition.
Table 2 shows different trade-offs among speaker similarity, emotion recognizability, emotion-embedding alignment, and text preservation. Matcha-Base obtained high SECS and low ASR-WER values, indicating that the neutral-adapted frozen backbone preserved speaker identity and text content well, but its low SER-UAR showed limited emotional controllability without an explicit emotion control path. GST-Based maintained a similarly high SECS but yielded only a modest SER-UAR improvement, suggesting that global style token conditioning was less effective for stable emotion rendering in this frozen backbone setting.
Full-Update improved the SER-UAR compared with Matcha-Base and GST-Based, but this improvement was accompanied by a lower SECS and higher ASR-WER. This pattern reflects a stronger trade-off between emotional adaptation and acoustic preservation when the backbone is updated. Compared with the internal frozen backbone baselines, the DTRF substantially improved the emotion-related metrics; the SER-UAR increased from 0.302 for Matcha-Base and 0.378 for GST-Based to 0.606, and E2V-CS increased from 0.104 and 0.132 to 0.497, with confidence intervals that were clearly separated from those of the two frozen internal baselines for these emotion-related metrics in
Table 2. These results indicate that the proposed residual paths introduced effective emotion-related information under the frozen backbone adaptation setting.
The comparison with external systems should be interpreted more cautiously because these systems differ in training data, architecture, speaker conditioning interface, and waveform decoder. The DTRF achieved an SER-UAR comparable to F5-TTS and obtained the highest point estimate for E2V-CS among synthesized systems. However, the confidence intervals of several high-performing systems overlapped, so these results do not support an unambiguous ranking among external baselines. Instead, they indicate that the DTRF reached competitive emotion recognizability and emotion-embedding alignment while maintaining a higher SECS than F5-TTS, DiEmo-TTS, and Emosphere++ in this evaluation. Since E2V-CS is an auxiliary diagnostic metric, these objective results are interpreted together with the independent SER-UAR and subjective EMOS scores. Overall, the objective results support the main claim that dual-track residual adaptation improves emotional rendering over frozen internal baselines while preserving a reasonable balance among speaker similarity, intelligibility, and emotion recognizability.
In the subjective evaluation,
Table 3 shows that the DTRF obtained overall NMOS and EMOS point estimates that were competitive with the strongest external systems (e.g., overall NMOS of 3.95 for the DTRF versus 3.90 for F5-TTS). Matcha-Base obtained the highest mean SMOS, consistent with the absence of explicit emotional transformation; however, its EMOS scores remained low across all four emotions. The DTRF yielded mean EMOS values that were among the highest across the evaluated emotion categories. Given the limited listening test scale (10 participants, one utterance per model–emotion condition, and no significance testing), these small MOS differences are treated as descriptive rather than as evidence of statistically significant superiority over F5-TTS, Emosphere++, or DiEmo-TTS.
4.2.2. Residual Strength Control
Table 4 evaluates the effect of the residual scaling factor
on angry and sad, two representative emotions with distinct prosodic characteristics. The full four-emotion comparison is reported in
Table 2 and
Table 3. When
, both the acoustic residual and the duration residual were disabled. The resulting NMOS, SMOS, SECS, and ASR-WER values remained close to Matcha-Base, indicating that removing the residual scale returned the system to a neutral backbone-like operating condition. The EMOS values under
also remained low, which is consistent with limited target emotion expression.
These results indicate that provided an inference-time residual-strength control mechanism for emotion rendering, rather than a calibrated perceptual emotion intensity scale. From to , emotion-related scores such as the EMOS, SER-Conf., and E2V-CS increased substantially, while the NMOS and SMOS remained relatively stable in the angry and sad listening tests. From to , additional gains in emotion-related scores became smaller or less consistent across metrics, whereas degradation in the NMOS, SMOS, SECS, and ASR-WER became more pronounced. In our experiments, – offered a relatively stable operating range, while corresponded to an extrapolative setting that emphasized emotional salience at the cost of acoustic stability and intelligibility.
To examine whether the residual scaling produces structured affective movement, we visualized the synthesized speech using the post hoc emotion2vec-based VAD projection described above. For readability, the VAD visualizations focused on happy, angry, and sad, together with neutral as the reference category; surprise was included in the main objective and subjective evaluations but omitted from these trajectory plots.
Figure 5 visualizes the VAD distribution of real speech from the filtered ESD set; VAD coordinates were derived from emotion2vec+ posterior scores and a fixed emotion-to-VAD prototype table. Angry speech was associated with higher arousal and dominance and lower valence, whereas sad speech lied in the low-VAD region.
For generated speech, we used an additional visualization grid,
, to examine the movement of synthesized samples around the standard operating point. This visualization grid was used only to show the trajectory trend and not intended to exactly match the discrete settings in
Table 4. As illustrated in
Figure 6, samples generated with
were closer to the neutral region. As
increased, the trajectories tended to move in emotion-dependent directions; happy shifted upward in the VAD space, angry showed higher arousal and dominance while maintaining lower valence, and sad moved toward lower values across the three affective dimensions. These trends are broadly consistent with the quantitative residual strength results in
Table 4, although the VAD projection is a post hoc diagnostic rather than a direct measure of perceptual intensity control.
4.2.3. EDA Diagnostics
To examine the contribution of the duration branch, we conducted an inference-time diagnostic analysis on angry and sad, two representative emotions with distinct prosodic characteristics, at several residual-strength settings. The full four-emotion comparison is reported in
Table 2 and
Table 3. The trained EDA output was disabled during inference while the acoustic Emo-ControlNet remained active. The full and EDA-disabled conditions used the same checkpoint, target emotion, speaker, transcript, and acoustic residual scale. The resulting comparisons therefore isolated the temporal effect introduced by the EDA under an unchanged acoustic residual setting.
For the utterance-level duration change, each generated waveform is summarized by
where
denotes the waveform duration generated by the full model and
denotes the duration generated when the EDA output is disabled. The corresponding percentage duration change is
Table 5 reports this utterance-level comparison across
.
As shown in
Table 5, the EDA changed the utterance-level duration under all tested non-zero
values. The full model produced shorter utterances than the EDA-disabled configuration for both emotions, as indicated by the negative
values. The duration shift was stronger for angry and became slightly larger as
increased. For sad, the shift was smaller and gradually weakened at larger
values. This pattern suggests emotion-dependent temporal modulation rather than a fixed global duration offset.
To provide more direct evidence for the duration-level effect of the EDA, we further analyze the predicted phoneme-duration sequences at . Full DTRF was compared with an EDA-disabled setting in which the acoustic Emo-ControlNet remained active but the EDA output was set to zero (). For each generated utterance, we recorded the predicted phoneme-duration sequence after discretization and computed the mean phoneme duration and speaking rate in phonemes per second. We also report the mean and phoneme-level standard deviation of the predicted log-duration residual for Full DTRF to examine whether the EDA applied phoneme-dependent duration modulation rather than a single uniform offset.
As shown in
Table 6, enabling the EDA changed the phoneme-level timing statistics and speaking rate relative to EDA-disabled inference. For angry speech, Full DTRF reduced the mean phoneme duration from 82.39 ms to 70.49 ms (−14.4%) and increased the speaking rate from 12.27 to 14.27 phonemes/s (+16.3%). For sad speech, the adjustment was weaker; the mean phoneme duration decreased by 6.2%, and the speaking rate increased by 6.6%. These results are consistent with the utterance-level analysis in
Table 5 and indicate that the EDA affects duration prediction at the phoneme-sequence level rather than only changing the final waveform duration.
The mean was negative for both angry and sad, indicating an overall shortening tendency under the current setting, with a stronger average residual for angry. Meanwhile, the phoneme-level standard deviations of were clearly non-zero for both emotions, suggesting that the EDA predicted phoneme-dependent duration residuals instead of applying a single uniform utterance-level offset. Although both emotions showed an overall shortening tendency relative to EDA-disabled inference, the magnitude differed across emotions, with angry receiving substantially stronger duration compression than sad. This pattern suggests that the current EDA mainly learns emotion-dependent rate adjustment rather than a fixed emotion-specific duration rule. It should be emphasized that these diagnostics measure the effect of enabling the EDA relative to EDA-disabled inference under the same acoustic residual setting, rather than measuring absolute agreement with ground-truth emotional prosody. Since this diagnostic focuses on predicted duration sequences, more detailed pause-level, F0, and energy-based prosody analyses are left for future work.
4.3. Ablation Study
Table 7 analyzes the contribution of the main components in the DTRF using speaker similarity and emotion-embedding alignment. Removing the Emotional Duration Adapter (DTRF (a)) reduced the E2V-CS from 0.497 to 0.425, while the SECS remained close to the full model. This indicates that the duration branch contributes to emotion-related conditioning without substantially changing the speaker-identity structure. The inference-time duration diagnostics in
Table 5 and
Table 6 further show that the trained EDA affected both the utterance-level timing and phoneme-level duration statistics, supporting its role in temporal modulation.
Removing the Emo-ControlNet (DTRF (b)) caused a larger decline in the E2V-CS from 0.497 to 0.186, while the SECS increased to 0.905. This contrast indicates that the acoustic residual branch provides the main emotion-related acoustic modification. Without this branch, the generated speech remains closer to the frozen backbone and therefore preserves more speaker-related cues, but its emotion-embedding alignment is substantially weakened. The remaining positive E2V-CS suggests that the duration branch and the residual emotion representation still provided a limited degree of emotion-related conditioning, although they were not sufficient to replace the acoustic control branch.
Replacing the relative emotion residual with direct absolute feature injection (DTRF (c)) preserved part of the emotion-embedding alignment but reduced the SECS from 0.886 to 0.852. This result suggests that neutral-anchor residualization helps reduce speaker–emotion entanglement and improves speaker-identity preservation during emotional adaptation. Overall, the ablation results were consistent with the intended division of roles among the three components; the EDA contributed to duration-related emotion conditioning, the Emo-ControlNet provided the main acoustic residual control, and the neutral-anchor residual representation improved the balance between emotion rendering and speaker preservation.
4.4. Inference Efficiency
To complement the trainable parameter analysis, we further profiled the wall-clock inference efficiency of the internal systems. The benchmark was conducted on a single NVIDIA GeForce RTX 4090 GPU with a batch size of one. All internal systems used the same acoustic frontend, HiFi-GAN vocoder, text-preprocessing configuration and explicit Euler flow-matching solver with 50 function evaluations. The same 500-utterance synthesis test set was used. Before timing, 50 warm-up utterances were synthesized and excluded from the statistics. For each timed utterance, we measured the mel-spectrogram generation time, vocoder time, end-to-end synthesis time, generated waveform duration, real-time factor (RTF), and peak allocated GPU memory during inference. The end-to-end time is the sum of the mel-spectrogram generation time and vocoder time. The RTF was computed as the end-to-end synthesis time divided by the generated waveform duration. Per-utterance timing started after text preprocessing; file I/O, text preprocessing, and GST prototype construction were excluded from the measured latency. The reported peak memory corresponds to the PyTorch allocated GPU memory during timed inference using PyTorch version 2.6.0+cu118 with CUDA 11.8 and does not represent the total device-level memory.
As shown in
Table 8, Matcha-Base and GST-Based exhibited comparable inference costs, with end-to-end RTF values of 0.125 and 0.122, respectively. The DTRF increased the mel-spectrogram generation time compared with Matcha-Base because the acoustic residual branch and the Emotional Duration Adapter were evaluated during synthesis. The end-to-end synthesis time of the DTRF was 566.2 ms on average, corresponding to an RTF of 0.240. Although this was higher than the frozen Matcha-Base backbone, the RTF remained below 1.0, indicating faster-than-real-time synthesis on a single RTX 4090 GPU.
Full-Update and DTRF show comparable inference cost, consistent with their similar inference architectures during synthesis. Therefore, the advantage of DTRF is not lower inference latency than Full-Update. Rather, the DTRF achieved a comparable inference speed while avoiding full backbone updating during emotional adaptation. These results indicate that the DTRF trades a moderate inference overhead relative to the frozen neutral backbone for stronger emotional controllability, while retaining the practical benefit of frozen backbone adaptation with fewer trainable parameters.
4.5. Qualitative Analysis
To provide a qualitative view of the generated representations, we visualized a representative subset of speaker and emotion embeddings using t-SNE for neutral, happy, angry, and sad; surprise was omitted from this visualization but included in the quantitative results. This analysis was intended as supplementary visualization rather than an exhaustive evaluation across all emotion categories.
As shown in
Figure 7a, the ECAPA-TDNN [
50] speaker embeddings formed compact clusters by speaker identity, with different emotional categories distributed within each speaker region. This pattern is consistent with the SECS results and suggests that the generated samples retained speaker-dependent characteristics across emotional conditions.
Figure 7b shows the corresponding emotion2vec embedding space. The generated samples were organized primarily by emotional state across speakers, indicating that the synthesized speech contained emotion-related structure in the conditioning representation space. Since emotion2vec was also used to construct the emotion residuals, this visualization was treated as a supplementary representation-level analysis and considered together with the independent SER-UAR, human EMOS ratings, and VAD trajectories. Overall, the t-SNE results are interpreted as supplementary representation-level consistency checks, rather than independent evidence of perceived emotional quality.
5. Discussion
This study investigated emotional adaptation for a frozen flow-matching TTS backbone, with a focus on preserving speaker-related acoustic priors while improving controllable emotional rendering. The experimental results suggest that the DTRF provides a practical trade-off between emotional expressiveness and acoustic preservation. Compared with full-parameter updating, which directly modifies the shared acoustic model and may lead to speaker identity drift or reduced pronunciation stability, the DTRF confines emotion adaptation to residual control paths. This design helps retain the structure of the neutral-adapted backbone while introducing emotion-related acoustic and prosodic variation.
The results also highlight the importance of duration modeling in emotional speech synthesis. Emotional expression is not only reflected in spectral characteristics but also in phoneme duration, speaking rate, and rhythm. The ablation results show that removing the Emotional Duration Adapter reduced emotion-embedding alignment, while the inference-time duration diagnostics indicate that the duration branch changed both the utterance-level timing and phoneme-level duration statistics. These findings support the motivation for a dual-track design in which acoustic vector field control is complemented by duration-level temporal prosody control.
The residual-strength experiments further suggest that the scalar factor provides a practical interface for moving generated speech from a neutral-like operating point toward target emotion expression. At , the system remained close to the neutral backbone behavior, while increasing generally strengthened emotion-related scores such as the EMOS, SER-Conf., and E2V-CS. However, larger values of also tended to reduce speaker similarity, naturalness, and ASR-based text preservation. This pattern reflects an expressiveness–fidelity trade-off rather than a fully disentangled perceptual intensity control mechanism; moderate residual scaling (–) offered a relatively stable operating region in our angry and sad tests, whereas stronger extrapolative scaling emphasized emotional salience at the cost of acoustic stability and intelligibility.
The current study also has several limitations. First, the current evaluation followed a seen speaker but text-disjoint setting. The test utterances and transcripts were disjoint from the training data, which reduced text leakage and evaluated emotional rendering on unseen linguistic content. However, the ESD speakers were already introduced during neutral backbone adaptation. Therefore, the present results mainly characterize seen-speaker emotional adaptation under a frozen backbone setting, rather than speaker-disjoint or fully zero-shot emotional speech synthesis. Evaluating speaker transfer requires additional protocols such as leave-one-speaker-out training, speaker-disjoint ESD splits, or enrollment-based adaptation with unseen-speaker reference speech.
Second, the current control interface used an utterance-level residual-strength factor . Although provides a practical way to scale the learned acoustic and duration residuals, it should not be interpreted as a calibrated or fully disentangled perceptual emotion-intensity axis. It also does not yet provide explicit word-, phoneme-, or frame-level emotion trajectory control. Third, the subjective evaluation was limited in scale; it relied on 10 participants, one utterance per model–emotion condition in the main comparison, and mean MOS reporting without pairwise significance testing. Therefore, small differences among closely matched systems—such as the DTRF, F5-TTS, Emosphere++, and DiEmo-TTS—should not be over-interpreted as definitive ranking evidence, and the MOS results are best read together with the larger-scale objective metrics. Fourth, the emotion control evaluation focused on the English subset of ESD, leaving multilingual transfer, non-verbal affective events, and robustness under recording condition mismatching for future work. Finally, the current duration diagnostics provided evidence of utterance-level and phoneme-level temporal modulation, but they did not fully characterize prosody. Additional pause-level, F0-based, and energy-based analyses, together with spectral distortion measures, would provide a more complete prosodic and acoustic evaluation. Although we added inference efficiency measurements on a single RTX 4090 GPU, broader deployment-oriented profiling across different hardware platforms, optimized inference settings, streaming scenarios, and alternative vocoders remains an important direction for future work.
Future work will focus on multi-resolution emotion trajectories, non-verbal affective event modeling, speaker-disjoint and leave-one-speaker-out evaluation, and more perceptually calibrated intensity control under broader cross-lingual and robustness-oriented protocols.
6. Conclusions
In this work, we presented the DTRF, a dual-track residual framework for controllable emotional speech synthesis built upon a frozen Matcha-TTS backbone. The framework combines neutral-anchor emotion residualization, a zero-initialized Emo-ControlNet, and an Emotional Duration Adapter to introduce emotional variation through both acoustic vector field control and duration-level temporal control.
Experiments on the English subset of ESD suggest that the DTRF improves emotion-related metrics and subjective emotion ratings over the internal full-parameter updating and style token conditioning baselines, while maintaining a practical balance among naturalness, speaker similarity, and intelligibility. The trainable emotion control modules contain approximately 27 M parameters, corresponding to 29.61% of the full model parameters and a 70.39% reduction compared with full-parameter updating. The scalar factor provides an inference-time residual-strength control interface, enabling gradual movement from neutral-like speech toward target emotion expression while also revealing an expressiveness–fidelity trade-off at larger scaling values. The added inference efficiency analysis further shows that the DTRF introduces a moderate runtime overhead compared with the frozen Matcha-Base backbone but remains faster than real time on a single RTX 4090 GPU, with an end-to-end RTF of 0.240.
Representation-level analyses in VAD and embedding spaces showed trends that are consistent with the objective and subjective evaluations, while serving as supplementary visualizations rather than independent perceptual evidence. At the same time, the current system was evaluated mainly under a seen speaker, text-disjoint setting, and it does not yet address speaker-disjoint generalization, multilingual transfer, local time-varying emotion control, or explicit non-verbal affect modeling. Future work will focus on more disentangled and perceptually calibrated intensity control, hierarchical and local time-varying emotion trajectories, multi-condition affect encoders, broader speaker-disjoint and cross-lingual evaluations, and robustness-oriented protocols for practical emotional TTS deployment.
Author Contributions
Conceptualization, Y.D. and Y.G.; methodology, Y.D. and Y.G.; software, Y.G.; validation, Y.D. and Y.G.; formal analysis, Y.D., Y.G. and W.Y.; investigation, Y.D. and Y.G.; resources, Y.D. and F.C.; data curation, Y.G. and W.Y.; writing—original draft preparation, Y.D. and Y.G.; writing—review and editing, Y.D., Y.G., W.Y. and F.C.; visualization, Y.G.; supervision, Y.D. and F.C.; project administration, Y.D. and F.C. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Institutional Review Board Statement
Ethical review and approval were waived for this study because the listening test involved only noninvasive perceptual ratings of speech samples, and no sensitive personal information was collected.
Informed Consent Statement
Informed consent was obtained from all participants involved in the listening test.
Data Availability Statement
The public datasets used in this study were ESD and VCTK, which are available from their original providers under their respective licenses. Experimental configuration summaries, split descriptions, metric definitions, and aggregate evaluation results supporting the findings of this study are available at
https://github.com/Lemonnys/DTRF-Emotional-TTS-Supplementary (accessed on 25 June 2026).
Acknowledgments
The authors thank the participants involved in the listening test for their time and feedback.
Conflicts of Interest
The authors declare no conflicts of interest.
Abbreviations
The following abbreviations are used in this manuscript:
| ASR | Automatic speech recognition |
| ASR-WER | Automatic Speech Recognition Word Error Rate |
| CFM | Continuous flow matching |
| Conv1d | One-dimensional convolution |
| DTRF | Dual-Track Residual Framework |
| E2V-CS | Emotion2vec Cosine Similarity |
| EDA | Emotional Duration Adapter |
| EMOS | Emotion mean opinion score |
| ESD | Emotional Speech Dataset |
| GST | Global style token |
| HiFi-GAN | High-fidelity generative adversarial network |
| LayerNorm | Layer normalization |
| MAS | Monotonic alignment search |
| MLP | Multilayer perceptron |
| MOS | Mean opinion score |
| NFE | Number of function evaluations |
| NMOS | Naturalness mean opinion score |
| ODE | Ordinary differential equation |
| OT-CFM | Optimal Transport Conditional Flow Matching |
| ReLU | Rectified linear unit |
| RTF | Real-time factor |
| SECS | Speaker Encoder Cosine Similarity |
| SER | Speech emotion recognition |
| SiLU | Sigmoid linear unit |
| SMOS | Speaker similarity mean opinion score |
| TTS | Text-to-speech |
| UAR | Unweighted average recall |
| VAD | Valence–arousal–dominance |
| VCTK | Voice Cloning Toolkit Corpus |
| WER | Word error rate |
References
- Inoue, S.; Zhou, K.; Wang, S.; Li, H. Hierarchical Control of Emotion Rendering in Speech Synthesis. IEEE Trans. Affect. Comput. 2025, 16, 3316–3328. [Google Scholar] [CrossRef]
- Shen, K.; Ju, Z.; Tan, X.; Liu, E.; Leng, Y.; He, L.; Qin, T.; Zhao, S.; Bian, J. NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers. In Proceedings of the Twelfth International Conference on Learning Representations, Vienna, Austria, 7–11 May 2024. [Google Scholar]
- Ju, Z.; Wang, Y.; Shen, K.; Tan, X.; Xin, D.; Yang, D.; Liu, E.; Leng, Y.; Song, K.; Tang, S.; et al. NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models. In Proceedings of the 41st International Conference on Machine Learning; PMLR, Proceedings of Machine Learning Research: Norfolk, MA, USA, 2024; Volume 235, pp. 22605–22623. [Google Scholar]
- Lipman, Y.; Chen, R.T.Q.; Ben-Hamu, H.; Nickel, M.; Le, M. Flow Matching for Generative Modeling. In Proceedings of the International Conference on Learning Representations, Kigali, Rwanda, 1–5 May 2023. [Google Scholar]
- Tong, A.; Fatras, K.; Malkin, N.; Huguet, G.; Zhang, Y.; Rector-Brooks, J.; Wolf, G.; Bengio, Y. Improving and Generalizing Flow-Based Generative Models with Minibatch Optimal Transport. Trans. Mach. Learn. Res. 2024, 1–34. Available online: https://openreview.net/forum?id=CD9Snc73AW (accessed on 7 May 2026).
- Le, M.; Vyas, A.; Shi, B.; Karrer, B.; Sari, L.; Moritz, R.; Williamson, M.; Manohar, V.; Adi, Y.; Mahadeokar, J.; et al. Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale. Adv. Neural Inf. Process. Syst. 2023, 36, 14005–14034. [Google Scholar] [CrossRef]
- Guo, Y.; Du, C.; Ma, Z.; Chen, X.; Yu, K. VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching. In Proceedings of the ICASSP 2024—2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); IEEE: New York, NY, USA, 2024; pp. 11121–11125. [Google Scholar] [CrossRef]
- Mehta, S.; Tu, R.; Beskow, J.; Székely, É.; Henter, G.E. Matcha-TTS: A Fast TTS Architecture with Conditional Flow Matching. In Proceedings of the ICASSP 2024—2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); IEEE: New York, NY, USA, 2024; pp. 11341–11345. [Google Scholar] [CrossRef]
- Eskimez, S.E.; Wang, X.; Thakker, M.; Li, C.; Tsai, C.H.; Xiao, Z.; Yang, H.; Zhu, Z.; Tang, M.; Tan, X.; et al. E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS. In Proceedings of the 2024 IEEE Spoken Language Technology Workshop (SLT); IEEE: New York, NY, USA, 2024; pp. 682–689. [Google Scholar] [CrossRef]
- Chen, Y.; Niu, Z.; Ma, Z.; Deng, K.; Wang, C.; Zhao, J.; Yu, K.; Chen, X. F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long
38 Papers); Association for Computational Linguistics: Stroudsburg, PA, USA, 2025; pp. 6255–6271. [Google Scholar] [CrossRef]
- Wu, H.; Wang, X.; Eskimez, S.E.; Thakker, M.; Tompkins, D.; Tsai, C.H.; Li, C.; Xiao, Z.; Zhao, S.; Li, J.; et al. Laugh Now Cry Later: Controlling Time-Varying Emotional States of Flow-Matching-Based Zero-Shot Text-to-Speech. In Proceedings of the 2024 IEEE Spoken Language Technology Workshop (SLT); IEEE: New York, NY, USA, 2024; pp. 690–697. [Google Scholar] [CrossRef]
- Wu, P.; Ling, Z.; Liu, L.; Jiang, Y.; Wu, H.; Dai, L. End-to-End Emotional Speech Synthesis Using Style Tokens and Semi-Supervised Training. In Proceedings of the 2019 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC); IEEE: New York, NY, USA, 2019; pp. 623–627. [Google Scholar] [CrossRef]
- Byun, S.W.; Lee, S.P. Design of a Multi-Condition Emotional Speech Synthesizer. Appl. Sci. 2021, 11, 1144. [Google Scholar] [CrossRef]
- Dahmani, S.; Colotte, V.; Girard, V.; Ouni, S. Learning Emotions Latent Representation with CVAE for Text-Driven Expressive Audiovisual Speech Synthesis. Neural Netw. 2021, 141, 315–329. [Google Scholar] [CrossRef] [PubMed]
- Cai, X.; Dai, D.; Wu, Z.; Li, X.; Li, J.; Meng, H. Emotion Controllable Speech Synthesis Using Emotion-Unlabeled Dataset with the Assistance of Cross-Domain Speech Emotion Recognition. In Proceedings of the ICASSP 2021—2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); IEEE: New York, NY, USA, 2021; pp. 5734–5738. [Google Scholar] [CrossRef]
- Zhao, W.; Yang, Z. An Emotion Speech Synthesis Method Based on VITS. Appl. Sci. 2023, 13, 2225. [Google Scholar] [CrossRef]
- Zhou, K.; Sisman, B.; Rana, R.; Schuller, B.W.; Li, H. Speech Synthesis With Mixed Emotions. IEEE Trans. Affect. Comput. 2023, 14, 3120–3134. [Google Scholar] [CrossRef]
- Zhou, K.; Sisman, B.; Rana, R.; Schuller, B.W.; Li, H. Emotion Intensity and Its Control for Emotional Voice Conversion. IEEE Trans. Affect. Comput. 2023, 14, 31–48. [Google Scholar] [CrossRef]
- Wu, X.; Cao, Y.; Lu, H.; Liu, S.; Kang, S.; Wu, Z.; Liu, X.; Meng, H. Exemplar-Based Emotive Speech Synthesis. IEEE/ACM Trans. Audio Speech Lang. Process. 2021, 29, 874–886. [Google Scholar] [CrossRef]
- Lei, Y.; Yang, S.; Wang, X.; Xie, L. MsEmoTTS: Multi-Scale Emotion Transfer, Prediction, and Control for Emotional Speech Synthesis. IEEE/ACM Trans. Audio Speech Lang. Process. 2022, 30, 853–864. [Google Scholar] [CrossRef]
- Zhu, X.; Lei, Y.; Li, T.; Zhang, Y.; Zhou, H.; Lu, H.; Xie, L. METTS: Multilingual Emotional Text-to-Speech by Cross-Speaker and Cross-Lingual Emotion Transfer. IEEE/ACM Trans. Audio Speech Lang. Process. 2024, 32, 1506–1518. [Google Scholar] [CrossRef]
- Li, T.; Wang, X.; Xie, Q.; Wang, Z.; Xie, L. Cross-Speaker Emotion Disentangling and Transfer for End-to-End Speech Synthesis. IEEE/ACM Trans. Audio Speech Lang. Process. 2022, 30, 1448–1460. [Google Scholar] [CrossRef]
- Liu, R.; Liang, K.; Hu, D.; Li, T.; Yang, D.; Li, H. Noise Robust Cross-Speaker Emotion Transfer in TTS through Knowledge Distillation and Orthogonal Constraint. IEEE Trans. Audio Speech Lang. Process. 2025, 33, 812–827. [Google Scholar] [CrossRef]
- Wellington, S.; Liu, X.; Yamagishi, J. Quantifying Source Speaker Leakage in One-to-One Voice Conversion. In Proceedings of the 2024 International Conference of the Biometrics Special Interest Group (BIOSIG); IEEE: New York, NY, USA, 2024; pp. 1–6. [Google Scholar] [CrossRef]
- Yang, D.; Liu, S.; Huang, R.; Weng, C.; Meng, H. InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt. IEEE/ACM Trans. Audio Speech Lang. Process. 2024, 32, 2913–2925. [Google Scholar] [CrossRef]
- Hassani, S.M.; Kangavari, M.R. Mix-MaxETTS: A Text-to-Emotional Speech Synthesis Model Based on a Deep Encoder–Decoder Structure for the Transfer of Secondary Emotions. ETRI J. 2025, early view. [Google Scholar] [CrossRef]
- Ueda, L.H.; Lima, J.G.T.; Corrêa, P.R.; Simões, F.O.; Neto, M.U.; Costa, P.D.P. SelfTTS: Cross-Speaker Style Transfer through Explicit Embedding Disentanglement and Self-Refinement Using Self-Augmentation. arXiv 2026, arXiv:2603.22252. [Google Scholar] [CrossRef]
- Ma, Z.; Zheng, Z.; Ye, J.; Li, J.; Gao, Z.; Zhang, S.; Chen, X. emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2024; Association for Computational Linguistics: Stroudsburg, PA, USA, 2024; pp. 15747–15760. [Google Scholar] [CrossRef]
- Cho, D.H.; Oh, H.S.; Kim, S.B.; Lee, S.W. DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech. In Proceedings of the Annual Conference of the International Speech Communication Association (INTERSPEECH 2025), Rotterdam, The Netherlands, 17–21 August 2025. [Google Scholar] [CrossRef]
- Jeong, J.; Lee, Y.; Kwon, M.; Uh, Y. TTS-CtrlNet: Time Varying Emotion Aligned Text-to-Speech Generation with ControlNet. arXiv 2025, arXiv:2507.04349. [Google Scholar] [CrossRef]
- Yu, C.; Lu, H.; Hu, N.; Yu, M.; Weng, C.; Xu, K.; Liu, P.; Tuo, D.; Kang, S.; Lei, G.; et al. DurIAN: Duration Informed Attention Network for Speech Synthesis. In Proceedings of the Annual Conference of the International Speech Communication Association (INTERSPEECH 2020), Shanghai, China, 25–29 October 2020; pp. 2027–2031. [Google Scholar] [CrossRef]
- Ren, Y.; Hu, C.; Tan, X.; Qin, T.; Zhao, S.; Zhao, Z.; Liu, T.Y. FastSpeech 2: Fast and High-Quality End-to-End Text to Speech. In Proceedings of the International Conference on Learning Representations, Virtual, 3–7 May 2021. [Google Scholar]
- Kalyan, P.; Rao, P.; Jyothi, P.; Bhattacharyya, P. Emotion Arithmetic: Emotional Speech Synthesis via Weight Space Interpolation. In Proceedings of the Annual Conference of the International Speech Communication Association (INTERSPEECH 2024), Kos, Greece, 1–5 September 2024. [Google Scholar] [CrossRef]
- Liang, S.; Zhou, R.; Yuan, Q. ECE-TTS: A Zero-Shot Emotion Text-to-Speech Model with Simplified and Precise Control. Appl. Sci. 2025, 15, 5108. [Google Scholar] [CrossRef]
- Ilharco, G.; Ribeiro, M.T.; Wortsman, M.; Gururangan, S.; Schmidt, L.; Hajishirzi, H.; Farhadi, A. Editing Models with Task Arithmetic. In Proceedings of the International Conference on Learning Representations, Kigali, Rwanda, 1–5 May 2023. [Google Scholar]
- Moss, H.B.; Aggarwal, V.; Prateek, N.; González, J.; Barra-Chicote, R. BOFFIN TTS: Few-Shot Speaker Adaptation by Bayesian Optimization. In Proceedings of the ICASSP 2020—2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); IEEE: New York, NY, USA, 2020; pp. 7639–7643. [Google Scholar] [CrossRef]
- Kwon, K.J.; So, J.H.; Lee, S.H. Parameter-Efficient Fine-Tuning for Low-Resource Text-to-Speech via Cross-Lingual Continual Learning. In Proceedings of the Annual Conference of the International Speech Communication Association (INTERSPEECH 2025), Rotterdam, The Netherlands, 17–21 August 2025; pp. 1613–1617. [Google Scholar] [CrossRef]
- Hu, E.J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W. LoRA: Low-Rank Adaptation of Large Language Models. In Proceedings of the International Conference on Learning Representations, Virtual, 25–29 April 2022. [Google Scholar]
- Rücklé, A.; Geigle, G.; Glockner, M.; Beck, T.; Pfeiffer, J.; Reimers, N.; Gurevych, I. AdapterDrop: On the Efficiency of Adapters in Transformers. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing; Association for Computational Linguistics: Stroudsburg, PA, USA, 2021; pp. 7930–7946. [Google Scholar] [CrossRef]
- Zhang, L.; Rao, A.; Agrawala, M. Adding Conditional Control to Text-to-Image Diffusion Models. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2023; pp. 3836–3847. [Google Scholar] [CrossRef]
- Kim, J.; Kim, S.; Kong, J.; Yoon, S. Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment Search. Adv. Neural Inf. Process. Syst. 2020, 33, 8067–8077. [Google Scholar]
- Popov, V.; Vovk, I.; Gogoryan, V.; Sadekova, T.; Kudinov, M. Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech. In Proceedings of the 38th International Conference on Machine Learning; PMLR, Proceedings of Machine Learning Research: Norfolk, MA, USA, 2021; Volume 139, pp. 8599–8608. [Google Scholar]
- Kim, J.; Kong, J.; Son, J. Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech. In Proceedings of the 38th International Conference on Machine Learning; PMLR, Proceedings of Machine Learning Research: Norfolk, MA, USA, 2021; Volume 139, pp. 5530–5540. [Google Scholar]
- Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional Networks for Biomedical Image Segmentation. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention; Springer: Berlin/Heidelberg, Germany, 2015; pp. 234–241. [Google Scholar] [CrossRef]
- Zhou, K.; Sisman, B.; Liu, R.; Li, H. Emotional Voice Conversion: Theory, Databases and ESD. Speech Commun. 2022, 137, 1–18. [Google Scholar] [CrossRef]
- Yamagishi, J.; Veaux, C.; MacDonald, K. CSTR VCTK Corpus: English Multi-Speaker Corpus for CSTR Voice Cloning Toolkit (Version 0.92); Dataset; The University of Edinburgh: Edinburgh, UK, 2019. [Google Scholar] [CrossRef]
- Chen, S.; Wang, C.; Chen, Z.; Wu, Y.; Liu, S.; Chen, Z.; Li, J.; Kanda, N.; Yoshioka, T.; Xiao, X.; et al. WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing. IEEE J. Sel. Top. Signal Process. 2022, 16, 1505–1518. [Google Scholar] [CrossRef]
- Radford, A.; Kim, J.W.; Xu, T.; Brockman, G.; McLeavey, C.; Sutskever, I. Robust Speech Recognition via Large-Scale Weak Supervision. In Proceedings of the 40th International Conference on Machine Learning; PMLR, Proceedings of Machine Learning Research: Norfolk, MA, USA, 2023; Volume 202, pp. 28492–28518. [Google Scholar]
- Cho, D.H.; Oh, H.S.; Kim, S.B.; Lee, S.W. EmoSphere++: Emotion-Controllable Zero-Shot Text-to-Speech via Emotion-Adaptive Spherical Vector. IEEE Trans. Affect. Comput. 2025, 16, 2365–2380. [Google Scholar] [CrossRef]
- Desplanques, B.; Thienpondt, J.; Demuynck, K. ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification. In Proceedings of the Annual Conference of the International Speech Communication Association (INTERSPEECH 2020), Shanghai, China, 25–29 October 2020; pp. 3830–3834. [Google Scholar] [CrossRef]
Figure 1.
Overallarchitecture of DTRF with frozen flow-matching TTS backbone and dual-track emotion residual paths.
Figure 1.
Overallarchitecture of DTRF with frozen flow-matching TTS backbone and dual-track emotion residual paths.
Figure 2.
Training pipeline of DTRF under OT-CFM framework.
Figure 2.
Training pipeline of DTRF under OT-CFM framework.
Figure 3.
Inference pipeline of DTRF with residual-strength control via .
Figure 3.
Inference pipeline of DTRF with residual-strength control via .
Figure 4.
Acoustic residual injection via the Emo-ControlNet in the frozen flow prediction network.
Figure 4.
Acoustic residual injection via the Emo-ControlNet in the frozen flow prediction network.
Figure 5.
Valence–arousal–dominance (VAD) distribution of ground-truth ESD reference speech.
Figure 5.
Valence–arousal–dominance (VAD) distribution of ground-truth ESD reference speech.
Figure 6.
Valence–arousal–dominance (VAD) trajectories of synthesized speech under different residual strength settings.
Figure 6.
Valence–arousal–dominance (VAD) trajectories of synthesized speech under different residual strength settings.
Figure 7.
t-SNE visualization of generated speech embeddings for speaker and emotion representations.
Figure 7.
t-SNE visualization of generated speech embeddings for speaker and emotion representations.
Table 1.
ESD emotion control split statistics based on the current file lists under the seen speaker evaluation setting.
Table 1.
ESD emotion control split statistics based on the current file lists under the seen speaker evaluation setting.
| Split | Speakers | Utterances | Emotions | Unique Texts |
|---|
| Train | 10 | 13,300 | 4 | 480 |
| Validation | 10 | 200 | 4 | 157 |
| Test | 10 | 500 | 4 | 284 |
Table 2.
Objective evaluation results comparing DTRF with representative baseline models. Values are reported as point estimates with 95% confidence intervals in brackets.
Table 2.
Objective evaluation results comparing DTRF with representative baseline models. Values are reported as point estimates with 95% confidence intervals in brackets.
| Model | SECS ↑ | SER-UAR ↑ | E2V-CS ↑ | ASR-WER (%) ↓ |
|---|
| GT | | | | |
| Matcha-Base | | | | |
| GST-Based | | | | |
| Full-Update | | | | |
| Emosphere++ | | | | |
| DiEmo-TTS | | | | |
| F5-TTS | | | | |
| DTRF | | | | |
Table 3.
Subjective evaluation results reported as mean MOS ± 95% confidence interval. EMOS evaluated across four distinct emotion categories.
Table 3.
Subjective evaluation results reported as mean MOS ± 95% confidence interval. EMOS evaluated across four distinct emotion categories.
| Model | NMOS ↑ | SMOS ↑ | EMOS ↑ |
|---|
| Overall | Overall | Angry | Sad | Happy | Surprise |
|---|
| GT | | | | | | |
| Matcha-Base | | | | | | |
| GST-Based | | | | | | |
| Full-Update | | | | | | |
| Emosphere++ | | | | | | |
| DiEmo-TTS | | | | | | |
| F5-TTS | | | | | | |
| DTRF | | | | | | |
Table 4.
Impact of the residual-strength scaling factor on subjective and objective metrics. The results reflect an expressiveness–fidelity trade-off among emotional salience, speaker similarity, naturalness, and stability.
Table 4.
Impact of the residual-strength scaling factor on subjective and objective metrics. The results reflect an expressiveness–fidelity trade-off among emotional salience, speaker similarity, naturalness, and stability.
| Model | Emotion | NMOS ↑ | SMOS ↑ | EMOS ↑ | SECS ↑ | SER-Conf. ↑ | E2V-CS ↑ | ASR-WER (%) ↓ |
|---|
| Matcha-Base | Angry | | | | 0.913 | 0.224 | −0.249 | 2.04% |
| Sad | | | | 0.913 | 0.312 | 0.115 | 1.92% |
| DTRF () | Angry | | | | 0.905 | 0.246 | −0.221 | 2.13% |
| Sad | | | | 0.910 | 0.336 | 0.129 | 1.69% |
| DTRF () | Angry | | | | 0.886 | 0.604 | 0.497 | 4.19% |
| Sad | | | | 0.884 | 0.716 | 0.492 | 3.75% |
| DTRF () | Angry | | | | 0.852 | 0.648 | 0.664 | 5.12% |
| Sad | | | | 0.855 | 0.784 | 0.630 | 4.97% |
| DTRF () | Angry | | | | 0.848 | 0.652 | 0.731 | 7.58% |
| Sad | | | | 0.839 | 0.796 | 0.704 | 6.84% |
Table 5.
Utterance-level duration modulation produced by the Emotional Duration Adapter. is the log-duration ratio between the full model and the same model with EDA disabled. Dur. gives the corresponding percentage change in utterance duration.
Table 5.
Utterance-level duration modulation produced by the Emotional Duration Adapter. is the log-duration ratio between the full model and the same model with EDA disabled. Dur. gives the corresponding percentage change in utterance duration.
| Emotion | Setting | | Dur. (%) |
|---|
| Angry | | −0.153 | −14.2 |
| −0.163 | −15.0 |
| −0.170 | −15.6 |
| Sad | | −0.064 | −6.2 |
| −0.061 | −5.9 |
| −0.047 | −4.6 |
Table 6.
Phoneme-duration and speaking rate statistics for EDA diagnosis at .
Table 6.
Phoneme-duration and speaking rate statistics for EDA diagnosis at .
| Emotion | System | Mean Dur. (ms) | Rel. Change | Rate (phn/s) | Rel. Change | Mean log D | Std. log D |
|---|
| Angry | EDA-disabled | 82.39 | — | 12.27 | — | — | — |
| Angry | Full DTRF | 70.49 | −14.4% | 14.27 | +16.3% | −0.082 | 0.448 |
| Sad | EDA-disabled | 82.39 | — | 12.27 | — | — | — |
| Sad | Full DTRF | 77.28 | −6.2% | 13.08 | +6.6% | −0.035 | 0.534 |
Table 7.
Ablation study on the core components of the DTRF using speaker-similarity and emotion-embedding metrics.
Table 7.
Ablation study on the core components of the DTRF using speaker-similarity and emotion-embedding metrics.
| Config. | R-Res | C-Net | D-Adp | SECS ↑ | E2V-CS ↑ |
|---|
| DTRF | ✓ | ✓ | ✓ | 0.886 | 0.497 |
| DTRF (a) | ✓ | ✓ | × | 0.883 | 0.425 |
| DTRF (b) | ✓ | × | ✓ | 0.905 | 0.186 |
| DTRF (c) | × | ✓ | ✓ | 0.852 | 0.428 |
Table 8.
Inference efficiency of internal systems on a single NVIDIA GeForce RTX 4090 GPU. Values are reported as mean ± standard deviation over 500 utterances.
Table 8.
Inference efficiency of internal systems on a single NVIDIA GeForce RTX 4090 GPU. Values are reported as mean ± standard deviation over 500 utterances.
| Model | NFE | Mel Gen. (ms) ↓ | Vocoder (ms) ↓ | E2E Time (ms) ↓ | RTF ↓ | Audio Dur. (s) | Peak Infer. Alloc. Mem. (GB) ↓ |
|---|
| Matcha-Base | 50 | | | | | | 0.219 |
| GST-Based | 50 | | | | | | 0.225 |
| Full-Update | 50 | | | | | | 0.240 |
| DTRF | 50 | | | | | | 0.240 |
| Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |