1. Introduction
Multimodal sentiment analysis (MSA) aims to infer human affective states by jointly exploiting linguistic, acoustic, and visual cues. Compared with text-only sentiment analysis, MSA is better aligned with real-world human communication, where sentiment is expressed not only through words but also through prosody, facial movements, pauses, gaze, and other nonverbal signals [
1]. This makes MSA particularly relevant to opinion video understanding, human–computer interaction, online education, intelligent customer service, affective dialogue systems, and social media analytics.
From a machine learning perspective, MSA is a representative multimodal learning problem because it requires models to handle heterogeneous input spaces, asynchronous temporal structures, modality imbalance, and cross-modal complementarity [
2]. Text usually carries explicit semantic information, while audio and visual streams provide paralinguistic and behavioral evidence that may strengthen, weaken, or even contradict the textual signal. Therefore, the central technical issue in MSA is not merely whether multiple modalities are available, but how to integrate them in a way that preserves task-relevant information while suppressing noisy or redundant cross-modal interactions.
Recent surveys show that MSA research has evolved from simple early/late fusion to tensor fusion, hierarchical fusion, attention-based fusion, pretrained-language-model-based fusion, self-supervised multimodal learning, and missing-modality-aware learning [
3]. In parallel, affective computing studies have demonstrated that multimodal signals can improve affect recognition, but the gain depends strongly on the quality of feature extraction, temporal alignment, fusion strategy, and the robustness of the model under noisy or incomplete inputs [
4]. These observations indicate that high-performing MSA models should not only pursue stronger feature interaction but also provide stable predictions under practical deployment constraints.
A major limitation of many existing MSA methods is the assumption that all modalities are available and reliable during both training and testing. In realistic scenarios, however, one or more modalities may be missing because of speech recognition failure, camera occlusion, low-quality audio, privacy restrictions, sensor malfunction, transmission loss, or platform-side data filtering [
5]. Missing-modality learning has therefore become an important research direction, because a model trained only on complete multimodal samples can suffer a substantial performance drop when partial modalities are absent at inference time [
6]. Recent tag-assisted and representation-alignment methods further show that explicit missing-pattern modeling and complete-to-incomplete representation alignment can improve robustness under uncertain modality availability [
7,
8].
Another practical challenge is sample-dependent modality reliability. Multimodal sentiment labels are inherently subjective, and textual, acoustic, and visual streams may provide inconsistent or even contradictory evidence. For example, a speaker may express positive words with a negative tone, or a visually neutral expression may accompany emotionally loaded language. Under such cases, deterministic fusion may over-amplify unreliable nonverbal cues and produce unstable predictions. Therefore, this study treats nonverbal reliability primarily as a modality-contribution regulation problem rather than as a fully probabilistic uncertainty-estimation problem. Acoustic and visual signals are introduced as residual correction cues, and their contribution is controlled by a learned residual reliability score. This design allows nonverbal modalities to refine the text-based sentiment estimate when they provide useful affective evidence, while reducing their influence when they are noisy, missing, or potentially misleading.
To address these issues, this study proposes UCMR-Net, a residual fusion framework for multimodal sentiment intensity prediction that uses text as the primary semantic anchor and incorporates adaptive residual weighting and masked-modality distillation as auxiliary design components. The model follows a text-anchored design: textual representations are treated as the primary semantic backbone, while acoustic and visual modalities are used as residual correction signals. This design is motivated by the observation that sentiment polarity and intensity in opinion videos are often explicitly expressed through language, whereas audio and visual modalities provide complementary but sometimes noisy contextual cues.
Specifically, UCMR-Net integrates four design principles. First, it uses contextual text representations as the semantic anchor for multimodal prediction. Second, it encodes acoustic and visual temporal streams as nonverbal residual evidence. Third, it performs text-anchored residual fusion to refine the text-based sentiment estimate without uniformly amplifying all nonverbal information. Fourth, it uses adaptive residual weighting, counterfactual masked-modality distillation, and auxiliary polarity/intensity supervision to stabilize prediction behavior.
The main contributions of UCMR-Net and the corresponding empirical evaluation framework are summarized as follows. First, we develop UCMR-Net as an integrative and application-oriented text-anchored residual fusion framework that treats language as the primary semantic backbone and models acoustic and visual streams as adaptive residual modifiers for continuous sentiment intensity estimation. Second, we formulate the learned residual coefficient as a residual reliability score for modality-contribution regulation and evaluate its empirical value through calibration analysis, residual reliability–error correlation, and modality corruption tests. Third, we provide a controlled experimental protocol including five random seeds, unified baseline evaluation, statistical significance testing, and ablations that separately isolate text anchoring, residual fusion, adaptive residual weighting, counterfactual distillation, and multi-task supervision. Fourth, we evaluate incomplete-input behavior under a unified missing-modality protocol and random missing-rate protocol, and further assess generalization using CMU-MOSEI external validation. These analyses position UCMR-Net as a text-dominant residual fusion model with strong complete-modality performance and moderate robustness when nonverbal modalities are missing or corrupted. The methodological contribution is therefore primarily integrative and application-oriented: UCMR-Net combines text anchoring, residual fusion, adaptive residual contribution regulation, masked-modality distillation, and auxiliary supervision within a unified framework, rather than claiming a fundamentally new fusion primitive.
The rest of this paper is organized as follows.
Section 2 reviews related studies on multimodal sentiment fusion, missing-modality robust learning, and adaptive modality-contribution learning.
Section 3 introduces the proposed UCMR-Net framework, including the problem definition, model architecture, module design, and training objective.
Section 4 reports the experimental setup and results.
Section 5 discusses the empirical findings, limitations, and future research directions.
3. Methodology
3.1. Overall Research Workflow
This study develops an application-oriented multimodal sentiment prediction framework for text-anchored sentiment intensity estimation. The overall workflow consists of six stages: dataset preparation, multimodal preprocessing, modality-specific encoding, adaptive text-anchored residual fusion, multi-task optimization, and unified experimental evaluation. To avoid inconsistent evaluation across complete and incomplete inputs, all complete-modality, missing-modality, ablation, calibration, and baseline experiments are conducted under the same data split, preprocessing pipeline, feature normalization parameters, checkpoint selection rule, and metric script. For stochastic reliability, the main model, all rerun baselines, and controlled ablation configurations are evaluated using five random seeds. Repeated-run results are reported as mean ± standard deviation unless otherwise specified.
Figure 1 illustrates the overall workflow from data preparation to unified result analysis.
3.2. Problem Definition
Given a multimodal utterance-level sample i, the input consists of three modality streams: textual input , acoustic input , and visual input . The target variable is the continuous sentiment intensity score , where negative values indicate negative sentiment, positive values indicate positive sentiment, and values around zero indicate neutral sentiment. The primary task is regression-based sentiment intensity prediction, while binary polarity prediction and ordinal sentiment-level prediction are introduced as auxiliary tasks.
The general prediction function is defined as:
where
denotes the modality availability mask, and
represents all trainable parameters. The binary polarity label is defined as:
For ordinal auxiliary learning, the continuous sentiment score is discretized into ordered sentiment-intensity classes:
The model is optimized to minimize regression error, preserve rank-consistent sentiment intensity, and improve polarity-level discriminability.
3.3. Proposed Model: UCMR-Net
The proposed model is named UCMR-Net. It is designed as a text-anchored residual fusion network with adaptive residual weighting and masked-modality distillation for continuous multimodal sentiment intensity prediction. UCMR-Net is built on a text-anchored assumption: in opinion video sentiment analysis, textual content usually carries the densest semantic evidence, whereas acoustic and visual streams provide complementary paralinguistic corrections. The model first obtains a contextual text representation from BERT, then encodes acoustic and visual temporal streams using BiLSTM-based encoders, and finally performs residual correction through a learned residual reliability score. The reliability score is used to regulate nonverbal residual contribution and is not interpreted as a direct estimate of aleatoric uncertainty, epistemic uncertainty, or predictive uncertainty. The model jointly outputs a continuous sentiment score, a binary polarity label, and an ordinal sentiment-intensity label. The overall architecture of UCMR-Net, including modality encoders, masked-modality compensation, residual fusion, and prediction heads, is shown in
Figure 2.
3.4. Text Encoder
The textual stream is encoded using BERT-base-uncased, which provides contextualized bidirectional token representations. For an input token sequence
, BERT generates hidden states:
The final text representation combines the [CLS] token representation, attention pooling, and mean pooling, followed by a projection layer to match the hidden dimension used by the acoustic and visual encoders:
where the projected text vector is used as the semantic anchor for subsequent residual fusion. BERT is used because pretrained bidirectional contextual representations have been shown to provide strong language understanding capacity for downstream NLP tasks [
31].
Figure 3 presents the BERT-based text encoding module and shows how token-level contextual representations are aggregated into the final textual feature vector.
3.5. Acoustic and Visual Temporal Encoders
The acoustic and visual streams are represented as time-series feature sequences. In the implemented feature pipeline, the acoustic feature dimension is 74 and the visual feature dimension is 35. Each stream is normalized and then fed into a bidirectional LSTM encoder. LSTM is used because it is suitable for modeling temporal dependencies in sequential data [
32].
For the visual stream:
where
. This design ensures that the acoustic and visual representations are projected into the same latent dimensionality as the textual representation before fusion, as illustrated in
Figure 4.
3.6. Text-Anchored Residual Fusion
The central component of UCMR-Net is the text-anchored residual fusion module. Rather than treating the three modalities as equally reliable at all times, UCMR-Net uses text as the main semantic anchor and uses acoustic and visual representations as residual correction signals. This is particularly suitable for CMU-MOSI, where sentiment labels are assigned to opinion segments and textual statements often provide the most explicit sentiment evidence.
The acoustic and visual representations are first projected into residual feature space:
The multimodal residual signal is then computed as:
where the residual estimator is implemented by fully connected layers, layer normalization, GELU activation, and dropout. The fused representation is obtained by adding a reliability-weighted nonverbal residual correction to the textual representation:
Here, the residual weight is produced by a positive scalar transformation and normalized across residual branches. In this formulation, the scalar is interpreted as a learned residual reliability controller rather than a probabilistic uncertainty estimate. A lower residual reliability score increases the contribution of the corresponding residual branch, whereas a higher reliability-risk score suppresses potentially noisy nonverbal correction. The empirical validity of this score is examined through its correlation with the absolute prediction error and through modality corruption tests, rather than assumed from the Softplus transformation alone.
This fusion design makes the model different from conventional early fusion and tensor fusion models. It explicitly prioritizes the textual semantic backbone while allowing nonverbal modalities to refine sentiment intensity estimation. The internal structure of the text-anchored residual fusion module is presented in
Figure 5.
3.7. Missing-Modality Compensation and Counterfactual Distillation
The framework can be evaluated under incomplete multimodal conditions. When one or more modalities are unavailable, the missing stream is replaced by a controlled placeholder representation, and the available modalities are used to generate a prediction. At the framework level, this operation is implemented as mask-based missing-modality compensation. In UCMR-Net, the missing pathway uses unavailable-stream masking, zero-filled modality placeholders, and counterfactual distillation rather than an independent generative decoder.
Let
denote the prediction under complete modalities and
denote the prediction under a missing-modality mask. The counterfactual distillation term is defined as:
The counterfactual distillation term encourages masked-input predictions to remain close to complete-input predictions under controlled missing-modality masks. This objective is intended to reduce prediction drift caused by modality unavailability, but it does not guarantee correctness when the removed modality contains decisive or contradictory sentiment evidence. Therefore, the effectiveness of masked-modality learning is evaluated empirically under a unified missing-modality protocol covering complete input, missing text, missing audio, missing visual, unimodal-only input, and random missing rates.
Figure 6 summarizes the missing-modality compensation and counterfactual distillation mechanism used to constrain prediction consistency under masked-modality conditions.
3.8. Adaptive Residual Reliability and Calibration Analysis
UCMR-Net uses a learned residual reliability score to regulate the contribution of acoustic and visual residual corrections. This score is not treated as a formal probabilistic uncertainty estimate. Instead, it is analyzed as a sample-dependent gating signal whose empirical behavior can be assessed through error correlation, calibration metrics, and modality corruption tests. For a modality representation, the residual reliability score is computed as:
where a small constant is used for numerical stability. Calibration is evaluated using ECE, MCE, Brier score, and NLL. Because raw neural confidence is often miscalibrated, post-hoc temperature scaling is additionally applied to the binary polarity head. The evaluation reports both raw and temperature-scaled calibration results, as well as the correlation between residual reliability-risk scores and absolute regression errors.
Figure 7 shows how adaptive residual reliability is connected with residual weighting and multi-task prediction heads.
3.9. Multi-Task Prediction Heads
The fused representation
is passed to three task-specific heads. The regression head predicts continuous sentiment intensity:
The binary head predicts sentiment polarity:
The ordinal head predicts ordered sentiment-intensity categories:
The binary and ordinal tasks serve as auxiliary supervision signals and help the model distinguish polarity and intensity levels beyond pure regression.
3.10. Objective Function
The total loss combines regression loss, correlation-oriented loss, counterfactual distillation, and auxiliary multi-task loss. The main regression loss is composed of smooth L1 loss, Concordance Correlation Coefficient (CCC) loss [
33], and Pearson correlation loss:
The Pearson correlation coefficient is:
The concordance correlation coefficient is:
The binary polarity loss is:
The ordinal classification loss is:
The final training objective is:
where
and
control the contribution of counterfactual distillation and auxiliary multi-task supervision.
3.11. Training Strategy
The training procedure follows a two-phase strategy. In Phase 1, the BERT backbone is frozen to stabilize the multimodal fusion layers, with a higher missing rate and stronger distillation weight. In Phase 2, BERT is partially unfrozen for fine-tuning, while the missing rate and distillation weight are reduced. This design prevents early-stage overfitting of the large text encoder and allows the multimodal fusion module to first learn stable cross-modal correction behavior. The detailed two-phase training configuration of UCMR-Net is summarized in
Table 1.
Implementation and reproducibility details. All experiments were implemented in Python 3.10 using PyTorch and HuggingFace Transformers. The 74-dimensional acoustic features and 35-dimensional visual features were obtained from the aligned CMU-MOSI feature release distributed through the CMU Multimodal SDK/MultiBench-style pipeline, corresponding to COVAREP acoustic descriptors and FACET visual descriptors. The optimizer was AdamW. The learning rate was set to for the unfrozen BERT layers and 1 × 10−3 for newly initialized non-text modules, with weight decay of , dropout of 0.30, batch size of 32, and early stopping patience of 8 epochs based on validation MAE. The model was trained with five random seeds (2026, 2027, 2028, 2029, and 2030). Checkpoints were selected exclusively according to validation MAE, and the test set was used only for final evaluation after the model configuration was fixed. Main baseline and controlled-ablation results are summarized as mean ± standard deviation across the five validation-selected seed-level runs. Separately, a locked ensemble was formed by equally averaging the five seed-level predictions and was used only for the post-hoc missing-modality, calibration, residual-reliability, modality-corruption, error-distribution, and qualitative analyses explicitly identified as ensemble-based.
4. Experimental Results
4.1. Dataset Description
The experiments are conducted on the CMU-MOSI dataset, a widely used benchmark for multimodal sentiment analysis. CMU-MOSI contains opinion-level video segments with aligned textual, acoustic, and visual information. Each utterance is annotated with a continuous sentiment intensity score in the range
, where
denotes strongly negative sentiment and
denotes strongly positive sentiment [
34]. CMU-MOSEI is additionally used in
Section 4.9 as an external validation benchmark for assessing generalization beyond CMU-MOSI [
35]. The datasets are accessed and processed using the CMU Multimodal SDK and MultiBench-style feature pipelines [
36].
The evaluation pipeline used in this study provides prediction and label arrays for the CMU-MOSI test split containing 686 samples. Based on the test labels used in the evaluation pipeline, the test set contains 379 negative samples, 30 neutral samples, and 277 positive samples. Regression metrics are computed on all 686 samples, while binary polarity metrics follow the Non0 protocol and are computed on the 656 non-neutral samples.
Table 2 summarizes the dataset scale, split setting, modality composition, target variables, and test-set label distribution.
To further visualize the label composition of the CMU-MOSI test set,
Figure 8 shows the distribution of negative, neutral, and positive samples.
4.2. Data Preprocessing
The preprocessing procedure follows the implementation used in the UCMR-Net training pipeline. Text is reconstructed from word-level features by removing special pause tokens and concatenating valid word tokens. Acoustic and visual features are cleaned by replacing invalid numerical values, including NaN and positive or negative infinity, with zeros. Z-score normalization is estimated from the training set and then applied to the validation and test sets to avoid test-set information leakage. Extremely large normalized values are clipped to the range for numerical stability.
The z-score normalization is defined as:
where
and
are computed from the training split only. The preprocessing operations applied to each modality and label type are summarized in
Table 3.
4.3. Baseline Models
The baseline selection covers representative multimodal sentiment analysis models and incomplete-modality methods. Early-fusion and late-fusion LSTM baselines are included as simple sequential fusion references. TFN and LMF represent tensor-based multimodal fusion [
37,
38]. MulT introduces cross-modal attention for unaligned multimodal language sequences [
39]. MISA separates modality-invariant and modality-specific representations [
40]. MAG-BERT integrates acoustic and visual information into pretrained language models through multimodal adaptation gates [
41]. Self-MM introduces self-supervised modality-specific labels for multi-task multimodal sentiment learning [
42]. MMIM maximizes mutual information hierarchically to preserve task-relevant multimodal information [
43]. TFR-Net, UMDF, and CorrKD are included as representative incomplete-modality or missing-modality robust methods [
44,
45,
46].
All key baselines were evaluated under the same CMU-MOSI split, text/audio/visual feature pipeline, preprocessing statistics, validation-based checkpoint selection rule, random seed protocol, and metric script as UCMR-Net. Public implementations were used when a verifiable author-provided or paper-associated repository was available; otherwise, the corresponding model was reimplemented according to its original methodological specification. Adaptations were restricted to the data interface, input dimensionality, prediction head, and unified evaluation protocol, while model-specific architectural components and original auxiliary objectives were retained.
Table 4 summarizes the baseline set and unified comparison status, and
Table 5 reports the implementation source, adaptation details, major hyperparameters, and configuration file location for each rerun model. Detailed implementation-source records, adaptation notes, and per-model configuration files are provided in the anonymized reproducibility repository.
To furthr improve the reproducibility of the unified comparison,
Table 5 summarizes the implementation source, adaptation details, major hyperparameters, and configuration file location for each rerun model. All models used the same CMU-MOSI split, training-set-derived preprocessing statistics, validation-based checkpoint selection criterion, five random seeds, and evaluation script, while model-specific architectural components and original auxiliary objectives were retained whenever applicable.
4.4. Evaluation Metrics
The primary regression metrics are MAE, RMSE, and Pearson correlation. MAE measures the average absolute prediction error, while RMSE penalizes large errors more strongly [
47]. Pearson correlation evaluates the linear consistency between predicted and true sentiment scores. For regression evaluation, MAE, RMSE, and Pearson correlation are computed on all 686 CMU-MOSI test samples. For binary polarity evaluation, Acc-2 and F1 follow the Non0 protocol commonly used in CMU-MOSI comparisons: samples with y = 0 are excluded, and polarity is defined as negative when y < 0 and positive when y > 0. Therefore, binary metrics are computed on 656 non-neutral test samples rather than on the full 686-sample test set. Unless otherwise stated, F1 denotes the positive-class F1 under this Non0 protocol. Calibration is assessed using ECE, MCE, Brier score, and NLL. ECE and reliability-based calibration analysis are widely used for modern neural networks [
48], and the Brier score is a classical probabilistic forecast verification metric [
49].
The mean absolute error is:
The root mean square error is:
The expected calibration error is:
The maximum calibration error is:
The negative log-likelihood is:
Table 6 summarizes the evaluation metrics used in this study, including their metric type, optimization direction, and analytical role.
Reporting convention: Repeated-run CMU-MOSI training results are reported as mean ± standard deviation (SD) across five random seeds unless otherwise stated. Post-hoc analyses computed from the locked validation-selected ensemble predictions are reported as single values because between-seed variability is not defined for those derived analyses. External-validation and cross-dataset transfer results are reported according to the evaluation basis specified in the corresponding table notes.
4.5. Unified Baseline Comparison and Five-Seed Main Results
To ensure a fair model comparison, representative baselines were evaluated under the same split, preprocessing pipeline, feature normalization, validation-based checkpoint selection, five-seed protocol, and metric script as UCMR-Net.
Table 7 reports the mean ± standard deviation across five independent runs for all rerun models. UCMR-Net achieves the lowest MAE and RMSE and the highest Pearson correlation among the evaluated methods, while the improvements in Acc-2 and F1 are comparatively smaller.
The unified comparison shows that UCMR-Net achieves the lowest MAE (0.699 ± 0.009) and RMSE (0.997 ± 0.011) and the highest Pearson correlation (0.802 ± 0.006) among the evaluated models. The improvements in Acc-2 and F1 are comparatively smaller, so the interpretation emphasizes regression and sentiment-intensity estimation rather than broad categorical superiority.
Table 8 reports the seed-level CMU-MOSI results and the corresponding mean ± standard deviation under the five-seed evaluation protocol. Reporting both the central tendency and between-run variability provides a stability-aware estimate of model performance under random initialization and training variation.
To test whether the regression gains are reliable across repeated runs,
Table 9 summarizes paired significance tests between UCMR-Net and strong baselines. The MAE comparisons provide the primary statistical evidence for improvements over BERT-only, MAG-BERT, CorrKD, and LNLN. The Corr comparison with LNLN (
p = 0.044) is treated as secondary evidence and is not used to support a strong superiority claim, while the Acc-2 and F1 differences are not statistically significant.
To visualize both model-level comparison and run-to-run stability,
Figure 9 presents the MAE/Corr comparison across evaluated baselines and the seed-level variation of UCMR-Net under the five-seed protocol.
Figure 9a shows the baseline comparison, whereas
Figure 9b illustrates the seed-level variability of UCMR-Net and the stability of its mean performance across repeated runs.
4.6. Controlled Ablation Study
Table 10 presents the controlled ablation study, with all results reported as mean ± standard deviation across five random seeds. All configurations use the same data split, preprocessing pipeline, training schedule, random seed set, validation-based checkpoint selection rule, and evaluation protocol. Only the investigated component is modified in each ablation.
The largest and most consistent degradation is observed when text anchoring is removed, with MAE increasing from 0.699 ± 0.009 to 0.797 ± 0.018. Removing residual fusion produces the second-largest degradation, increasing MAE to 0.739 ± 0.014. Removing adaptive residual weighting, counterfactual distillation, or individual auxiliary supervision components results in smaller but consistent performance decreases. The increased variability of the configuration without text anchoring further suggests that the textual backbone contributes not only to predictive accuracy but also to training stability.
Figure 10 visualizes the delta MAE values from
Table 10, making the relative contribution of each component explicit.
The visualization indicates that removing text anchoring produces the largest degradation, followed by residual fusion, whereas adaptive residual weighting and auxiliary objectives provide smaller but consistent contributions.
4.7. Unified Missing-Modality and Random Missing-Rate Analysis
The missing-modality evaluation uses the same trained model family, normalization parameters, ensemble workflow, and metric script as the complete-modality test. As shown in
Table 11, the Complete condition matches the main five-seed result, ensuring that complete and incomplete input settings are compared under an aligned evaluation pipeline.
Performance remains close to the complete-input result when audio or visual information is removed, but it decreases substantially when text is unavailable. This pattern supports a cautious conclusion: UCMR-Net is robust to the absence of nonverbal streams to a moderate extent, but it is not a modality-symmetric missing-modality model.
To further examine robustness under stochastic input degradation,
Table 12 reports random missing rates from 10% to 50%. Performance decreases gradually as the missing rate increases, with MAE rising from 0.699 at 0% missing rate to 0.829 at 50% missing rate and Corr decreasing from 0.802 to 0.711.
Figure 11 summarizes the unified missing-modality and random missing-rate analyses by combining the single-modality removal results with the random missing-rate trend.
The visualization shows smooth degradation under random missingness and a substantially larger performance drop when the textual stream is unavailable.
4.8. Calibration, Residual Reliability, and Modality Corruption Analysis
The calibration results indicate that raw UCMR-Net provides only limited calibration improvement over the main baseline in ECE, MCE, Brier score, and NLL, and these differences are insufficient to support a claim of comprehensive probabilistic calibration.
Table 13 therefore reports both raw and temperature-scaled calibration metrics.
After post-hoc temperature scaling, ECE decreases from 0.121 to 0.064, MCE from 0.229 to 0.142, Brier score from 0.136 to 0.128, and NLL from 2.337 to 2.286. These results indicate improved post-hoc calibration, while the raw residual reliability score should not be interpreted as a complete uncertainty estimator.
Table 14 further examines whether the residual reliability-risk score is related to prediction error. The positive Spearman and Pearson correlations indicate that larger risk scores tend to accompany larger absolute errors, and the quartile analysis shows a clear increase in MAE from the lowest to the highest reliability-risk group.
To test whether the learned residual weights respond to degraded modality quality,
Table 15 reports audio and visual corruption results. As corruption increases, MAE and RMSE increase while the corresponding mean residual weight decreases, suggesting that adaptive residual weighting responds to lower nonverbal reliability.
To further interpret the calibration and reliability analyses,
Figure 12 visualizes the confidence calibration behavior of raw UCMR-Net and temperature-scaled UCMR-Net, while
Figure 13 presents the empirical behavior of the residual reliability-risk score by combining error stratification across reliability-risk quartiles with residual-weight responses under audio and visual corruption.
The reliability diagram shows that post-hoc temperature scaling improves the alignment between predicted confidence and empirical accuracy, whereas raw UCMR-Net outputs remain less well calibrated.
The observed quartile and corruption patterns support the interpretation of the learned score as a residual reliability controller, because higher reliability-risk scores are associated with larger errors and corrupted nonverbal streams receive lower residual weights.
4.9. CMU-MOSEI External Validation and Transfer Analysis
To evaluate whether the proposed text-anchored residual fusion strategy generalizes beyond CMU-MOSI, we further conducted experiments on the CMU-MOSEI dataset [
35].
Table 16 reports the standard CMU-MOSEI validation results under the same metric definitions.
Under the standard CMU-MOSEI training/validation/test setting, UCMR-Net achieves MAE = 0.548, RMSE = 0.752, Corr = 0.792, Acc-2 = 85.80%, and F1 = 85.50%, outperforming the evaluated baselines in regression metrics.
Table 17 reports a stricter MOSI-to-MOSEI transfer setting. As expected, all models show degraded performance under cross-dataset transfer, but UCMR-Net remains the strongest among the compared models.
Figure 14 presents a visual comparison of the standard CMU-MOSEI validation results and the MOSI-to-MOSEI transfer results, highlighting both in-domain external validation and cross-dataset transfer degradation.
The visualization indicates that UCMR-Net remains competitive on CMU-MOSEI, whereas the MOSI-to-MOSEI transfer setting shows a larger performance degradation than the standard in-domain evaluation.
4.10. Error Distribution and Qualitative Case Analysis
A qualitative case analysis was conducted to examine when nonverbal residuals help or harm text-based prediction.
Table 18 summarizes representative subsets covering consistent cues, modality contradiction, strongly negative samples, near-neutral samples, and cases where residual correction hurts performance.
UCMR-Net reduces MAE most clearly when text, audio, and visual cues are consistent and for strongly negative samples, indicating that nonverbal residuals can strengthen or correct the text-based estimate. However, near-neutral boundary samples and residual-hurts cases remain difficult: in these subsets, nonverbal residuals may introduce additional ambiguity or over-correct the textual anchor.
Figure 15 visualizes delta MAE across qualitative subsets with signed bars, showing when residual fusion improves or degrades prediction performance.
The signed-bar pattern illustrates a practical model boundary: nonverbal residuals are beneficial in several cue-consistent or strongly negative cases, but they are not uniformly helpful across all sentiment conditions.
Table 19 summarizes the prediction error distribution under the five-seed evaluation protocol. The percentile-level absolute errors characterize the remaining sample-level difficulty and link the aggregate metrics with the qualitative error cases.
To complement the percentile-level error statistics,
Figure 16 presents the sample-level prediction analysis and absolute-error distribution under the five-seed evaluation protocol.
The visualization shows where remaining errors concentrate and provides a sample-level counterpart to the qualitative subsets in
Table 18.
4.11. Sentiment-Intensity Interval Analysis
To further examine whether the model behaves differently across sentiment regions, the test samples are divided into negative, neutral, positive, strong-negative, weak-negative, weak-positive, and strong-positive intervals.
Table 20 reports interval-wise MAE and binary consistency scores.
The interval analysis indicates that UCMR-Net is relatively stable in positive sentiment regions but still struggles with strongly negative sentiment intensity. The neutral category has the smallest number of samples and the lowest binary consistency, which is expected because polarity boundaries around zero are inherently ambiguous.
Figure 17 compares interval-wise MAE and binary consistency across sentiment regions, allowing the intensity-estimation and polarity-consistency patterns to be examined together.
The interval-wise visualization indicates that strong-negative and near-neutral regions remain important sources of residual error despite strong aggregate performance.
4.12. Summary of Experimental Findings
The experiments support four main conclusions. First, UCMR-Net achieves strong complete-modality regression performance on CMU-MOSI under a five-seed protocol, with MAE = 0.699 ± 0.009 and Corr = 0.802 ± 0.006 (mean ± SD). Second, the unified baseline comparison and paired tests provide the primary evidence for regression improvements over strong baselines; the Corr comparison with LNLN at p = 0.044 is treated as secondary evidence, and the binary polarity gains are not statistically significant. Third, the controlled ablation study shows that removing text anchoring causes the largest degradation, while removing residual fusion produces the second-largest degradation. Fourth, the unified missing-modality experiments show an asymmetric robustness pattern: the model remains comparatively stable when audio or visual streams are removed, but performance degrades substantially when text is unavailable.
Overall, the experimental evidence supports UCMR-Net as a text-anchored residual fusion model with strong complete-modality performance, moderate robustness under nonverbal missing or corrupted conditions, and empirically useful residual reliability analysis. This evidence does not claim full modality symmetry or formal probabilistic uncertainty estimation.
5. Discussion
The results show that UCMR-Net achieves stable complete-modality regression performance under the five-seed CMU-MOSI protocol. The unified baseline comparison reduces the likelihood that the observed performance differences are attributable to heterogeneous data splits, feature pipelines, or metric scripts. Paired tests provide the primary statistical support for the MAE improvements over strong baselines, whereas the Corr comparison with LNLN at p = 0.044 is treated as secondary evidence and the binary polarity differences are not statistically significant.
The most important empirical finding is the contribution of text-anchored residual fusion. The controlled ablation study confirms that removing text anchoring causes the largest MAE degradation, and removing residual fusion produces the second-largest degradation. This supports the design assumption that text can serve as the semantic backbone in opinion-level MSA, while acoustic and visual signals function as adaptive residual modifiers.
The missing-modality experiments clarify the scope of robustness. UCMR-Net remains comparatively stable when audio or visual streams are removed, but performance degrades substantially when text is unavailable. Therefore, the model is better described as a text-anchored residual fusion model than as a modality-symmetric missing-modality model. This pattern is consistent with the architecture and the observed missing-text degradation.
The calibration and reliability analyses indicate that raw UCMR-Net does not by itself establish comprehensive probabilistic calibration. However, temperature scaling substantially improves calibration metrics, and the learned residual reliability-risk score correlates with absolute prediction error and responds to modality corruption. These findings support the use of adaptive residual weighting as a practical modality-contribution regulation mechanism, while leaving formal uncertainty estimation as a direction for future probabilistic modeling.
External validation on CMU-MOSEI suggests that the text-anchored residual fusion strategy generalizes beyond CMU-MOSI, but the MOSI-to-MOSEI transfer setting remains challenging. The qualitative case analysis further shows that nonverbal residuals are helpful when multimodal cues are consistent or when strongly negative evidence is present, but they may hurt near-neutral boundary samples or cases where nonverbal evidence over-corrects the textual anchor. Together, these findings characterize the main strengths, robustness boundaries, and remaining error modes of the proposed framework.
6. Conclusions
This study proposed UCMR-Net, a text-anchored residual fusion framework for multimodal sentiment intensity prediction. The model uses contextual language representations as the primary semantic backbone and introduces acoustic and visual streams as adaptive residual correction signals. A learned residual reliability score is used to regulate nonverbal contribution, while counterfactual masked-modality distillation and multi-task supervision provide auxiliary constraints for prediction stability and polarity discrimination.
Under a unified five-seed CMU-MOSI protocol, UCMR-Net achieves MAE = 0.699 ± 0.009, RMSE = 0.997 ± 0.011, Corr = 0.802 ± 0.006, Acc-2 = 85.82 ± 0.62%, and F1 = 85.60 ± 0.62% (mean ± SD). Unified baseline comparisons and statistical tests show that the proposed model provides significant regression improvements over strong baselines, although gains in binary polarity metrics are more modest. Controlled ablations confirm that text anchoring and residual fusion are the dominant architectural contributors. Unified missing-modality results further show that the model is moderately robust when audio or visual signals are unavailable, but performance declines markedly when text is removed.
External validation on CMU-MOSEI, together with calibration, reliability-error, modality corruption, and qualitative case analyses, characterizes the model’s generalization behavior, reliability-related patterns, robustness boundaries, and sample-level error characteristics. Overall, UCMR-Net offers a practical and extensible framework for text-centered multimodal sentiment intensity prediction, while formal uncertainty estimation and fully modality-symmetric missing-modality robustness remain important directions for future research.