Robust Incomplete Multimodal Sentiment Analysis via Iterative Proxy Correction
文本信息缺失或损坏时,这个AI不靠一次性猜测,而是反复修正猜测结果,从而更准确地判断情绪
现实场景中的多模态情绪分析经常要处理缺失或受损的语音、视频、文本信息,这会拖累识别准确率。以往方法在文本信息不足时,会用音频视频信息生成一个替代表示(代理),但只生成一次就直接投入使用,初期误差会一路传导下去。这篇论文提出让这个代理表示分多个阶段逐步修正,在MOSI、MOSEI、SIMS三个数据集上都比此前最强基线LNLN表现更稳定、更好。
他们做了什么
- 问题背景:现实中文本、音频、视频信息常常缺失或带噪声,导致情绪分析模型判断不准
- 已有方法的局限:文本信息不足时会用音频视频信息生成'代理'来代替文本,但只生成一次就立刻使用,导致初始错误被放大并传递下去
- 提出的方法:先只用音频和视频信息生成初步代理表示,再通过带门控机制的残差修正方式,分多个步骤逐步完善这个代理
- 修正后的代理会根据对观测到的文本信息可信度的估计分数,与文本表示自适应地融合后再用于最终判断
- 训练阶段用完整无损的文本表示作为语义锚点,引导每一步修正朝正确方向进行
- 结果:在MOSI、MOSEI、SIMS三个基准上,该方法整体优于或持平于此前最强基线LNLN,且在信息缺失程度加重时性能下降更平缓
| Method | MOSI | MOSEI | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Acc-7 | Acc-5 | Acc-2 | F1 | MAE | Corr | Acc-7 | Acc-5 | Acc-2 | F1 | MAE | Corr | |
| MISA | 29.85 | 33.08 | 71.49/70.00 | 71.28/70.33 | 1.085 | 0.524 | 40.84 | 39.39 | 71.27/75.82 | 63.85/68.73 | 0.780 | 0.503 |
| Self-MM | 29.55 | 34.67 | 70.51/69.26 | 66.60/67.54 | 1.070 | 0.512 | 44.70 | 45.38 | 73.89/77.42 | 68.92/72.31 | 0.695 | 0.498 |
| MMIM | 31.30 | 33.77 | 69.14/67.06 | 66.65/64.04 | 1.077 | 0.507 | 40.75 | 41.74 | 73.32/75.89 | 68.72/70.32 | 0.739 | 0.489 |
| CENET | 30.38 | 37.25 | 71.46/67.73 | 68.41/64.85 | 1.080 | 0.504 | 47.18 | 47.83 | 74.67/77.34 | 70.68/74.08 | 0.685 | 0.535 |
| TETFN | 30.30 | 34.34 | 69.76/67.68 | 65.69/63.29 | 1.087 | 0.507 | 40.30 | 47.70 | 69.76/67.68 | 65.69/63.29 | 1.087 | 0.508 |
| ALMT | 30.30 | 33.42 | 70.40/68.39 | 72.57/71.80 | 1.083 | 0.498 | 40.92 | 41.64 | 76.64/77.54 | 77.14/78.03 | 0.674 | 0.481 |
| LNLN | 34.26 | 38.27 | 72.55/70.94 | 72.73/71.25 | 1.046 | 0.527 | 45.42 | 46.17 | 76.30/78.19 | 77.77/79.95 | 0.692 | 0.530 |
| Ours | 34.52 | 38.55 | 73.27/71.93 | 73.03/71.93 | 1.046 | 0.532 | 47.10 | 47.94 | 78.20/78.94 | 78.46/80.00 | 0.664 | 0.595 |

| Method | Acc-5 | Acc-3 | Acc-2 | F1 | MAE | Corr |
|---|---|---|---|---|---|---|
| MISA | 31.53 | 56.87 | 72.71 | 66.30 | 0.539 | 0.348 |
| Self-MM | 32.28 | 56.75 | 72.81 | 68.43 | 0.508 | 0.376 |
| MMIM | 31.81 | 52.76 | 69.86 | 66.21 | 0.544 | 0.339 |
| CENET | 22.29 | 53.17 | 68.13 | 57.90 | 0.589 | 0.107 |
| TETFN | 33.42 | 56.91 | 73.58 | 68.67 | 0.505 | 0.387 |
| ALMT | 20.00 | 45.36 | 69.66 | 72.76 | 0.561 | 0.364 |
| LNLN | 34.64 | 57.14 | 72.73 | 79.43 | 0.514 | 0.397 |
| Ours | 35.13 | 58.28 | 73.05 | 76.29 | 0.498 | 0.404 |
| Method | MOSI | MOSEI | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Acc-7 | Acc-5 | Acc-2 | F1 | MAE | Corr | Acc-7 | Acc-5 | Acc-2 | F1 | MAE | Corr | |
| T | 45.58 | 51.70 | 84.91 / 82.75 | 84.84 / 82.68 | 0.731 | 0.790 | 52.39 | 53.81 | 85.75 / 84.46 | 85.7 / 84.14 | 0.548 | 0.769 |
| A | 22.84 | 23.08 | 58.49 / 57.38 | 61.27 / 59.64 | 1.371 | 0.280 | 41.38 | 41.38 | 63.84 / 71.28 | 51.85 / 59.93 | 0.830 | 0.152 |
| V | 22.98 | 24.83 | 58.59 / 56.80 | 51.78 / 53.90 | 1.376 | 0.230 | 42.46 | 42.46 | 65.30 / 71.02 | 61.06 / 58.99 | 0.811 | 0.244 |
| T+A | 45.53 | 51.31 | 85.16 / 83.19 | 84.97 / 83.07 | 0.739 | 0.790 | 52.29 | 53.68 | 85.77 / 84.65 | 85.73 / 84.20 | 0.549 | 0.764 |
| T+V | 45.40 | 51.31 | 84.86 / 82.60 | 84.70 / 82.56 | 0.737 | 0.790 | 52.48 | 53.96 | 86.08 / 84.85 | 85.99 / 84.85 | 0.545 | 0.768 |
| A+V | 22.40 | 24.30 | 59.25 / 58.31 | 56.45 / 57.04 | 1.350 | 0.198 | 42.52 | 42.52 | 65.38 / 71.22 | 60.40 / 60.12 | 0.812 | 0.244 |
| Random | 34.52 | 38.55 | 73.27/71.93 | 73.03/71.93 | 1.046 | 0.532 | 47.10 | 47.94 | 78.20/78.94 | 78.46/80.00 | 0.664 | 0.595 |

| Method | MOSI | SIMS | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Acc-7 | Acc-5 | Acc-2 | F1 | MAE | Corr | Acc-5 | Acc-3 | Acc-2 | F1 | MAE | Corr | |
| FULL | 34.52 | 38.55 | 73.27/71.93 | 73.03/71.93 | 1.046 | 0.532 | 35.13 | 58.28 | 73.05 | 76.29 | 0.498 | 0.404 |
| w/o Proxy | 34.53 | 38.57 | 72.29 / 71.89 | 72.29 / 71.84 | 1.053 | 0.527 | 30.93 | 54.98 | 70.05 | 62.08 | 0.569 | 0.238 |
| w/o Iterative | 34.27 | 38.32 | 73.19 / 71.84 | 72.37 / 71.69 | 1.051 | 0.536 | 30.98 | 55.23 | 70.58 | 69.74 | 0.565 | 0.241 |
| w/o ℒcorr | 34.23 | 38.25 | 72.37 / 71.29 | 71.48 / 70.45 | 1.051 | 0.529 | 31.05 | 55.03 | 70.09 | 62.12 | 0.570 | 0.236 |
为什么重要
语音识别误差、网络不稳定、隐私限制等原因导致部分信息缺失的情况在实际应用中很常见。让情绪分析在这种不完整输入下依然稳定可靠,对客服中心、评论分析、对话式AI等场景中做出可信判断很重要。
本文术语
- 多模态情绪分析(MSA) · 结合文本、音频、视频等多种信息来推断情绪状态的任务
- 代理(proxy) · 为替代缺失或受损信息,用其他信息生成的辅助表示
- 门控残差修正 · 通过门控机制控制修正幅度,逐步在原有基础上添加修正量的方法
- 可信度分数(reliability score) · 衡量观测到的文本信息可信程度的0到1之间的数值
- MAE、F1、Acc · 分别代表预测误差大小、分类平衡性指标、以及预测准确率的评估指标
论文原文摘要(英文)
Multimodal sentiment analysis aims to infer affective states by integrating language, visual, and acoustic cues. However, real-world multimodal inputs are often incomplete or corrupted, which can weaken cross-modal complementarity and introduce misleading information into downstream fusion. Existing proxy-based methods for incomplete MSA commonly rely on one-shot proxy construction to compensate for degraded language information, but the generated proxy may be coarse or unreliable at initialization. Prematurely injecting such a proxy into multimodal reasoning can propagate initial errors and compromise sentiment prediction. To address this limitation, we propose an iterative proxy correction framework for robust incomplete MSA. Our method constructs a language-oriented proxy from non-language modalities and progressively refines it under multimodal context through gated residual correction. The corrected proxy is then adaptively fused with the observed language representation according to an estimated language reliability score, allowing the model to balance proxy-based compensation and trustworthy linguistic evidence. In addition, we introduce a stage-wise latent correction objective that uses the complete language representation as a training-time semantic anchor to stabilize the proxy refinement trajectory. Extensive experiments on MOSI, MOSEI, and SIMS under diverse missing-modality settings demonstrate that the proposed framework consistently outperforms competitive baselines and achieves robust sentiment prediction under incomplete inputs.
在 arXiv 阅读最新论文
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?让AI编程助手去修复真实科学软件,连最强的那个也有一半以上任务没做对
- When Saying No Makes Better Videos: Designing Dual Gatekeeping for Pedagogically Grounded AI Content Creation会说'不行'的AI才能做出更好的教学视频
- LoRA-GA$^2$: Low Rank Adaptation with Multi-step Gradient Adaptive Alignment在正式微调前先偷看几步训练的梯度,让LoRA的初始化更聪明
- TESTNAV: Pareto-Guided Search for Compositional Robustness Testing测试AI模型面对多种叠加干扰时不必穷举所有组合,也能找出真正危险的失败案例
- Natural Language Code Retrieval for 1C:Enterprise: An Open Benchmark and Efficient Bi-Encoder首个用俄语提问就能搜索1C企业软件代码的公开基准和专用AI模型问世
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving把稀疏注意力从论文原型变成能真正上线服务的加速方案
- NepOOC-M: Bilingual Nepali-English Benchmark and Comparative Analysis of Multimodal Architectures for OOC Detection尼泊尔语假新闻检测:只看文字就能追平图文结合模型
- Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM AgentsAI助手在该向你提问的时候,却更愿意自己去核实事实
METAL LAB 最新报道
图片来源: Zhifa Geng et al., arXiv:2608.19971, arxiv-nonexclusive