Dynamic Gated Cross-Modal Fusion with Sarcastic-aware Contrastive Regularization for Multimodal Sarcasm Detection
文字和图片表面上看起来很搭,实则暗藏矛盾,这就是讽刺,这套AI能识别出来
识别图文帖子中讽刺意味的AI系统常常会被表面一致但实际含义相反的文字图片组合迷惑。这篇论文提出一个框架,能针对每条帖子动态判断文字和图片哪个更关键,并引入一种对比学习方法,把讽刺样本里的表面一致性当作陷阱而非真实证据来处理。在MMSD和MMSD2.0两个基准数据集上,该方法的表现持续超过现有的强基线模型。
他们做了什么
- 由于有些讽刺帖子主要靠文字线索,有些则靠图片矛盾,作者设计了一个动态门控融合模块,对文字和图片这两种模态的信息进行双向过滤,并针对每个样本单独调整两者的权重
- 讽刺类图文对常常字面上看起来一致,实际含义却相反,为此论文提出了SaCR(讽刺感知对比正则化),对非讽刺样本提高图文相似度、对讽刺样本则压低相似度,防止模型被表面一致性误导
- 使用预训练的视觉-语言模型CLIP提取文字和图像特征,再通过双向交叉注意力机制在数值层面加门控,过滤掉无用或有误导性的跨模态信号
- 整个模型通过多目标损失进行端到端训练,同时优化最终分类损失、文本/图像单模态辅助分类损失以及SaCR对比损失
- 在MMSD和MMSD2.0两个数据集上,该方法的F1分数均优于对比的基线模型;消融实验显示,去掉跨模态交互模块或动态融合门会导致性能下降最明显

| Dataset | Split | Total | Sarcastic | Non-sarcastic |
|---|---|---|---|---|
| MMSD | Train | 19,816 | 8,642 | 11,174 |
| Val | 2,410 | 959 | 1,451 | |
| Test | 2,409 | 959 | 1,450 | |
| MMSD2.0 | Train | 19,816 | 9,576 | 10,240 |
| Val | 2,410 | 1,042 | 1,368 | |
| Test | 2,409 | 1,037 | 1,372 |

| Modality | Method | MMSD | MMSD2.0 | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Acc.% | P% | R% | F1% | Acc.% | P% | R% | F1% | ||
| Text | TextCNN | 80.03 | 74.29 | 76.39 | 75.32 | 71.61 | 64.62 | 75.22 | 69.52 |
| SMSD | 80.90 | 76.46 | 75.18 | 75.82 | 73.56 | 68.45 | 71.55 | 69.97 | |
| BERT | 83.60 | 78.50 | 82.51 | 80.45 | 76.52 | 74.48 | 73.09 | 73.91 | |
| Image | ResNet | 64.76 | 54.41 | 70.80 | 61.53 | 65.50 | 61.17 | 54.39 | 57.58 |
| ViT | 67.83 | 57.93 | 70.07 | 63.40 | 72.02 | 65.26 | 74.83 | 69.72 | |
| Multimodal | DIP | 89.59 | 87.76 | 86.58 | 87.17 | 80.96 | 78.02 | 77.56 | 77.79 |
| Multi-view CLIP | 88.33 | 82.66 | 88.65 | 85.55 | 85.64 | 80.33 | 88.24 | 84.10 | |
| MoBA | 88.96 | 82.84 | 88.12 | 85.40 | 85.83 | 80.42 | 88.67 | 84.34 | |
| G2SAM | 90.48 | 87.95 | 89.02 | 88.48 | 79.43 | 72.04 | 78.07 | 78.07 | |
| TFCD | 89.57 | 84.83 | 89.43 | 88.13 | 86.54 | 82.46 | 87.95 | 84.31 | |
| DGLF | 89.43 | 85.81 | 89.27 | 87.51 | 86.82 | 81.90 | 89.85 | 85.69 | |
| LLaVA+RAG | 89.97 | 89.26 | 89.58 | 89.42 | 86.43 | 87.00 | 86.30 | 86.34 | |
| ESAM | 90.11 | 86.87 | 89.54 | 88.19 | 85.87 | 83.12 | 86.05 | 84.56 | |
| GPT-5.4 (zeroshot) | 71.05 | 76.51 | 75.50 | 71.01 | 72.85 | 78.79 | 75.78 | 72.55 | |
| Ours | 92.62 | 91.96 | 92.82 | 92.33 | 89.66 | 89.36 | 89.74 | 89.51 |

| Variant | MMSD | MMSD2.0 | ||
|---|---|---|---|---|
| Acc.% | F1% | Acc.% | F1% | |
| Full | 92.62 | 92.33 | 89.66 | 89.36 |
| w/o CMI | 87.91 | 87.42 | 85.10 | 84.97 |
| w/o BiXAtt (only v→t) | 87.48 | 87.03 | 81.86 | 81.70 |
| w/o BiXAtt (only t→v) | 81.84 | 81.07 | 80.53 | 80.37 |
| w/o VGate | 92.08 | 91.77 | 88.71 | 88.59 |
| w/o DFGate | 90.35 | 89.80 | 88.34 | 88.28 |
| w/o SaCR | 91.28 | 91.00 | 88.54 | 88.49 |
| w/o UniAux | 91.31 | 90.92 | 89.16 | 89.03 |
为什么重要
任何需要从社交媒体帖子中读取语气或情感的系统(如内容审核、舆情分析、聊天机器人应答)如果漏掉被表面一致文字图片掩盖的讽刺,就可能完全误判说话人的真实意图,因此专门针对这一失效模式的方法具有直接的实用价值。文字图片表面一致却可能掩盖矛盾意图这一发现,也可能对讽刺检测之外的其他多模态理解任务有参考意义。
本文术语
- 多模态讽刺检测(MSD) · 通过同时分析文字、图片等不同形式的内容来判断说话者是否在讽刺
- CLIP · 一种预训练的视觉-语言模型,能把文字和图像映射到同一个可比较的空间
- 门控(gate) · 一种可学习的机制,用0到1之间的数值控制信号通过的多少
- 对比正则化 · 一种辅助训练方法,让相似的表示彼此靠近,不相似的表示彼此远离
- Grad-CAM · 一种可视化技术,能显示模型做判断时主要依据图像的哪些区域
论文原文摘要(英文)
Multimodal sarcasm detection aims to identify sarcastic intent from multimodal content, where inconsistencies between literal meaning and contextual cues often signal irony. This task has attracted increasing research attention. However, accurate detection remains challenging due to instance-dependent modality contributions and misleading semantic consistency, where surface-level alignment masks underlying contradictory intent. Existing methods often rely on fixed fusion strategies and treat sarcasm as generic cross-modal mismatch, limiting their ability to capture subtle sarcasm cues and instance-specific modality interactions. To address these challenges, we propose a novel MSD framework that integrates Dynamic Gated Cross-Modal Fusion with Sarcastic-aware Contrastive Regularization (SaCR). Specifically, a bidirectional gated interaction module performs cross-modal feature filtering and adaptively calibrates textual and visual contributions at the instance level. A dynamic fusion gate further balances modality importance to generate more robust multimodal representations. Furthermore, SaCR is introduced as a label-aware contrastive regularization objective that encourages semantic consistency for non-sarcastic samples while suppressing misleading consistency in sarcastic cases. The proposed framework is trained end-to-end with a multi-objective learning strategy that jointly optimizes multimodal classification and auxiliary unimodal supervision. Extensive experiments on MMSD and MMSD2.0 demonstrate that the proposed method consistently outperforms strong baselines.
在 arXiv 阅读最新论文
- LoRA-GA$^2$: Low Rank Adaptation with Multi-step Gradient Adaptive Alignment在正式微调前先偷看几步训练的梯度,让LoRA的初始化更聪明
- Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM AgentsAI助手在该向你提问的时候,却更愿意自己去核实事实
- Robust Incomplete Multimodal Sentiment Analysis via Iterative Proxy Correction文本信息缺失或损坏时,这个AI不靠一次性猜测,而是反复修正猜测结果,从而更准确地判断情绪
- Generating Diverse Personas for User Simulators to Test Interview Dialogue Systems要测试访谈式对话系统需要大量不同性格的虚拟用户,这项研究用大语言模型自动生成这些虚拟用户人设
- Rethinking Patch Based Multivariate Time Series Forecasting with Semantic Structured Partitioning别再机械切分时间序列,按语义把它切成有意义的块
- Reliable Financial Named Entity Recognition under Domain ShiftAI在正式文件里学到的自信,一到推特上就变得不可信
- Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis让AI分析脑影像数据时,把“为什么这个结论可信”也一并记录下来
- GenMatch: An End-to-End Generative Matching Framework for Micro-View Order-Dispatching in Ride-Hailing滴滴把打车派单从预测-计算-匹配三段式流程改成一次生成完成,线上效果提升明显
METAL LAB 最新报道
图片来源: Hao Guo et al., arXiv:2608.19942, arxiv-nonexclusive