Robust Multimodal Large Language Models Against Modality Conflict

cs.AI updates on arXiv.org 07月11日 12:04

Robust Multimodal Large Language Models Against Modality Conflict

本文从模态冲突视角探讨MLLM在视觉-语言任务中的幻觉现象，构建了Multimodal Modality Conflict（MMMC）数据集，提出三种方法缓解模态冲突导致的幻觉，并通过实验验证了效果。

arXiv:2507.07151v1 Announce Type: cross Abstract: Despite the impressive capabilities of multimodal large language models (MLLMs) in vision-language tasks, they are prone to hallucinations in real-world scenarios. This paper investigates the hallucination phenomenon in MLLMs from the perspective of modality conflict. Unlike existing works focusing on the conflicts between model responses and inputs, we study the inherent conflicts in inputs from different modalities that place MLLMs in a dilemma and directly lead to hallucinations. We formally define the modality conflict and construct a dataset named Multimodal Modality Conflict (MMMC) to simulate this phenomenon in vision-language tasks. Three methods based on prompt engineering, supervised fine-tuning, and reinforcement learning are proposed to alleviate the hallucination caused by modality conflict. Extensive experiments are conducted on the MMMC dataset to analyze the merits and demerits of these methods. Our results show that the reinforcement learning method achieves the best performance in mitigating the hallucination under modality conflict, while the supervised fine-tuning method shows promising and stable performance. Our work sheds light on the unnoticed modality conflict that leads to hallucinations and provides more insights into the robustness of MLLMs.

Fish AI Reader

AI辅助创作，多种专业模板，深度分析，高质量内容生成。从观点提取到深度思考，FishAI为您提供全方位的创作支持。新版本引入自定义参数，让您的创作更加个性化和精准。

FishAI

鱼阅，AI 时代的下一个智能信息助手，助你摆脱信息焦虑

联系邮箱 441953276@qq.com

相关标签

MLLM 模态冲突幻觉现象数据集方法研究

相关文章

MS MARCO Web Search: A Large-Scale Information-Rich Web Dataset Featuring Millions of Real Clicked Query-Document Labels

This Week In Machine Learning & AI - 5/27/16: The White House on AI & Aggressive Self-Driving Cars

CinePile: A Novel Dataset and Benchmark Specifically Designed for Authentic Long-Form Video Understanding

‘RAG Me Up’: A Generic AI Framework (Server + UIs) that Enables You to Do RAG on Your Own Dataset Easily

HuggingFace Releases ? FineWeb: A New Large-Scale (15-Trillion Tokens, 44TB Disk Space) Dataset for LLM Pretraining

Unlocking the Language of Proteins: How Large Language Models Are Revolutionizing Protein Sequence Understanding

MAGPIE: A Self-Synthesis Method for Generating Large-Scale Alignment Data by Prompting Aligned LLMs with Nothing

Midjourney: ↩️ @kortizart To the best of our knowledge; you are not in our dataset. Here's a "portrait by Karla Ortiz" vs a "portrait by artist". FW...

Hugging Face: We're excited to welcome @argilla_io to the Hugging Face team! ? Time to democratise good Machine Learning, one dataset at a time!...

Hugging Face: Hugging Face is hosting a demo site for @iclr_conf authors to find and claim their papers and discuss those papers on dedicated pages Th...