On the Out-Of-Distribution Generalization of Multimodal Large Language Models

Zhang, Xingxuan; Li, Jiansheng; Chu, Wenjing; Hai, Junjia; Xu, Renzhe; Yang, Yuqing; Guan, Shikai; Xu, Jiazheng; Cui, Peng

Computer Science > Computer Vision and Pattern Recognition

arXiv:2402.06599 (cs)

[Submitted on 9 Feb 2024]

Title:On the Out-Of-Distribution Generalization of Multimodal Large Language Models

Authors:Xingxuan Zhang, Jiansheng Li, Wenjing Chu, Junjia Hai, Renzhe Xu, Yuqing Yang, Shikai Guan, Jiazheng Xu, Peng Cui

View PDF HTML (experimental)

Abstract:We investigate the generalization boundaries of current Multimodal Large Language Models (MLLMs) via comprehensive evaluation under out-of-distribution scenarios and domain-specific tasks. We evaluate their zero-shot generalization across synthetic images, real-world distributional shifts, and specialized datasets like medical and molecular imagery. Empirical results indicate that MLLMs struggle with generalization beyond common training domains, limiting their direct application without adaptation. To understand the cause of unreliable performance, we analyze three hypotheses: semantic misinterpretation, visual feature extraction insufficiency, and mapping deficiency. Results identify mapping deficiency as the primary hurdle. To address this problem, we show that in-context learning (ICL) can significantly enhance MLLMs' generalization, opening new avenues for overcoming generalization barriers. We further explore the robustness of ICL under distribution shifts and show its vulnerability to domain shifts, label shifts, and spurious correlation shifts between in-context examples and test data.

Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2402.06599 [cs.CV]
	(or arXiv:2402.06599v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2402.06599

Submission history

From: Xingxuan Zhang [view email]
[v1] Fri, 9 Feb 2024 18:21:51 UTC (9,524 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:On the Out-Of-Distribution Generalization of Multimodal Large Language Models

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:On the Out-Of-Distribution Generalization of Multimodal Large Language Models

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators