Contrastive Visual-Linguistic Pretraining

Shi, Lei; Shuang, Kai; Geng, Shijie; Su, Peng; Jiang, Zhengkai; Gao, Peng; Fu, Zuohui; de Melo, Gerard; Su, Sen

Computer Science > Computer Vision and Pattern Recognition

arXiv:2007.13135 (cs)

[Submitted on 26 Jul 2020]

Title:Contrastive Visual-Linguistic Pretraining

Authors:Lei Shi, Kai Shuang, Shijie Geng, Peng Su, Zhengkai Jiang, Peng Gao, Zuohui Fu, Gerard de Melo, Sen Su

View PDF

Abstract:Several multi-modality representation learning approaches such as LXMERT and ViLBERT have been proposed recently. Such approaches can achieve superior performance due to the high-level semantic information captured during large-scale multimodal pretraining. However, as ViLBERT and LXMERT adopt visual region regression and classification loss, they often suffer from domain gap and noisy label problems, based on the visual features having been pretrained on the Visual Genome dataset. To overcome these issues, we propose unbiased Contrastive Visual-Linguistic Pretraining (CVLP), which constructs a visual self-supervised loss built upon contrastive learning. We evaluate CVLP on several down-stream tasks, including VQA, GQA and NLVR2 to validate the superiority of contrastive learning on multi-modality representation learning. Our code is available at: this https URL.

Subjects:	Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)
Cite as:	arXiv:2007.13135 [cs.CV]
	(or arXiv:2007.13135v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2007.13135

Submission history

From: Peng Gao [view email]
[v1] Sun, 26 Jul 2020 14:26:18 UTC (2,953 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.CV

< prev | next >

new | recent | 2020-07

Change to browse by:

cs
eess
eess.IV

References & Citations

DBLP - CS Bibliography

listing | bibtex

Lei Shi
Kai Shuang
Shijie Geng
Peng Su
Zhengkai Jiang

…

export BibTeX citation

Computer Science > Computer Vision and Pattern Recognition

Title:Contrastive Visual-Linguistic Pretraining

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Contrastive Visual-Linguistic Pretraining

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators