Exploiting Auxiliary Caption for Video Grounding

Li, Hongxiang; Cao, Meng; Cheng, Xuxin; Zhu, Zhihong; Li, Yaowei; Zou, Yuexian

Computer Science > Computer Vision and Pattern Recognition

arXiv:2301.05997 (cs)

[Submitted on 15 Jan 2023 (v1), last revised 24 Mar 2024 (this version, v3)]

Title:Exploiting Auxiliary Caption for Video Grounding

Authors:Hongxiang Li, Meng Cao, Xuxin Cheng, Zhihong Zhu, Yaowei Li, Yuexian Zou

View PDF HTML (experimental)

Abstract:Video grounding aims to locate a moment of interest matching the given query sentence from an untrimmed video. Previous works ignore the {sparsity dilemma} in video annotations, which fails to provide the context information between potential events and query sentences in the dataset. In this paper, we contend that exploiting easily available captions which describe general actions, i.e., auxiliary captions defined in our paper, will significantly boost the performance. To this end, we propose an Auxiliary Caption Network (ACNet) for video grounding. Specifically, we first introduce dense video captioning to generate dense captions and then obtain auxiliary captions by Non-Auxiliary Caption Suppression (NACS). To capture the potential information in auxiliary captions, we propose Caption Guided Attention (CGA) project the semantic relations between auxiliary captions and query sentences into temporal space and fuse them into visual representations. Considering the gap between auxiliary captions and ground truth, we propose Asymmetric Cross-modal Contrastive Learning (ACCL) for constructing more negative pairs to maximize cross-modal mutual information. Extensive experiments on three public datasets (i.e., ActivityNet Captions, TACoS and ActivityNet-CG) demonstrate that our method significantly outperforms state-of-the-art methods.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2301.05997 [cs.CV]
	(or arXiv:2301.05997v3 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2301.05997

Submission history

From: Hongxiang Li [view email]
[v1] Sun, 15 Jan 2023 02:04:02 UTC (1,121 KB)
[v2] Tue, 28 Mar 2023 10:49:59 UTC (1,346 KB)
[v3] Sun, 24 Mar 2024 05:46:10 UTC (1,918 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Exploiting Auxiliary Caption for Video Grounding

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Exploiting Auxiliary Caption for Video Grounding

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators